The Applied AI Summit, which began in 2020 as the NLP Summit, has always been the conference for practitioners: the state of the art in industry, as a complement to the conferences that present the academic state of the art. For six years, people presented variants of how to build useful things: fine-tuning deep learning models, optimizing RAG pipelines, scaling real-world systems at a reasonable cost.
This year I was surprised to find that about two thirds of the submissions were about how to validate and test AI systems. They came from every industry and from companies of every size.
Building now takes weeks but validation still takes quarters
Developing software has become far easier and faster. You can now build complex software without experts in the minutiae of front-end libraries, DevOps tools, or back-end scaling and partitioning, which is what developers used to specialize in. One engineer can build an end-to-end, scalable, industry-specific customer service platform, HR platform, or medical decision support system in a few weeks.
One of this year’s keynotes, on the infrastructure behind production agents from ThoughtSpot, describes the same speed for AI agents: adding one takes a configuration entry and weeks of effort, where it used to take a codebase and quarters.
What remains hard is validating what was built: someone still has to write the test cases, a domain expert still has to say what a correct answer is, and a person still has to read the failures. A 2026 economics working paper describes the two curves: the cost to automate a task keeps falling, and the cost to verify the output is limited by how much people can check.
A summit session on comprehension debt describes the same gap inside engineering teams: AI generating code, prompts, and integrations faster than engineers can understand or validate them.
Letting the AI check itself does not close the gap. A BlackRock session on governing AI-assisted software development covers AI writing both the code and its tests from the same flawed assumptions, so defects pass validation.
Someone still has to show that the system won’t make a catastrophic financial or legal mistake. An insurance underwriting platform must not start writing policies that will bankrupt the insurer, or refuse to write them because of illegal bias. A customer service agent must not anger customers, or give away enough refunds to make the company unprofitable.
Mistakes are already common in production, because organizations seem to release AI systems without testing them enough. In McKinsey’s 2025 survey of 1,993 people, 51% of those whose organizations use AI had seen at least one negative consequence from it. Nearly one-third of all respondents had seen one that came from AI being inaccurate.
This is the pattern among AI early adopters in industry this year: the software development team spends most of its time being the QA team. (Yep, it’s not just you.) It’s the reverse of what was hard about building software until very recently. The summit’s theme became show your work and the proof it asks for has to come from your own team. The 50+ sessions will stream online from October 13 to 15 and are free to attend.
Three kinds of borrowed proof that don’t cover your system
The natural shortcut is to borrow the proof: assume that the published studies, the model’s benchmark scores, or the vendor’s own validation is enough. However, each one describes something other than the system you are running. The evidence, mostly from healthcare:
Published studies rarely involve patients. Of 4,609 studies of large language models (LLMs) in clinical medicine published from 2022 to 2025, 23% used patient data, and 19 studies, or 0.4%, were prospective randomized trials, the kind that follows patients forward and compares outcomes with and without the tool.
Benchmarks rarely say whether a gap between two models is larger than chance. A review of 445 LLM benchmarks found that 16% ran any statistical test, so a two-point lead on a leaderboard may be noise.
A vendor’s number reflects only the vendor’s test data. AUC scores a risk model from 0.5, a coin flip, to 1.0, a perfect ranking of patients. When one health system tested a widely used sepsis model on 38,455 of its own hospitalizations, it measured 0.63, against the 0.76 to 0.83 its developer reported.
Among US hospitals that use predictive models, 61% check them for accuracy on their own data and 44% check them for bias, according to a Health Affairs analysis of 2,425 hospitals. The rest run on borrowed proof: they put software into production that nobody tested on their own patients.
Five kinds of proof from your own team
If the proof can’t be borrowed, the team that deploys has to produce its own. There are five things to show, and this year’s summit has sessions on each.
1. The system finishes the whole job
A score on individual steps overstates what users get. Errors multiply: an agent that is right 90% of the time at each of ten steps completes the whole task about 35% of the time.
HealthAdminBench, a set of 135 healthcare administration workflows such as prior authorizations and appeals, shows the size of the gap. The best score any agent reached on individual steps was 83%, and the best rate of completing a workflow from start to finish was 36%.
On PhysicianBench, 100 tasks adapted from specialist consultations that take an average of 27 tool calls each, the best of 13 agents succeeded 46% of the time. Both results come from a keynote on extending MedHELM from single tasks to multi-step workflows, from Stanford Health Care and Pacific AI.
A second gap opens when a person joins the loop. In a randomized study of 1,298 people, LLMs working alone identified the right medical condition in 94.9% of cases. People using those same LLMs got there in about a third of cases, no better than the control group. The model’s score was accurate, and the person and the model together still failed.
A session on real-world evaluation of AI agents, from Expedia Group, starts from the same observation: accuracy, precision, and recall, the standard scores for a model, often can’t tell you whether an agent helped a user reach a goal. It sets out how to measure task completion, decision quality, and user satisfaction.
2. The answer was reached without shortcuts
In regulated settings an answer you can’t trace to evidence is an answer you can’t defend, and a system that is right by luck on one case will be wrong on the next one that looks the same.
A test built on clinical audits shows how much luck is in a score. Sixteen AI agents reviewed 750 patient cases built from hospital records and answered audit questions with yes, no, or one of two ways of saying the record can’t settle it. Scored on the answer alone, they were right 65.3% to 76.1% of the time.
When an answer counted only if the agent had taken none of the shortcuts a board of clinicians had ruled out, every score fell by 4.8 to 14.8 points, and the ranking of the agents changed. The results come from a keynote on CliniCARE-Bench, built and contributed as open source by Scale AI.
Every agent also committed to yes or no more readily than the records justified. For an audit that is the dangerous direction: a confident answer gets acted on, and an undetermined one gets a second look.
A keynote on evaluating generative models for financial research, from BlackRock, applies the same idea to analyst reports. Each claim is checked against the source it cites, which catches an existing filing cited for something it does not say. The model’s stated confidence is compared with how often it was right, and the results decide which outputs go to an analyst for review.
3. A clean test could have caught a problem
A clean result is evidence only if the test was large enough to catch a problem. However, five in six of the 445 benchmarks in the review above gave a score with no statistical test behind it.
A keynote on auditing a hiring AI shows what a large enough test looks like. Opptly scores job candidates with a model it built, and New York City’s Local Law 144 requires a bias audit by someone independent of the developers. The audit created 148,940 pairs of resumes, identical except for one marker of identity such as a name or a graduation year, across 43 protected groups, and compared the scores within each pair.
A first run on 10% of the data showed what looked like bias among the highest scores. The full run showed it was noise in a small number of extreme cases, and every group was selected between 91% and 92% of the time. A small sample can show a pattern that isn’t there and can miss one that is, so a result needs its sample size and margin of error beside it.
The same reasoning sets the size of PacificMed, a set of clinical safety and fairness benchmarks derived from 201 legal instruments, which my team at Pacific AI built. Each finished benchmark has 800 or 1,155 cases, sized to report a score to within about 3 to 3.5 points.
4. The system works where you deploy it
A deployment has its own languages, specialties, and data. An average is dominated by the cases that were easy to collect, and failures cluster in the ones that weren’t.
Only 12.8% of prompts got the same safety label in every language from the same model, in a test of 10 leading LLMs on 6,000 prompts about caste, religion, gender, health, and politics in 12 Indic languages that more than 1.2 billion people speak. The result comes from a keynote on IndicSafe, a multilingual safety evaluation, from Oracle. If your safety testing was done in English, it only describes your English users.
The same is true inside one hospital. A keynote on clinical NLP that clinicians will use, from Sheba Medical Center, covers how an aggregate score hides the departments where a model fails, and how measuring each specialty separately changes what gets deployed. It is true between hospitals too, which is what the sepsis model’s 0.63 showed.
5. The system still works today
A model you rent from a provider can change with no change to your code, so a validation describes the system on the day you ran it.
Researchers at Stanford and UC Berkeley measured this on GPT-4: its accuracy at identifying prime numbers was 84.0% in March 2023 and 51.1% in June 2023, under the same product name. Few systems are built to notice. A session reviewing 49 published studies of healthcare AI agents, from a data scientist at Cigna, found that 48 of them (!) had no mechanism for detecting or responding to drift.
A keynote on governing AI that changes after deployment, from Hologic and NNIT, argues that an AI system must be kept in a validated state after launch. It proposes three controls: predefined change boundaries, continuous monitoring for drift, and targeted reassessment.
A keynote on continuous AI governance reruns the release test in live use and raises an alert when the score falls, for example when a fairness test that scored 0.80 before release drops more than 5%. A session on fail-closed drift containment, from the Florida Institute of Technology, replays a fixed set of probe questions against the hosted model and stops sending it work when the answers change.
Both need a log of what the model did, and the open-source OpenTelemetry conventions for generative AI give you a shared format for one.
Proof at lower cost through risk tiers and shared tests
The strongest argument against all of this is cost: a single round of one academic leaderboard for AI agents cost about $40,000 to run. Here are two ways to lower the bill.
The first is matching the depth of testing to the risk. A keynote on reviewing AI tools before they reach practice, from Mayo Clinic, describes one review standard for vendor products and tools built in-house, with the depth of evaluation set by a risk tier. It draws on frameworks that have brought more than 100 software innovations into practice at Mayo, and it takes up when a vendor’s own evidence is enough.
The second is sharing the tests. The clinical benchmarks in this post, HealthAdminBench, PhysicianBench, CliniCARE-Bench, and PacificMed, are all being contributed to MedHELM, the open-source project for evaluating medical AI, so the next team does not write them again. Outside healthcare, Inspect, an open-source evaluation framework from the UK AI Security Institute and Meridian Labs, does the same job.
Regulators are easing their own requirements. The HTI-5 proposed rule would remove 34 of the 60 certification criteria for health IT and drop the model card requirements for AI, the standard disclosure of how a model was built and tested. The EU delayed its main high-risk AI obligations from August 2026 to December 2027.
Neither removes the question a customer, a board, or a patient asks after something goes wrong, which is how you knew the system worked. When a mandated test goes away, the testing moves to whoever deploys.
Five steps to take before your next AI release
Each of the five kinds of proof turns into one step you can take on a system you already run:
Measure the whole task with the people who use it, and report that number beside the model’s score.
Score how each answer was reached, and count the cases where the system should have declined to answer.
State how large a problem your test could have detected before you report that it found none.
Break results out by language, specialty, and site, and test on your own data before you trust a vendor’s figure.
Rerun your release tests in production on a schedule, and decide in advance what happens when a score drops.
The Applied AI Summit: Free & online from October 13 to 15
Every session linked above is part of the Applied AI Summit, which is free, online, and in its seventh year. Each day starts at 12:00 pm Eastern, every session stays available on demand for two weeks, and the schedule has the times.
Pick the one of the five you could not show for your own system today and start with the session on it. Register for all three days of the Applied AI Summit.
Frequently asked questions
What does “show your work” mean for an AI system?
It means producing your own evidence that the system works, is safe, and follows the rules it is subject to, in a form someone else can check. This post breaks that into five things: the whole task gets finished, the answer was reached without shortcuts, a clean test could have caught a problem, the system works in your languages and settings, and it still works after release.
Why can’t I rely on a model’s benchmark scores or a vendor’s validation?
Because they describe a different system from the one you run. Of 4,609 published studies of LLMs in clinical medicine, 19 were prospective randomized trials, and a widely used sepsis model scored an AUC of 0.63 at one health system against 0.76 to 0.83 reported by its developer. Testing on your own data is the only measurement of your deployment.
How often should a deployed AI system be retested?
On a schedule, and whenever the model, the prompt, or the rules change. GPT-4’s accuracy on one task fell from 84.0% to 51.1% between March and June 2023 under the same product name, and 48 of 49 published healthcare agent studies had no drift detection.
Are the Applied AI Summit sessions available to watch on demand?
Yes. If you register, you can watch each session on the summit’s online platform about 12 hours after it streams, and it stays available for two weeks.
Does passing an evaluation or audit mean an AI system complies with a regulation?
No. A benchmark score or an audit result is evidence that a governance program can present to a regulator, an auditor, or a customer. What it means for a specific deployment under a specific law is a question for your own compliance counsel.



