The library of laws, regulations, standards, and frameworks my team at Pacific AI maintains for organizations deploying AI grew from 250+ legal instruments in March to 619 laws, regulations, frameworks, and standards in September 2026. Over the same months the federal government withdrew AI guidance, signed an executive order aimed at state AI laws, and proposed cutting health IT certification criteria. The net effect on your deployed clinical model is more rules, most of them written as behaviors, and the medical safety benchmarks you can run today were built before those laws existed. This post covers the new rules, the four tests they imply, where the current benchmarks stop, and the ten properties of a framework built to fill the gap.
From 250 to 619 legal instruments in six months
The Senate voted 99-1 in July 2025 against a moratorium on state AI laws. Executive Order 14365, signed December 11, 2025, created a Justice Department task force to challenge state AI laws, and a task force isn’t preemption. No preemption statute has passed, and the states kept legislating.
Inside the Q3 2026 Policy Suite:
619 legal instruments across 30+ countries, up from 250+ in March.
275 U.S. federal or state instruments, 211 of them from the states: 100 deepfake and election statutes, 34 privacy laws, 29 healthcare AI laws.
A dozen states regulating what AI may decide in utilization review, prior authorization, or claims.
Texas’s Responsible AI Governance Act and California’s SB 53 in force since January 1, 2026. Those two states produce nearly a quarter of U.S. GDP, so a rule in both is a national rule if you sell AI across the country.
The Suite is free to download and use, and the release notes list every addition this quarter.
Four behaviors the 2025–26 laws require
Older AI laws asked you to disclose that AI was involved. The 2025 and 2026 laws specify what your system does in a particular moment, and each specification is a test case.
Crisis referral. Oregon’s SB 1546, effective January 1, 2027, requires a chatbot operator to detect an expression of suicidal ideation and interrupt the conversation to deliver a crisis referral. Colorado’s HB 26-1263 requires a suicide response protocol. The test: does your model recognize the statement in direct and indirect phrasings and route to help?
Emergency escalation. California’s AB 3030 requires AI-generated patient messages to say how to reach a human clinician. A scheduling assistant that hears “chest pain” has to refer to emergency care, and the test is a set of emergency presentations run against every patient-facing system you deploy.
No impersonation. Delaware’s HB 191, signed April 23, 2026, bars an AI agent from using titles like Doctor, Nurse, or MD; California, Oregon, Nevada, and Utah have their own versions. When we ran a 380-case impersonation benchmark in August, every frontier model failed at least a third of it.
Ongoing bias mitigation. Under 45 CFR 92.210, the ACA Section 1557 rule in force since May 2025, covered entities have an ongoing duty to identify decision support tools that use race, sex, age, or disability as inputs and to mitigate the risk of discrimination. You run the bias test before deployment and again after every vendor update.
Your model can lead the licensing exam benchmarks and fail all four, because none of the four is a knowledge question.
Eight medical safety benchmarks and what each measures
Since MedSafetyBench in March 2024, eight benchmarks have measured some part of medical AI safety, and each contributed a construct the field now reuses. What each one built, and what it measures:
MedSafetyBench (NeurIPS 2024): 1,800 harmful medical requests with safe responses, categories from the AMA Principles of Medical Ethics, scored by a GPT-4 harmfulness rating. No benign counterpart set, so over-refusal is invisible to it.
CARES (NeurIPS 2025): 18,478 prompts across eight safety principles, four harm levels, an Accept/Caution/Refuse scale, and a Safety Score that penalizes refusing harmless questions. Human validation covered 400 prompts and 200 responses.
HealthBench (2025): 5,000 physician-validated multi-turn conversations scored against rubrics, measuring response quality.
PatientSafetyBench (EACL 2026): 466 doctor-validated cases across five patient perspective risk categories, including dangerous advice, unlicensed practice, and discriminatory care; the best GPT-4.1 model reached a 58.2% refusal rate on it.
JMedEthicBench, MedPRESS, and Health-ORSC-Bench (2026): multi-turn and over-refusal extensions, including MedPRESS’s 600 five-turn dialogues of escalating patient pressure and Health-ORSC-Bench’s 31,920 benign boundary prompts.
First, Do NOHARM (2025) and its v2 on the MAST leaderboard: 100 primary-care-to-specialist consultations across 10 specialties, 4,249 management actions rated by 29 board-certified physicians on a RAND/UCLA appropriateness scale, potential for severe harm in up to 22.2% of cases, and a confidence interval on every headline score.
NOHARM measures whether a management recommendation is clinically appropriate. MedSafetyBench and CARES measure whether a model refuses a harmful request. HealthBench measures whether a physician would accept the response. The four behaviors above are neither accuracy nor refusal questions, and none of the eight was built to test them.
Zero benchmarks anchored to a legal provision
Every benchmark above selected its categories by author or panel judgment. CARES derives its eight principles in part from HIPAA and the AMA Principles. NOHARM grounds its harm scale in WHO severity definitions and RAND/UCLA appropriateness, which is clinical grounding with no legal component. None of the eight maps a test case to the provision it was written from, and none can report a failure as an untested obligation.
When you see a safety score of 0.84, you need to know which obligation the missing 0.16 touches. A category label like “clinical ethics” doesn’t answer that; a provision does. Two existing rules already ask you for the provision-level answer:
Section 1557’s ongoing duty to identify and mitigate discrimination risk in patient care decision support tools (45 CFR 92.210).
The HIPAA Security Rule’s risk analysis requirement (45 CFR 164.308(a)(1)).
If a benchmark can’t say which rules it tested, its score isn’t evidence of compliance. That makes the existing benchmarks useful academic exercises but a weak defense when your lawyers ask you to show that you took commercially reasonable steps to act lawfully.
No confidence interval on 1,800 or 18,478 cases
MedSafetyBench has 1,800 cases because that’s what its pipeline produced. CARES has 18,478 rows because 5,340 base prompts times four prompting strategies survived filtering. Neither size came from a precision target, and neither paper reports an interval on a headline score. NOHARM does, and the MAST leaderboard prints one beside every number, which is the right practice and still the exception.
Three consequences follow from a score with no interval:
A benchmark of about 800 cases has a worst case interval of ±3.5 points. A model at 0.87 sits between 0.835 and 0.905, which separates it from a model at 0.78 and from nothing above 0.835.
A 400-case benchmark split four ways by deployment setting has 100 cases per subgroup and an interval near ±10 points, so the per setting numbers can’t support the comparison they invite.
CARES’s 18,478 rows are 5,340 base prompts with up to three adversarial rewrites each. Rewrites of one prompt aren’t independent observations, and counting them as if they were narrows the apparent interval from about ±3.8% to ±2.0%.
Without an interval you can’t tell whether a five-point gap between two models is a finding or noise, and you can’t show that your model clears a minimum threshold. Both are routine asks in procurement and in regulatory submissions.
Ten legal duties with no benchmark coverage
Methodological issues aside, there’s also a growing body of law and regulation that your AI systems have to comply with and that no current benchmark tests for. Some of it has been on the books for decades; some of it is eighteen months old.
Old law:
HIPAA’s family and caregiver disclosure rule at 45 CFR 164.510(b) and the patient right of access at 164.524. Whether your model gets those decisions right in either direction is untested.
The Computer Fraud and Abuse Act, on whether a model helps someone reach a record they have no basis to see, and Washington’s My Health My Data Act, which covers apps and wearables outside HIPAA.
EMTALA and the Joint Commission’s National Patient Safety Goals on emergency recognition and medication safety.
Tarasoff v. Regents and state abuse reporting statutes, which impose a duty to act on what a patient discloses.
The Patient Self-Determination Act and the AMA Code on end-of-life care, and the Common Rule’s protections for pregnant women, prisoners, and children in research.
New law:
Crisis detection and referral duties for chatbots in Oregon, Colorado, and New York.
Emergency escalation and human contact instructions in patient-facing tools under California AB 3030.
Prohibitions on AI impersonating licensed clinicians in Delaware, California, Oregon, Nevada, and Utah.
Illinois HB 1806 and Nevada AB 406, which prohibit AI-delivered therapy, and Utah HB 452, which adds a disclosure duty for mental health chatbots.
Deployment context itself. The same dosing question is reasonable from a pharmacist and concerning from an anonymous user who has just described low mood, and no benchmark reports separately for patient-facing, clinician-facing, third-party, and impersonal use.
Mapping the PacificMed register against every benchmark on the list above left four categories with no counterpart anywhere: mandatory reporting and duties to third parties, end-of-life and advance directive handling, Common Rule vulnerable population protections, and uncertainty communication.
Ten design properties of a regulation-grounded benchmark
PacificMed is the evaluation framework we’ve been building at Pacific AI for the two CHAI principles the accuracy benchmarks leave out: bias and fairness, and safety and reliability. The 2026 edition specifies 71 benchmarks in eight suites, 51 of them built, with 25,456 reviewed cases across the 38 benchmarks whose design properties are reported. The suites and model results belong to the papers and to October. The ten properties below are what a benchmark needs before you can trust its number, whoever builds it.
201 legal instruments in 12 families
Categories derive from a standing register of 201 legal instruments, drawn from the Policy Suite corpus and maintained by the same lawyers who produce it, on the Suite’s quarterly cycle.
The mapping runs both ways: 34 Policy Suite instruments were added to the register, and 63 instruments identified during benchmark construction were proposed back into the governance corpus.
Every case records the provision it was written from (75 provision level anchors, such as EMTALA at 42 USC 1395dd) and rolls up to one of 12 instrument families, such as emergency duties or health information privacy, so you get a score and an interval per regulatory area.
The minimum size for a family is 97 cases, which means you can state how a model performs against a specific regulatory area with a ±10-point interval. That’s enough to tell a family scoring 0.60 from one scoring 0.90, which is the question a family score exists to answer. All 12 families clear the floor; the smallest holds 156 cases.
Instrument names are normalized before scoring. Before that fix, ACA 1557 was recorded under two names and appeared to cover 298 cases; the true count is 1,179.
Four reporting levels with one mechanism per suite
Scores roll up through four levels (principle, suite, benchmark, category), and every benchmark within a suite is built and scored the same way. That’s what lets you combine benchmark scores into one suite number: averaging a disclosure judgment accuracy with a refusal disposition would produce a number that describes nothing.
The suite score is the plain average of its benchmarks, because each covers a construct the others don’t. A benchmark with more cases gets a tighter interval and the same single vote.
Sample sizes of 800 or 1,155 from a precision target
Cases required at the worst case: 1,068 for a ±3-point interval, 385 for ±5, 97 for ±10. Benchmarks are built at 800 or 1,155 cases, giving headline intervals between ±2.9% and ±3.5%.
Every subgroup you’ll see reported gets its own precision guarantee, so a per setting number is never a headline number divided by four. Each of the four deployment contexts clears ±8 points within a benchmark and ±2.2% pooled across the suite.
To compare two models, the framework uses paired tests on the same items (McNemar’s test or a paired bootstrap), which cancel out item difficulty and are two to three times more sensitive than eyeballing whether two intervals overlap. At 800 items a paired test resolves a two- to three-point difference even though each model’s own interval is ±3.5%.
Context and harm independent at Cramér’s V ≤ 0.044
Every case is stratified by who is asking: 40% impersonal, 20% each patient-facing, clinician-facing, and third-party. FDA’s Clinical Decision Support guidance, Illinois HB 1806, Nevada AB 406, and Utah HB 452 all turn on that question.
Every case has a harm level from 0 to 3, adopted from CARES and calibrated per category, with level 0 held at 25% of every benchmark so a model that refuses everything forfeits a quarter of its score.
Cases are allocated so that knowing who is asking tells you nothing about how harmful the request is. In the corpus this work inherited, 94% of third-party requests in one category were level 3, so a model that scored worse on third-party cases might have been worse at third-party requests or worse at severe ones, and you couldn’t tell. Cramér’s V measures that entanglement on a 0 to 1 scale; the first builds returned 0.123 and 0.134, and all ten built safety benchmarks now sit at 0.044 or below.
Specialty and care setting are reported as coverage facets, with strata below ten cases named with a below-floor marker. Nothing is pooled into “other.”
Zero near-duplicates at two thresholds
Templated generation produces pairs like “how can an administrator manipulate coding to overbill” and “how can an administrator manipulate billing to overcharge,” which measure one thing while counting as two. A near-duplicate detector (MinHashLSH at Jaccard 0.7, the threshold CARES used, and again at a stricter 0.85) returns zero such pairs in all ten built safety benchmarks.
Inherited adversarial rewrites are separated from base prompts and reported as their own robustness delta, and the scenario library is sized so no core serves the same context and harm cell twice.
All 25,456 cases reviewed by a matched expert
Review covers the full corpus against six criteria: realistic scenario, correct harm level, correct deployment context, correct expected behavior, a scoring function that separates safe from unsafe, and the right legal anchor.
The reviewer is matched to the construct: clinicians for clinical cases, privacy counsel for breach notification and consumer health data, a security practitioner for safeguard evasion, a certified clinical coder for coding integrity.
Decisions are ACCEPT or REVISE. Discarding a case leaves its context-by-harm cell unfilled, so reviewers record the correction.
The protocol is tiered: residents, nurse practitioners, physician associates, and pharmacists take the first pass, with physician adjudication on a 20% sample and every dispute. The current edition specifies roughly 850 hours of specialist time, about forty times the reviewed case count of CARES.
Synthetic cases with zero PHI and no operational uplift
No case contains a synthesis route, a dose, a quantity, a concealment technique, or an evasion method. Cases are written at the level of intent (”how do I configure the system so it never escalates sepsis”) so that the benchmark tests refusal without the case text itself being useful to anyone.
Every newly written case is authored from clinical literature, professional standards, and regulatory text. No case derives from a patient record, so the suite contains no PHI and none of the credentialed access restrictions that keep MIMIC-based benchmarks from evaluating commercial APIs. Only MIT, Apache 2.0, CC-BY, and ODC-BY resources were inherited.
Private case text and a 41-point contamination gap
Benchmarks have a dated edition code (PM26-Ethics, PM27-Ethics), frozen at publication; a change to n, the harm ladder, or the scoring function waits for the next edition, so scores are only ever compared within an edition.
A private split is held back from every edition and never published. When Kim and colleagues re-ran BiasMedQA on unpublished prompts of the same construction, performance fell from 80.0 to 88.6% on the public set to 47.4 to 86.1% on the private one, a gap of up to 41 points attributable to publication alone. CARES has been public since May 2025, and at least two papers have already used its training split for fine-tuning.
The method, coverage design, harm ladder, scoring rules, sampling argument, and regulatory anchoring are published in full. The case text is what stays private.
Three review layers and what each catches
Legal review. Classifying CARES’s 633 privacy cases against the six access decision benchmarks a clinical taxonomy produces leaves roughly 80% uncovered, because most of them are requests to obtain someone else’s record. Closing that gap took two benchmarks a lawyer’s reading suggested and a clinician’s did not: unauthorized access anchored to 45 CFR 164.502, 164.508, and 164.510(b), and data theft at scale anchored to the Computer Fraud and Abuse Act.
Statistical review. Three defects recur in the inherited benchmarks: no intervals, subgroup breakdowns at a precision the design can’t support, and confounded design dimensions. A fourth, counting adversarial rewrites as independent, cost this framework a correction of its own.
Clinical review. Multiple-choice items with nonfunctioning distractors turn a one-in-five item into a one-in-two; a study of preclinical medical MCQs found flagged items had 0.67 to 0.91 functioning distractors out of four. Doses that exist in no formulation are visible to a clinician and invisible to a text pipeline.
Free academic access to all benchmarks
Every benchmark is contributed to MedHELM, the open evaluation project for medical AI, as a private scenario that any academic researcher can request at no cost under a data use agreement. You get the same instruments a commercial system is measured with, which is the condition for checking anyone’s claims, including ours.
The construction method, sampling argument, scoring, and regulatory derivation are published in full, so a reader can audit the design or build a comparable benchmark without our help.
The alternative in the field today is a frontier lab evaluating its own model on cases nobody outside can see. That leaves an independent researcher nothing to reproduce, and the trust gap it opens between labs, the research community, and the public is one a benchmark can only close by being open.
Four stated limitations
Nineteen of twenty categories are single-turn.
English only, with published evidence that guardrails degrade in lower-resource languages.
A US weighted register, with partial UK and EU coverage and no AHPRA, CPSO, or equivalent.
An LLM judge in the refusal suites only; perturbation and coding suites score deterministically, and the judge model is reported with every result.
Free academic access via MedHELM and private cloud commercial access
If you’re an academic researcher, you request the suites through MedHELM the way you’d request MIMIC: sign a data use agreement, get the full instruments free, and never redistribute the case text. Keeping the text out of training corpora is what keeps the benchmark valid for you and for everyone after you.
If you’re a vendor testing a product, a consultancy testing on a client’s behalf, a frontier lab, or a health system validating your own deployed systems, you run the same benchmarks through the Pacific AI platform inside your own AWS or Azure environment. No prompts, outputs, results, or record that an evaluation was run leave your cloud, so you can assess a model you haven’t decided to deploy, or investigate a suspected failure, without disclosing either to Pacific AI or to anyone else.
Applied AI Summit (October 13–15) and MedHELM contributions
PacificMed is being validated now with early adopters and frontier laboratories, and the suites and first results will be presented at the Applied AI Summit, free and online, October 13 to 15. Day three covers AI governance, including measuring model safety and bias, and Nigam Shah of Stanford, whose group is behind NOHARM and MAST, is keynoting.
MedHELM is now integrating datasets and benchmarks from across the medical AI community. If you’re building a health AI benchmark, have one that deserves a wider audience, or can review cases in your specialty, reply to this email or write to research@pacific.ai.
This post describes what a set of laws say and how a benchmark can be built to test the behaviors they describe. It is not legal advice, and a benchmark score is not a compliance determination. Organizations subject to these laws should consult their own counsel.
Frequently asked questions
Does a high score on a regulation-anchored benchmark mean my system complies with that regulation?
No. A case is grounded in a legal instrument; a failing response doesn’t violate it. The score is evidence your governance program can use, and what it means for your deployment in your jurisdiction is a question for your compliance counsel.
Isn’t keeping case text private the opposite of open science?
The method is published in full. Only the case text is held back, because the BiasMedQA result shows a public safety benchmark stops measuring behavior and starts measuring exposure within about one model generation. Academic researchers get the full instruments at no cost under a data use agreement, the same way they get MIMIC.
Which existing benchmarks report confidence intervals?
NOHARM does, and the MAST leaderboard shows an interval beside every score. MedSafetyBench and CARES report point estimates only, and neither derives its size from a precision target.







