The argument · fourteen sources
Everything in clinical AI got free.
One thing didn’t.
In eighteen months the model got free and won. The specialized medical tool got beaten by the general one. The safety harness got patented — by someone else, who published the benchmarks. In July a 2.8-trillion-parameter open-weight model drew level with Opus 4.8 at about half the cost per task, and the refusals did not come with the weights. Every layer of the stack raced to zero.
One layer did not, and it is not a technical layer. When a country deployed an open model across more than three hundred hospitals, the liability did not land on the model. It landed on the doctors. It always lands on whoever signed.
What got free.
Not a forecast. Five measurements, taken between early 2025 and this month, each from a source that has no stake in this argument.
The model
It got free — and then it won.
DeepSeek R1, an open-weight model, topped Stanford’s MedHELM: a 66% win rate and a 0.75 macro-average across 35 benchmarks and 121 clinician-validated tasks. It runs inside a hospital firewall at zero marginal cost. Its strongest categories were clinical note generation (0.85) and patient communication (0.89) — the delegable work.
The medical tool
It got beaten by the general model.
Three general-purpose frontier models outperformed OpenEvidence and UpToDate Expert AI, clinical tools built for exactly this job, in all three evaluations: exam questions, HealthBench, and 100 real physician queries rated by blinded clinicians. On those real queries the clinical tools scored no better than Google’s search AI Overview. The assumption that a tool built for medicine must beat a tool built for everything did not survive the test.
The harness
It got patented — by someone else.
A physician-built neurosymbolic stack posts a 95.8% safety-intervention rate on psychosis-bench, against 30.2% for the bare Gemini 2.5 Flash model it mostly routes to. The benchmark was pre-registered, and its own report grades the formal verdict inconclusive because its two AI judges agreed too weakly (κ 0.249). Allowed US patent. Benchmarks published openly, naming the bare model. A harness is buildable, patentable, and now public. It is not a moat.
The evidence
It never showed up.
Organizational AI adoption reached 88%. The FDA authorized 258 AI-enabled medical devices in the first nine months of 2025, more than in any prior full year — yet of the AI devices it had cleared through 2024 that reported clinical studies, only 2.4% were supported by randomized-trial data. Of 500+ clinical-AI studies reviewed, roughly 5% used real clinical data. The regulator is clearing AI faster than the evidence is being generated.
The guardrail
It didn’t come with the weights.
Kimi K3 shipped on 16 July 2026: 2.8 trillion parameters, open weights, third on the Artificial Analysis Intelligence Index — level with Opus 4.8, at about half its cost per task. Its predecessor was independently evaluated as carrying “similar dual-use capabilities” to GPT-5.2 and Claude Opus 4.5, but with “significantly fewer refusals” on CBRNE-related requests. Refusal is a vendor policy, not a property of the capability. It is the one layer that does not travel with a downloaded file — and the one nobody can patch afterward.
The precedent · another industry ran the whole movie
Solar ran this cascade first. The ending is on record.
Clinical AI is not the first stack to race to zero. Solar hardware commoditized two decades earlier: module prices collapsed, one country came to dominate the commodity layer, and every scarcity prediction failed. What refused to fall was not technical — it was the permission layer around the install. That industry has now run long enough to show which part of the permission layer dies, and which part survives.
The commodity won
Clean power met all new demand.
In 2025, for the first time, growth in clean generation covered the entire growth in global electricity demand, and renewables overtook coal for the first time in a century — 33.8% of generation to coal’s 33.0%. Solar alone met 75% of demand growth. The hardware layer did not just get cheap. It won.
The cost that wouldn’t fall
The panel got free. The permission didn’t.
The panels cost roughly the same in every country. Yet an American rooftop system costs two to three times an Australian one, and the difference is almost entirely soft costs — permitting, inspection, interconnection, labor, customer acquisition: everything attached to people rather than silicon. Historically about 65% of a US system’s price, versus roughly 25% in Australia.
What survived automation
The paperwork died. The signature didn’t.
NREL’s SolarAPP+ now issues residential permits near-instantly in 450+ jurisdictions — the bureaucratic part of the permission layer was always automatable, and it is being automated. What the fastest solar regime on earth kept: in Australia, the federal rebate exists only if a licensed, insured, accredited installer signs the installation. Every removable cost was removed. The accountable human wasn’t.
That is the distinction this page turns on. Bureaucracy is not accountability. The cascade destroys the first — instant permits, automated paperwork — and comes to rest on the second: a named, licensed person signing the work. Solar reached that ending first. Clinical AI is running the same cascade, with the same two kinds of permission in its path.
The experiment · a country ran it
Nobody had to argue about where the liability goes. It was measured.
Beginning in early 2025, an open model was deployed inside the firewalls of more than 300 Chinese hospitals — within weeks, the fastest clinical-AI rollout ever attempted. The data never left the building. The model was free. The infrastructure was already built. It was, in every technical sense, the thing everyone says they want.
Then Nature Medicine reported what happened. The model produced “plausible but factually incorrect outputs.” Clinicians risked over-reliance. The regulatory framework was a gray area. And on the question that decides everything — who pays when it is wrong — the finding was unambiguous:
“Hospitals and doctors are more likely to bear diagnosis and treatment damage liability rather than medical product liability if misdiagnosis occurs.”— Nature Medicine, on DeepSeek’s deployment across Chinese hospitals
Free model. Local data. National scale. And the liability came to rest exactly where it always rests: on the licensed human who signed.
"AI is a great second opinion. Ultimately, clinical judgement prevails."— Dr. Sharon Heng, Consultant Ophthalmic Surgeon, Moorfields Eye Hospital
"Most health systems are monitoring AI the same way they governed a new MRI scanner in 2010: a subcommittee, a checklist, a quarterly meeting. For AI, it is dangerously inadequate."— Dr. Peter Pronovost, Chief Quality Officer, University Hospitals; MacArthur Fellow; Time 100 Most Influential
"Trust is currency; if finance leaders cannot audit the model, they will not rely on it."— Brian Overstreet, President & CEO, Arbital Health, on actuarially engineered AI for value-based care
The harness · necessary, and not sufficient
A harness can make a model safer. It cannot make it accountable.
The bare model is genuinely dangerous. On psychosis-bench, all eight language models tested showed a strong tendency to perpetuate a user’s delusion rather than challenge it, offering a safety intervention in only about a third of the turns that warranted one.
What the harness fixes
30.2% → 95.8%
Bare Gemini 2.5 Flash against a physician-built neurosymbolic stack that routes most queries to that same model, on the same 16 scenarios. The authors attribute the gap to the scaffolding, and both AI judges agree on its direction. Their pre-registered judge-agreement check still failed (κ 0.249), so they grade the formal verdict inconclusive. The gap looks real. The grading is the weak link.
What the harness cannot fix
“Never gives definitive diagnoses.”
That is the same system’s own description of itself — built by physicians, patented, benchmarked in public. It reached 95.8% and still tells its users it does not diagnose. Not a wall it hit. A boundary it chose, because the boundary is not technical.
The same boundary, drawn at industrial scale
This is not a small-team artifact. Hippocratic AI — roughly $404M raised, a $3.5B valuation — built the most sophisticated safety harness in healthcare AI: a multi-model “Polaris” constellation of supervisory models cross-checking a conversational one, safety-tested by thousands of licensed nurses and hundreds of physicians, under a stated rule that its agents must be safer than the humans doing the same job. And its scope is drawn, deliberately, at non-diagnostic, patient-facing work — the agents do not diagnose and do not prescribe, and the judgment call stays with a clinician. The best-funded harness in the industry stops at the same line the physician-built one does. When a company worth $3.5 billion spends its safety budget proving where it must escalate to a human, that line is not a limitation of the technology. It is the product.
The same boundary, drawn by the regulator
In 2026 the FDA's device center issued a discussion paper on generative-AI medical devices that grades them on two axes — how autonomously the AI acts, and how much harm a wrong output can do — and places a human checkpoint before any consequential action, with monitoring that continues after clearance. It also reminds sponsors that a foundation-model master file is not model approval: the company deploying the AI stays responsible for it. A discussion paper, not a rule — but the regulator is drawing the boundary at exactly the place the builders do: the accountable human.
The same boundary, drawn by the record itself
In September 2026 The Lancet published Mahmood and Topol on the new temporal foundation models that turn a whole patient journey into a computable representation, and opened with the caveat: the medical record is not the patient; it is the partial trace of a patient's encounters with a health system. A month earlier JAMA Network Open showed that across 3.6 million portal message threads, writing style — not clinical content — explained about half of the gap in how often care teams answered patients from historically marginalized groups. Both point at the same gap: a representation nobody can interrogate, and a triage nobody wrote down. Access logs say who looked. A signature with a written rationale says who is accountable, and why. That is the one part of the record an auditor can still check years later, and it is the part this layer exists to produce.
Why this one layer holds
You cannot compile a license.
A medical license does not attach to a company. It does not attach to a founder’s credentials. It does not attach to a model, a corpus, or a patent. It attaches to a person, signing a specific patient’s care, who can be held responsible for it.
That is why two physicians can build the best safety harness on the market and still write “not a licensed medical provider” on their own homepage. Their credentials did not transfer to their software. Nobody’s do. It is also why every other player — the phone maker, the search company, the EHR vendor, the open-source de-identifier, twenty-five state legislatures — arrives at the same line from a different direction and stops there.
The clearest version of that is written by the search company itself. Google DeepMind ships MedGemma — open-weight medical models, 4B and 27B, that read CT, MRI, histopathology and chest films, downloadable and runnable on your own hardware — and states in the same documentation that the outputs are “not intended to directly inform clinical diagnosis, patient management decisions, treatment recommendations, or any other direct clinical practice applications” and that all outputs “require independent verification, clinical correlation, and further investigation.” The best-resourced medical AI group on earth gave the model away and kept the disclaimer — not because the model is weak, but because the last step was never a modeling problem.
Google DeepMind, MedGemma model documentation
The line is not a harder model. It is not a better harness. It is a signature, and a commons cannot hold one.
See every player stopping at the same line →The load-bearing test · what breaks without it
Take the signature out and three systems lose their anchor at once.
The usual argument about clinical AI is a scoreboard argument: can the model out-diagnose the doctor. Assume it can. The interesting question is what the signature was holding up, and the answer is that it was never one thing. Liability, payment and regulatory category are three independent systems that happen to be bolted to the same person. None of them is about whether the care is good.
Liability
Malpractice has nothing to attach to.
The strongest case for autonomous care was made in JAMA this August by Emanuel and the Khoslas, who make the case that AI alone could outperform both physicians and physician-AI hybrids on cognitive medical tasks, and expect it ready for real-world use in some, maybe many, clinical workflows by 2030. Grant all of it. Malpractice still attaches to a licensed person who signed. Remove that person and the mechanism by which a harmed patient is made whole has no anchor \u2014 not because the care got worse, but because there is no longer anyone the claim can name.
Reimbursement
The payment codes are built around a person.
This is not a forecast. In the CY2027 Physician Fee Schedule proposed rule, CMS proposes that remote monitoring be furnished by direct employees of the billing practitioner \u2014 tightening, not loosening, the attachment between payment and an identifiable human. Care that no clinician furnished does not have a weak claim to a code. It has no code.
Regulation
The reviewer is not a safety check. It is the exemption.
This one is usually stated backwards. It is said that FDA clears decision-support tools partly because a clinician reviews the output. Under section 520(o)(1)(E), added by the 21st Century Cures Act, qualifying software is excluded from the definition of a device altogether \u2014 and the fourth criterion is that the clinician can independently review the basis for a recommendation rather than rely primarily on it. Remove the reviewing clinician and the product does not lose a safeguard. It changes regulatory category.
Read together they make a sharper point than any one of them. The same human is load-bearing in two opposite directions at once. To stay outside FDA’s device definition, a decision-support product needs a clinician who independently reviews the basis. To be paid under the remote monitoring device codes, the software has to be a device. One person, two regulatory regimes, pulling opposite ways — and every serious clinical AI product sits somewhere on that line.
We do not claim a physician review makes a model more accurate. On tasks where the model is already better, it may not. The claim is narrower and harder to dislodge: the review makes the output owned. Accuracy is a property of a system. Accountability is a property of a person.
There is a fourth system in this argument that we are not going to pretend to have solved. If AI carries more of the cognitive load, the skills clinicians use less will erode, and medical training is still preparing physicians for a job that is moving underneath it. Aviation designed around exactly this and it took decades. We build the accountability layer. Nobody, including us, has built the answer to deskilling.
The mechanism
So we made signing the design, not the policy.
Everyone agrees a human should be in the loop. A policy saying so is a sentence in a handbook. We made it a structural property of the system: there is no code path that emits an unsigned clinically-consequential output.
01
Drafted
Any model — Claude, GPT, Gemini, open-weight, on-device — drafts the output in seconds. The model is interchangeable by design, because the model is the commodity.
02
Held
The draft is intercepted, not advisory-flagged. No path sends it to a patient, a payer, or a record unsigned. This is the part a policy cannot do and an architecture can.
03
Attested
A named physician reviews and signs — NPI-bound, timestamped, hash-anchored to an externally verifiable chain. The output becomes actionable because someone is accountable for it.
Honest scope
What this argument does not claim.
- Not that we have the best model. We do not train one. The argument above is that the model is the commodity — which is only credible if we act like it. Any model, swapped at will.
- Not that we have the best harness. Others are further along on harness architecture and IP: the neurosymbolic system cited above holds an allowed US patent and publishes its benchmarks; ours publishes neither yet. If harness quality were the moat, we would not be claiming it.
- Not that AI is unsafe, or that it should be slowed. The measured gains are real — up to 83% less physician time on notes, with meaningful burnout reduction. The point is not to hold the technology back. It is to make its outputs actionable by making someone accountable for them.
- Not that open weights are the problem. The same independent evaluation we cite found the open model has no frontier-level autonomous cyber capability — the alarming version of that story is not supported, and we are not making it. Open weights are why a hospital can run this at zero marginal cost inside its own firewall, which is good. The narrow point is that refusal behavior is a vendor policy that does not travel with a downloaded file. That is an argument about where accountability has to sit, not an argument against the file.
- Not that friction is the moat. Most of any industry’s permission layer is removable bureaucracy, and it deserves to be removed — solar’s instant-permit software is a good thing, and so is automating prior-auth paperwork. If this position depended on paperwork surviving, it would be indefensible. The claim is narrower: when every removable cost is removed, what remains is the accountable signature. We build for that layer, not for the friction around it.
- Not that we have outcome data of our own. The evidence on this page is other people’s, deliberately — ten independent sources with no stake in our conclusion. Our own operating data is what comes next, and it will be published the same way.
Every layer of the stack is being built and given away — the model, the sensing, the de-identification, and now the safety harness itself, patented and benchmarked in the open. Except one. A commons can build a harness. It cannot hold a medical license or assume the liability. What cannot be copied is not the harness — it is the signature inside it.