HarnessHealth

The argument · ten sources

Everything in clinical AI got free.
One thing didn’t.

In eighteen months the model got free and won. The specialized medical tool got beaten by the general one. The safety harness got patented — by someone else, who published the benchmarks. This month a 2.8-trillion-parameter open-weight model drew level with Opus 4.8 at a third of the cost, and the refusals did not come with it. Every layer of the stack raced to zero.

One layer did not, and it is not a technical layer. When a country deployed an open model across more than three hundred hospitals, the liability did not land on the model. It landed on the doctors. It always lands on whoever signed.

What got free.

Not a forecast. Five measurements, taken between early 2025 and this month, each from a source that has no stake in this argument.

The model

It got free — and then it won.

DeepSeek R1, an open-weight model, topped Stanford’s MedHELM: a 66% win rate and a 0.75 macro-average across 35 benchmarks and 121 clinician-validated tasks. It runs inside a hospital firewall at zero marginal cost. Its strongest categories were clinical note generation (0.85) and patient communication (0.89) — the delegable work.

MedHELM, Nature Medicine

The medical tool

It got beaten by the general model.

General-purpose frontier models outperformed the specialized clinical tools built for exactly this job — including UpToDate, the standard reference in American medicine. The assumption that a tool built for medicine must beat a tool built for everything did not survive contact with the evidence.

Nature Medicine, reported by Eric Topol

The harness

It got patented — by someone else.

A physician-built neurosymbolic stack takes a small model from 30.2% to 95.8% safety-intervention on psychosis-bench. Allowed US patent. Benchmarks published openly, naming the bare model. A harness is buildable, patentable, and now public. It is not a moat.

asha-bench-public

The evidence

It never showed up.

Organizational AI adoption reached 88%. The FDA authorized 258 AI-enabled medical devices in 2025 — only 2.4% supported by randomized-trial data. Of 500+ clinical-AI studies reviewed, roughly 5% used real clinical data. The regulator is clearing AI faster than the evidence is being generated.

Stanford AI Index 2026

The guardrail

It didn’t come with the weights.

Kimi K3 shipped on 16 July 2026: 2.8 trillion parameters, open weights, third on the Artificial Analysis Intelligence Index — level with Opus 4.8, at roughly a third of the cost per task. Its predecessor was independently evaluated as carrying “similar dual-use capabilities” to GPT-5.2 and Claude Opus 4.5, but with “significantly fewer refusals.” Refusal is a vendor policy, not a property of the capability. It is the one layer that does not travel with a downloaded file — and the one nobody can patch afterward.

Independent safety evaluation of Kimi K2.5

The precedent · another industry ran the whole movie

Solar ran this cascade first. The ending is on record.

Clinical AI is not the first stack to race to zero. Solar hardware commoditized two decades earlier: module prices collapsed, one country came to dominate the commodity layer, and every scarcity prediction failed. What refused to fall was not technical — it was the permission layer around the install. That industry has now run long enough to show which part of the permission layer dies, and which part survives.

The commodity won

Clean power met all new demand.

In 2025, for the first time, growth in clean generation covered the entire growth in global electricity demand, and renewables overtook coal for the first time in a century — 33.8% of generation to coal’s 33.0%. Solar alone met 75% of demand growth. The hardware layer did not just get cheap. It won.

Ember, Global Electricity Review 2026

The cost that wouldn’t fall

The panel got free. The permission didn’t.

The panels cost roughly the same in every country. Yet an American rooftop system costs two to three times an Australian one, and the difference is almost entirely soft costs — permitting, inspection, interconnection, labor, customer acquisition: everything attached to people rather than silicon. Historically about 65% of a US system’s price, versus roughly 25% in Australia.

US Department of Energy, solar soft costs

What survived automation

The paperwork died. The signature didn’t.

NREL’s SolarAPP+ now issues residential permits near-instantly in 450+ jurisdictions — the bureaucratic part of the permission layer was always automatable, and it is being automated. What the fastest solar regime on earth kept: in Australia, the federal rebate exists only if a licensed, insured, accredited installer signs the installation. Every removable cost was removed. The accountable human wasn’t.

Clean Energy Regulator (Australia)

That is the distinction this page turns on. Bureaucracy is not accountability. The cascade destroys the first — instant permits, automated paperwork — and comes to rest on the second: a named, licensed person signing the work. Solar reached that ending first. Clinical AI is running the same cascade, with the same two kinds of permission in its path.

The experiment · a country ran it

Nobody had to argue about where the liability goes. It was measured.

Beginning in early 2025, an open model was deployed inside the firewalls of more than 300 Chinese hospitals — within weeks, the fastest clinical-AI rollout ever attempted. The data never left the building. The model was free. The infrastructure was already built. It was, in every technical sense, the thing everyone says they want.

Then Nature Medicine reported what happened. The model produced “plausible but factually incorrect outputs.” Clinicians risked over-reliance. The regulatory framework was a gray area. And on the question that decides everything — who pays when it is wrong — the finding was unambiguous:

“Hospitals and doctors are more likely to bear diagnosis and treatment damage liability rather than medical product liability if misdiagnosis occurs.”— Nature Medicine, on DeepSeek’s deployment across Chinese hospitals

Free model. Local data. National scale. And the liability came to rest exactly where it always rests: on the licensed human who signed.

"AI is a great second opinion. Ultimately, clinical judgement prevails."— Dr. Sharon Heng, Consultant Ophthalmic Surgeon, Moorfields Eye Hospital
"Most health systems are monitoring AI the same way they governed a new MRI scanner in 2010: a subcommittee, a checklist, a quarterly meeting. For AI, it is dangerously inadequate."— Dr. Peter Pronovost, Chief Quality Officer, University Hospitals; MacArthur Fellow; Time 100 Most Influential
"Trust is currency; if finance leaders cannot audit the model, they will not rely on it."— Brian Overstreet, President & CEO, Arbital Health, on actuarially engineered AI for value-based care

The harness · necessary, and not sufficient

A harness can make a model safer. It cannot make it accountable.

The bare model is genuinely dangerous. On psychosis-bench, language models showed a strong tendency to perpetuate a user’s delusion at every turn rather than challenge it — offering a safety intervention in only about a third of the turns that warranted one.

What the harness fixes

30.2% → 95.8%

The same small model, wrapped in a physician-built neurosymbolic stack, on the same test. The safety came from the scaffolding, not the weights. The harness is real, and it works.

What the harness cannot fix

“Never gives definitive diagnoses.”

That is the same system’s own description of itself — built by physicians, patented, benchmarked in public. It reached 95.8% and still tells its users it does not diagnose. Not a wall it hit. A boundary it chose, because the boundary is not technical.

The same boundary, drawn at industrial scale

This is not a small-team artifact. Hippocratic AI — roughly $404M raised, a $3.5B valuation — built the most sophisticated safety harness in healthcare AI: a multi-model “Polaris” constellation of supervisory models cross-checking a conversational one, pressure-tested by a panel of 7,500 licensed nurses paid to attack every release until it scored safer than a human before launch. And its scope is drawn, deliberately, at the scope of a registered nurse or less — it does not diagnose, does not prescribe, and hands off to a human for the judgment call. The best-funded harness in the industry stops at the same line the physician-built one does. When a company worth $3.5 billion spends its safety budget proving where it must escalate to a human, that line is not a limitation of the technology. It is the product.

The same boundary, drawn by the regulator

In 2026 the FDA's device center issued a discussion paper on generative-AI medical devices that grades them on two axes — how autonomously the AI acts, and how much harm a wrong output can do — and places a human checkpoint before any consequential action, with monitoring that continues after clearance. It also reminds sponsors that a foundation-model master file is not model approval: the company deploying the AI stays responsible for it. A discussion paper, not a rule — but the regulator is drawing the boundary at exactly the place the builders do: the accountable human.

Why this one layer holds

You cannot compile a license.

A medical license does not attach to a company. It does not attach to a founder’s credentials. It does not attach to a model, a corpus, or a patent. It attaches to a person, signing a specific patient’s care, who can be held responsible for it.

That is why two physicians can build the best safety harness on the market and still write “not a licensed medical provider” on their own homepage. Their credentials did not transfer to their software. Nobody’s do. It is also why every other player — the phone maker, the search company, the EHR vendor, the open-source de-identifier, twenty-five state legislatures — arrives at the same line from a different direction and stops there.

The line is not a harder model. It is not a better harness. It is a signature, and a commons cannot hold one.

See every player stopping at the same line →

The mechanism

So we made signing the design, not the policy.

Everyone agrees a human should be in the loop. A policy saying so is a sentence in a handbook. We made it a structural property of the system: there is no code path that emits an unsigned clinically-consequential output.

01

Drafted

Any model — Claude, GPT, Gemini, open-weight, on-device — drafts the output in seconds. The model is interchangeable by design, because the model is the commodity.

02

Held

The draft is intercepted, not advisory-flagged. No path sends it to a patient, a payer, or a record unsigned. This is the part a policy cannot do and an architecture can.

03

Attested

A named physician reviews and signs — NPI-bound, timestamped, hash-anchored to an externally verifiable chain. The output becomes actionable because someone is accountable for it.

Honest scope

What this argument does not claim.

  • Not that we have the best model. We do not train one. The argument above is that the model is the commodity — which is only credible if we act like it. Any model, swapped at will.
  • Not that we have the best harness. Others are further along on harness architecture and IP: the neurosymbolic system cited above holds an allowed US patent and publishes its benchmarks. Ours is a drafted provisional. If harness quality were the moat, we would not be claiming it.
  • Not that AI is unsafe, or that it should be slowed. The measured gains are real — up to 83% less physician time on notes, with meaningful burnout reduction. The point is not to hold the technology back. It is to make its outputs actionable by making someone accountable for them.
  • Not that open weights are the problem. The same independent evaluation we cite found the open model has no frontier-level autonomous cyber capability — the alarming version of that story is not supported, and we are not making it. Open weights are why a hospital can run this at zero marginal cost inside its own firewall, which is good. The narrow point is that refusal behavior is a vendor policy that does not travel with a downloaded file. That is an argument about where accountability has to sit, not an argument against the file.
  • Not that friction is the moat. Most of any industry’s permission layer is removable bureaucracy, and it deserves to be removed — solar’s instant-permit software is a good thing, and so is automating prior-auth paperwork. If this position depended on paperwork surviving, it would be indefensible. The claim is narrower: when every removable cost is removed, what remains is the accountable signature. We build for that layer, not for the friction around it.
  • Not that we have outcome data of our own. The evidence on this page is other people’s, deliberately — ten independent sources with no stake in our conclusion. Our own operating data is what comes next, and it will be published the same way.

Every layer of the stack is being built and given away — the model, the sensing, the de-identification, and now the safety harness itself, patented and benchmarked in the open. Except one. A commons can build a harness. It cannot hold a medical license or assume the liability. What cannot be copied is not the harness — it is the signature inside it.