AI industry

The check is the product

A scan result read null demyelination. The AI scribe wrote demyelination. One word went missing, and a clean bill of health became a marker for multiple sclerosis.

Nothing crashed. No alert fired. The note came out fluent, well-formatted and professional — and wrong. The patient caught it, because she happens to work in the health service herself and knew to question the wording. That is the part worth sitting with.

Healthwatch England went public with this last week. Their finding, in short: AI scribes used across GP surgeries and hospitals are getting drug names and diagnoses wrong, and patients are noticing errors that the professional in the room did not. One scribe swapped a prescribed drug for a different one with a similar name. A summary letter left out the consultant's instruction to get a repeat prescription — which would have quietly left someone without their migraine medication. Another recorded a doctor telling a patient to continue a drug that was never prescribed or discussed.

Twenty-seven different AI scribe systems are in use across the health service in England. There is no single national safety check they all had to clear first.

The regulator drew the line, and it isn't where you'd guess

In late July the MHRA published guidance clarifying the status of these tools. A system that transcribes a conversation, summarises it, drafts a letter, or suggests clinical codes for a professional to review is not regulated as a medical device. It only crosses that line when it starts suggesting diagnoses or treatment options, or finalises the record and places orders without a human reviewing it.

Read that carefully, because it's the most useful sentence in this whole story. The regulator drew the line at where the liability sits — on the professional who is assumed to be reading the output — not at where the risk sits. The entire safety case for twenty-seven products rests on one unstated assumption: that somebody checks.

Three things an operator should take from this

The failure mode is fluency, not breakage. A tool that crashes gets caught the same day. A tool that produces a confident, tidy, plausible wrong answer gets filed, forwarded and acted on. Every review habit your team has was built for the first kind of failure. None of it was built for the second.

The defect surfaced at the customer. In an operation, that is the most expensive place a defect can ever appear — and the only place where it also costs you trust. If the people downstream of your process are the ones finding your errors, you don't have a quality problem, you have a detection problem, and they're doing your inspection for free until they stop calling.

No external floor is coming. This is the general lesson, and it is not confined to healthcare. Regulation of these tools will keep landing behind deployment, and it will keep drawing the line at accountability rather than accuracy. Whatever standard your output is held to, you are going to have to set it yourself and hold it yourself.

What setting it yourself actually looks like

It isn't complicated, it's just unglamorous. Before you trust a tool's output, sample it against the source — a real count, on real work, written down, so you know your error rate instead of feeling it. Make the check a named step with an owner and a stop condition, not a general expectation that somebody's eyes are on it. And keep a human gate on anything that leaves the building: money, customers, compliance, permanently, by design rather than as a trial phase you eventually graduate from.

The public has already told anyone willing to listen what they want here. In Healthwatch's own polling of just over 4,000 adults, 69% said they'd feel more comfortable if the professional committed to checking the accuracy of what the scribe produced, and 81% of those with a recent appointment wanted to be told it was being used and asked first. Nobody is asking for a better model. They're asking for a verified one.

Everyone is buying the thing that writes the note. The thing that catches the missing word is the part you actually sell.

Sources: The Next Web — AI scribes are getting drug names and diagnoses wrong in NHS records (31 Aug 2026) · Covington Global Policy Watch — New UK guidance clarifies medical device status of AI scribes · Healthwatch England — What does the public think about AI scribe use in healthcare?

Set your own floor