A tool that drafts your clinic notes should be judged on the errors it hides, not the minutes it saves. Before you trust an accuracy claim, ask three questions the sales deck rarely answers: does the draft leave out facts that were said, does it assert things that were never said, and does it shift how certain your language sounds? Time saved is simple to count and simple to oversell. Note accuracy is harder to measure, and it is the number that decides whether the document is safe to sign.
Ambient documentation tools listen to a visit and return a draft note. The promise is easy to like: less typing, more eye contact, shorter evenings. Because the promise is about workflow, the evidence people quote tends to be about workflow too, and that is the first place to slow down. A product can genuinely give back time and still hand you a record that is harder to read or subtly wrong. The useful question is not "did it feel faster" but "what does it get wrong, how often, and how much does each mistake cost."
Key points#
- Judge an ambient scribe on accuracy, not time saved; the two do not track together.
- The three failure modes that matter are omission, fabrication, and certainty drift.
- Omission is the most dangerous error because nothing on the page looks wrong.
- A raw error rate is meaningless without severity grading; ask how many errors could change management.
- Prefer randomized comparisons and blinded note review over before-and-after satisfaction surveys.
The three ways a draft note goes wrong#
To evaluate accuracy you first have to name the failure modes, because each one hides in a different place.
Omission is a fact that was spoken and clinically relevant, then dropped. It is the hardest error to catch precisely because the page looks complete; the reader sees a clean note and has no way to know a detail is missing. Think of it as a false negative on a lab test, reassuring right up until it is wrong.
Fabrication, sometimes called hallucination or confabulation, is the mirror image: the note states something that never happened. A symptom the patient denied, an exam that was not performed, a value nobody said out loud. Because fabricated text reads as confident and specific, it is easy to accept at a glance.
Certainty drift is the subtle one and the most human to miss. A patient says a symptom is "a little" better and the note records "improved." A finding described as "possible" hardens into "present." No fact is added or removed, but the confidence of the language moves, and with it the clinical meaning.
A 2025 framework published in npj Digital Medicine put numbers on the first two. Across nearly 13,000 clinician-annotated sentences from AI-drafted clinical text, the authors measured a fabrication rate of roughly 1.5 percent and an omission rate of roughly 3.5 percent. Omissions ran about twice as common as fabrications, which fits the intuition that dropping detail is the natural failure of any system that summarizes. The more important figure was the severity grading: around 44 percent of the fabrications were judged clinically major, meaning they could plausibly change a diagnosis or a treatment decision. A low error rate reassures no one on its own. You have to know how many of the errors were the kind that could hurt a patient.
Why time saved is the metric to distrust first#
Documentation time is the number vendors reach for, and its appeal is obvious. It is objective, the electronic record captures it automatically, and it moves in the direction everyone wants.
A pragmatic randomized trial run at UCLA and published in NEJM AI in late 2025 shows how slippery even this friendly metric can be. Investigators randomized 238 outpatient physicians across 14 specialties to one of two commercial ambient scribes or to usual care, and tracked time spent in the note. One tool cut time-in-note by about 9.5 percent against control; the other produced no statistically significant reduction at all. Same class of product, same trial, opposite headline. When one number can split that far between two vendors, one number is not a verdict.
The study design matters as much as the result. Because it was randomized, each tool was compared against a concurrent control group rather than against how a clinician remembered feeling last quarter. Most claims you will encounter are not built this way. They are before-and-after satisfaction surveys, which flatter any new tool, since novelty and self-selection both push the rating in the same direction.
That same UCLA trial also catalogued the texture of the accuracy failures physicians reported: dropped facts, pronoun-resolution errors, missed negation and affirmation, speaker misattribution, and structural clutter. Negation failure deserves a second look. If a patient denies chest pain and the note records chest pain, that is not a typo. It inverts the clinical picture.
Note bloat is a failure mode too#
Length is its own kind of error. Ambient tools tend to be generous, and a generous note is not a better note. When every utterance becomes a documented finding, the signal a reader actually needs gets buried under padding. A scoping review of ambient documentation metrics, posted to medRxiv in early 2025, found that most evaluation frameworks measure surface similarity to a reference text using natural-language metrics such as ROUGE and BERTScore, while few capture whether the note is clinically useful or appropriately concise. A draft can score well on textual overlap and still be padded, over-hedged, and tiring to read. What deserves measuring is whether a clinician can find the important facts fast, not whether the text resembles a template.
A short checklist for reading any claim#
Bring the same questions to every study, demo, or vendor pitch.
- What was measured? If the only outcome is time or satisfaction, accuracy was never tested.
- What was the comparison? Randomized against a concurrent control beats before-and-after, which cannot separate the tool from novelty.
- Were omissions checked specifically? They are the errors that leave no mark on the page, so they need to be sought deliberately.
- Were errors graded for severity? A rate without a harm scale tells you volume, not risk.
- Who read the notes? Clinicians blinded to the tool are a far better judge than the people who built or bought it.
Notice, too, the point the scoping review makes unavoidable: the field still lacks a shared standard for the things that matter most, which means many published numbers are not comparable to one another.
Sources and further reading
Questions and answers
Does an AI scribe that saves time also improve accuracy?
Not necessarily. The two are separate outcomes, and the UCLA trial showed one tool saving meaningful time while another saved none. Time and accuracy have to be measured, and reported, on their own terms.
Which error is most dangerous?
Omission, because it is invisible. A fabricated detail can be caught on a careful read; a missing fact leaves the note looking complete and correct. That is why evaluations have to test for omissions on purpose.
Is the case against ambient documentation?
No. These tools can genuinely return time and attention to the visit, and the npj framework found that careful prompt and workflow design can push major-error rates below what is reported for unaided human notes. The argument is narrower: a tool that writes a medical record is a clinical instrument, and it earns the same scrutiny as any other, not "did it feel faster" but "what does it get wrong, how often, and how much does that cost.