Evidence explainer

Chronic disease in primary care

How Delirium Is Detected With the Confusion Assessment Method

The Confusion Assessment Method distills the definition of delirium into four bedside features. Its rule makes a positive result hard to fake and a negative result easy to trust too much.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. From a diagnostic manual to a bedside rule
  3. The four features and the rule that ties them together
  4. Reading the accuracy figures without over-reading them
  5. Why real-world numbers drift below the original ones
  6. Where the CAM tends to miss

A patient who was oriented yesterday is now drifting in and out of a conversation, losing the thread of a sentence, and cannot say the months of the year in reverse. The Confusion Assessment Method, or CAM, exists to turn that scene into a defensible diagnosis. It takes the formal definition of delirium, which normally lives inside a psychiatric interview, and compresses it into four features that any trained clinician can check at the bedside in a few minutes. The short version of how it works: a positive result needs an acute, fluctuating change plus inattention, and then one of two more findings. That structure makes a positive CAM worth acting on, and it also explains why a negative CAM should never be read as an all-clear.

Key points#

From a diagnostic manual to a bedside rule#

The problem the CAM was built to solve is a mismatch of setting. Delirium is common and dangerous in older hospitalized patients, yet the reference standard for confirming it was a psychiatric evaluation against the Diagnostic and Statistical Manual, a task most ward clinicians were not trained or positioned to carry out. A 1990 study by Sharon Inouye and colleagues, published in Annals of Internal Medicine, closed that gap. The team took the DSM-III-R criteria for delirium and translated them into a set of concrete items, then, before testing anything, named the four features they expected to do most of the diagnostic work.

That pre-specification is the real innovation. Instead of asking you to reconstruct a whole diagnostic definition from memory, the CAM names exactly which observations to make and exactly how to add them up. In its original validation on a small group of patients aged 65 and older at two academic sites, it reported sensitivity of 94 to 100 percent and specificity of 90 to 95 percent against the psychiatric standard. Those numbers describe the tool under study conditions, not a promise for any one patient, a distinction that becomes important later.

The four features and the rule that ties them together#

The algorithm rests on four observations:

  1. Acute onset with a fluctuating course.
  2. Inattention.
  3. Disorganized thinking.
  4. Altered level of consciousness.

The rule for combining them is deliberately strict. Delirium is registered only when features one and two are both present, and then only if feature three or feature four is also there. Think of it less as a checklist and more as a gate that will not open unless the required inputs line up.

Each feature has a practical way of being captured. Acute change and fluctuation usually come from someone who knew the patient before, a family member or the nurse who admitted them, because the whole point is a departure from that person's baseline. Inattention is tested directly, by asking the patient to recite something backward or by watching how easily the thread of a conversation slips away. Disorganized thinking surfaces as rambling, illogical, or unpredictable speech. Altered consciousness runs the whole range from drowsiness to a wired, hypervigilant state. Because the required inputs are specified in advance, two different assessors tend to reach the same answer, which is what makes the instrument reproducible.

Reading the accuracy figures without over-reading them#

The honest way to use any screening test is to keep its two error rates in separate hands. A systematic review and meta-analysis by Shi and colleagues, published in 2013 and pooling 22 studies with 2,442 patients, put the CAM's pooled sensitivity at about 82 percent and its pooled specificity at about 99 percent.

Specificity is the reassuring figure here. At roughly 99 percent, a positive CAM almost never fires on a patient who does not actually have delirium, so a positive result is a strong reason for you to treat the syndrome as real and start hunting for its cause. Sensitivity tells the cautionary half of the story. At roughly 82 percent, close to one in five patients who genuinely have delirium will still screen negative. A negative CAM lowers the odds of delirium, but it does not close the question, which is why the review framed the tool as an aid to your judgment rather than a replacement for it.

Why real-world numbers drift below the original ones#

The distance between the near-perfect 1990 figures and the more sober pooled estimates is not evidence that the instrument is broken. It is a general truth about tests: an accuracy number belongs to the exact conditions under which it was measured, including who ran it and how carefully. Sensitivity is the part that sags most once the CAM leaves the study that created it, especially when it is applied without formal cognitive testing or without training in what each feature really demands.

Work on standardizing the CAM for multicenter trials, reported in BMC Geriatrics in 2019, describes the calibration and rater training needed to keep results consistent across people and sites. The takeaway travels well beyond this one tool. Use it as a proper structured assessment and the CAM performs close to its reputation. Use it as a hurried yes-or-no impression without attention testing and it misses more cases than its headline figures suggest.

Where the CAM tends to miss#

A few blind spots follow directly from the design. Hypoactive delirium, the subdued, withdrawn presentation rather than the agitated one, is the easiest to walk past and accounts for a meaningful share of missed cases, because nothing about a calm patient prompts a second look. Baseline dementia muddies the picture too, since inattention and disorganized thinking can predate the acute illness; that is exactly why establishing your patient's usual baseline is built into the very first feature.

Remember what a positive result is and is not. The CAM identifies a syndrome; it does not name a cause. A positive screen is the opening of an evaluation for infection, medication effects, metabolic disturbance, urinary retention, and other reversible contributors, not the conclusion of one. Read that way, the tool does precisely the job it was designed for: it makes a hard diagnosis reproducible enough for non-psychiatrists to act on, while leaving the harder reasoning about why it is happening firmly in your hands.

Sources and further reading

  1. CAM original validation (Inouye 1990, Annals of Internal Medicine)
  2. CAM diagnostic accuracy meta-analysis (Shi 2013)
  3. CAM training and standardisation (BMC Geriatrics 2019)

Questions and answers

Does a negative CAM rule out delirium?

No. Pooled sensitivity is about 82 percent, so roughly one in five patients with delirium can still screen negative. A negative result lowers the probability but does not exclude the diagnosis, and a high-suspicion case should be reassessed with structured attention testing.

Why is a positive CAM more trustworthy than a negative one?

Because specificity is very high, around 99 percent in pooled data. A test with high specificity rarely produces false positives, so when the CAM is positive it strongly supports that delirium is present and should prompt a search for the cause.

Why does the CAM depend so much on training?

The features require actual observation, especially direct testing of attention. When the tool is used as a quick impression rather than a structured assessment, sensitivity falls, which is why standardization and rater training were developed for its use in trials.