Evidence explainer

Chronic disease in primary care

How Syncope Risk Scores Are Built and Validated

A syncope risk score is only as good as the studies behind it, and the real test is prospective validation in new patients, not a good fit to its own data.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. What the score is actually trying to solve
  3. Building the score: derivation
  4. Earning trust: validation in new patients
  5. The two numbers that matter more than a cutoff
  6. A short checklist for reading any risk score

A trustworthy syncope risk score is not the one with the cleverest math. It is the one that was tested on patients it had never seen and still sorted them correctly. That distinction, between a formula that fits its own data and a rule that holds up in the real world, is the single most useful thing to understand when you read any clinical decision tool. The Canadian Syncope Risk Score (CSRS), published by Thiruganasambandamoorthy and colleagues in CMAJ in 2016, is a good teaching case because its full life story has been published in the open: how it was built, how it was tested elsewhere, and where it falls short.

Key points#

What the score is actually trying to solve#

Syncope sends a lot of people to the emergency department, and most of them are fine. A small minority are on the edge of a dangerous arrhythmia, a heart attack, a pulmonary embolism, or another serious event. The clinical problem is a sorting problem: send the low-risk majority home safely, and hold the few who need watching. A risk score is an attempt to turn that judgment into a repeatable, auditable estimate of probability. It does not make the decision. It gives the clinician a calibrated number to reason from.

Building the score: derivation#

The CSRS was built from 4,030 adults who came to six Canadian teaching-hospital emergency departments across four cities with syncope. Within 30 days, 147 of them (3.6 percent) had a serious adverse event, a bundled outcome covering death, arrhythmia, myocardial infarction, structural heart disease, aortic dissection, pulmonary embolism, and serious hemorrhage. The investigators began with a wide list of candidate predictors drawn from the history, the physical, the electrocardiogram, and the labs, then ran multivariable logistic regression to see which ones carried weight on their own rather than merely riding along with others.

Nine predictors survived that filter. Each was turned into integer points scaled to its regression coefficient. A vasovagal predisposition scores minus 1; a history of heart disease plus 1; a systolic blood pressure below 90 or above 180 mm Hg plus 2; a troponin above the 99th percentile plus 2; an abnormal QRS axis plus 1; a QRS duration over 130 ms plus 1; a corrected QT over 480 ms plus 2; and the emergency physician's own impression, minus 2 for a vasovagal picture and plus 2 for a cardiac one. Totals range from minus 3 to plus 11, translating to a predicted risk from roughly 0.4 percent up to about 84 percent. Those totals were then binned into risk bands running from very low, through low and medium, up to high and very high.

Two details in the derivation separate a durable rule from one that will disappoint. The first is discrimination, measured here by a C-statistic of 0.88 (0.87 after correction for optimism), which means the score reliably ranks a patient who will deteriorate above one who will not. The second is the team's honesty about overfitting. They used bootstrap resampling to estimate a shrinkage factor of 0.91 and applied it to the coefficients, deliberately pulling the predictions toward the average so the rule would travel better to patients outside the original cohort. They also reported calibration, the match between predicted and observed risk, rather than reporting discrimination alone. That matters because a model can rank patients well and still be systematically off in the numbers it produces.

Earning trust: validation in new patients#

Derivation only shows that a rule fits the data it was carved from. Anyone can build a formula that describes its own patients. The question that decides whether a score deserves use is whether it survives contact with new people, and an internal bootstrap check is not a substitute for that.

The CSRS was tested prospectively in a separate Canadian cohort of 3,819 syncope patients across nine emergency departments, reported in JAMA Internal Medicine in 2020. Discrimination held up, with an area under the curve of 0.91, and at a threshold of minus 1 the score reached a sensitivity near 97.8 percent for 30-day serious outcomes, with fewer than 1 percent of patients in the low-risk bands going on to a serious event. A careful validation like this one reports calibration in the new cohort instead of assuming it carries over, because a sharp-looking rule that over- or under-predicts can mislead clinicians in either direction.

Then comes the harder test: does the score work in health systems built differently, with different case mix, referral patterns, and troponin assays? An international validation in Annals of Internal Medicine in 2022 enrolled patients aged 40 and older across emergency departments in eight countries on three continents. The CSRS discriminated well again (area under the curve 0.85, against 0.74 for the older OESIL score), and fewer serious events were missed among patients it triaged as low risk. That study also surfaced an honest limitation worth repeating: a stripped-down version using only the clinician's diagnostic impression performed almost as well as the full nine-item score. In other words, the tool sharpens clinical judgment rather than replacing it, and naming that dependency is exactly what good external validation is for.

The two numbers that matter more than a cutoff#

When a missed case can mean death or a lethal arrhythmia, sensitivity is the number to watch. A score that catches nearly everyone who will deteriorate is doing its job even when it also flags some who will not, because a short period of observation is a far smaller harm than discharging a patient who then arrests. This is why validation reports foreground sensitivity and negative predictive value at low thresholds rather than a single balanced cutoff.

But sensitivity has an obvious loophole: you can score a perfect 100 percent simply by calling every patient high risk. The correction for that is net benefit, formalized in decision-curve analysis. It weighs true positives against the harm of false positives across the whole range of thresholds a reasonable clinician might pick, and it asks whether the score beats two trivial defaults, admit everyone and admit no one, over the thresholds that reflect real stakes. Seen this way, a decision rule is a probability estimator feeding a value judgment, not a verdict. The question is never only "what is the score," but "does acting on this score, at a defensible threshold, do more good than the default." No score replaces individualized clinical assessment.

A short checklist for reading any risk score#

When you meet a new score, look past the headline accuracy figure and ask five things. Was it derived with proper variable selection, not just a hopeful list of associations? Was it shrunk to guard against overfitting? Was it validated prospectively in populations unlike the one it was built on? Was calibration checked, or only discrimination? And did anyone demonstrate net benefit at a threshold that matches the clinical stakes? The CSRS is worth studying because its published record answers all five in the open. A score that has only been derived, or only re-tested in its home setting, has not yet earned the same trust.

Sources and further reading

  1. CMAJ 2016 CSRS derivation
  2. JAMA Internal Medicine 2020 multicenter validation
  3. Annals of Internal Medicine 2022 international validation

Questions and answers

Does a high C-statistic mean a score is ready to use?

No. A high C-statistic means the score ranks patients well, but a rule can rank correctly and still produce miscalibrated probabilities. It also says nothing about whether the score holds up in new populations or improves decisions. Discrimination, calibration, external validation, and net benefit are separate questions.

Why prefer sensitivity over a balanced cutoff for syncope?

Because the cost of the two errors is lopsided. Holding a low-risk patient for a few hours is a minor harm; discharging someone who then has a fatal arrhythmia is not. When the downside of a miss is that severe, a rule tuned to catch nearly every serious case is the safer design, and net benefit analysis confirms whether that tuning actually helps.