Evidence explainer

Evidence and research methods

How a Perioperative Cardiac Risk Score Is Derived and Validated

A risk score is only as good as the study behind it. Reading a perioperative cardiac risk score well means asking how it was derived and where it was tested, not just what number it returns.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. The question a score is meant to answer
  3. Building the model from data
  4. Why developers hold data back
  5. Two things a good score must do
  6. Does the score travel?
  7. A predictor is not a lever
  8. A short checklist for reading any score

A perioperative cardiac risk score earns trust the same way any honest prediction tool does: someone follows a large group of patients, records who has the complication in question, and lets statistics sort out which baseline features actually separate the two groups. The number a score hands back is only as good as the study behind it, so the real skill you need is not memorizing the score but reading how it was made. The Revised Cardiac Risk Index, published by Thomas Lee and colleagues in Circulation in 1999, is the cleanest example to learn on. It squeezed thousands of surgical cases into six yes-or-no questions, then checked those questions against a fresh set of patients to see whether the predicted risk held up.

Key points#

The question a score is meant to answer#

Before any statistics, a prediction study has to pin down two things: who is being studied and what counts as the outcome. The original index enrolled 4,315 patients aged 50 or older undergoing elective major noncardiac surgery at a single teaching hospital. The outcome was a defined bundle of major cardiac complications: myocardial infarction, pulmonary edema, ventricular fibrillation or primary cardiac arrest, and complete heart block. Everything the score later claims to do is bounded by those two choices. A tool trained on that population and that endpoint is entitled to speak about that population and that endpoint, and no further.

Building the model from data#

To turn the cohort into a usable rule, the investigators used multivariable logistic regression. That method looks at many candidate factors together and keeps only those that still carry predictive information once the other factors are accounted for, which weeds out features that merely travel alongside a stronger one. Six factors survived: a high-risk type of surgery, a history of ischemic heart disease, a history of congestive heart failure, a history of cerebrovascular disease, preoperative insulin treatment, and a preoperative serum creatinine above 2.0 mg/dL. Each factor is worth one point, and the point total sorts patients into ascending risk classes. The elegance is the compression: a messy dataset becomes six questions a clinician can answer at the bedside in under a minute.

Why developers hold data back#

Here is the step that separates a credible score from a flattering one. The team did not build and grade the model on the same patients. They split the cohort, using 2,893 patients to derive the six-factor rule and holding back a separate 1,422 to test it. That split matters more than it looks. Any model tuned on a dataset will describe that dataset a little too well, because it has partly fit the random noise unique to those particular patients. Statisticians call this overfitting. A score that is only ever reported on the data that produced it is close to a self-portrait, and it flatters. Testing on patients the model never saw is the minimum honest bar, and testing in a wholly separate population is better still.

Two things a good score must do#

A score has to succeed at two separate jobs that are easy to blur together.

Discrimination asks whether the score ranks patients in the right order. If patient A scores higher than patient B, does A truly face the greater risk? The usual summary is the c-statistic, equivalent to the area under the receiver operating characteristic curve, which runs from 0.5 (a coin flip) to 1.0 (perfect ranking). In the index validation cohort, the rate of major cardiac complications rose steadily across classes, from 0.4% with no risk factors to 0.9% with one, about 7% with two, and roughly 11% with three or more. That orderly climb is discrimination you can see.

Calibration asks a different question: do the predicted percentages match reality? A score can rank flawlessly and still be systematically too high or too low. Suppose a model labels a group as 5% risk and the group actually has 15% events. It discriminates, but the numbers are wrong, and that matters because clinicians and patients act on the number itself, not on its rank in a list. A score that ranks well but predicts poorly can still steer a decision the wrong way. Honest evaluation reports both properties rather than letting one stand in for the other.

Does the score travel?#

Internal validation on a held-back sample is a floor, not a ceiling. The harder test is whether a score keeps working in hospitals, populations, and years its developers never touched. A 2010 systematic review by Ford, Beattie, and Wijeysundera pooled 24 studies covering roughly 792,740 patients and found that the Revised Cardiac Risk Index discriminated moderately well for mixed noncardiac surgery, with an area under the curve near 0.75. That durability across many settings is why the tool is still taught decades after it appeared.

The same review supplied the warning label. The index predicted cardiac events poorly after vascular surgery specifically, and it predicted death poorly. A score is validated for a defined population and a defined outcome, not as a universal instrument. Reaching for it in a different operation, a different endpoint, or a very different patient mix is precisely where prediction tools tend to fail without anyone noticing.

A predictor is not a lever#

The six factors earned their spots by association, not by proven mechanism. Preoperative insulin use, for instance, flags more advanced diabetes rather than causing cardiac arrest, and creatinine is a marker for kidney function and vascular disease. A prediction model answers how likely an outcome is; it does not explain why, and it makes you no promise that changing a factor will change the outcome. Treating a predictor as a lever, and pushing the number down in the hope the risk follows, confuses the map for the territory. Only an intervention study, not a risk index, can show that acting on a factor actually moves the outcome.

A short checklist for reading any score#

A handful of questions carry over to almost any clinical prediction tool. In what population was it derived? What exact outcome does it predict, and over what time window? Was it tested only internally, or also externally in independent cohorts? And is the reported performance about discrimination, calibration, or both? Answer those, and read the result as a probability that describes a group rather than a verdict pronounced over one person.

Sources and further reading

  1. Lee et al., Circulation 1999
  2. Ford et al., Ann Intern Med 2010

Questions and answers

Is a high c-statistic enough to trust a score?

No. The c-statistic captures ranking, not accuracy of the predicted percentages. A score with strong discrimination can still be poorly calibrated and quote risks that are too high or too low, so both should be reported.

Why can a validated score still be wrong for my situation?

Because validation is always for a specific population and outcome. A tool that performs well for mixed noncardiac surgery may perform poorly for vascular surgery or for predicting death, as the pooled evidence on this index showed. Individual perioperative decisions belong to a patient and their own clinicians.