Evidence explainer

Physician-scientist and medical humanities

Can Empathy Be Measured? What the Jefferson Scale Actually Captures

The Jefferson Scale measures a self-reported, mostly cognitive orientation toward understanding patients. It cannot grade an individual clinician.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. Start with the definition, not the questionnaire
  3. Inside the instrument
  4. Why psychometricians trust it
  5. The outcome question, and where the evidence thins
  6. The ceiling is built in
  7. Reading a score without overreading it

Empathy can be measured, but only in the disciplined sense in which any psychological quality is measured: you write a precise definition, build questions that sample it, and then check whether the resulting scores behave the way the definition predicts. Judged that way, the Jefferson Scale of Empathy (JSE), the instrument used more than any other in medicine, does its job well for one particular idea of empathy: a largely cognitive ability to grasp a patient's perspective and to signal that you have grasped it. The catch is what that number is not. It is a self-reported orientation, not a recording of what happens at the bedside, and that gap fixes a hard limit on how far any single score can be pushed.

Key points#

Start with the definition, not the questionnaire#

Every empathy score is really a claim wearing a number, and the claim is set before the first item is written. When researchers say they are measuring empathy, they have already decided which empathy. The developers of the JSE, Mohammadreza Hojat and colleagues at what is now the Sidney Kimmel Medical College at Thomas Jefferson University, made that decision explicit around 2001. They framed empathy in patient care as a cognitive attribute: understanding a patient's pain, suffering, and point of view, joined to an intention to help and the ability to communicate that understanding.

That framing does real work. The scale deliberately draws a line between empathy, treated as understanding, and sympathy, treated as sharing the patient's feelings. So the JSE is not weighing empathy in some universal, everyday sense. It is measuring the specific construct its authors chose to name empathy, and every finding has to be read inside that choice. Miss the definition and you will misread the score.

Inside the instrument#

Mechanically the JSE is modest: 20 statements answered on a 7-point agreement scale. It ships in three parallel versions, one for medical students, one for practicing health professionals, and one for students in other health professions, so the same construct can be tracked across a career. Its items load onto three components the authors describe as perspective taking, compassionate care, and a smaller factor they call standing in the patient's shoes.

Nothing about the format is exotic. What makes the JSE useful is that its narrow, stated definition keeps the questions pointed at one thing rather than sprawling across every quality a warm clinician might have.

Why psychometricians trust it#

By the yardsticks the field uses, the JSE holds up. The large nationwide study by Hojat and colleagues surveyed roughly 6,000 first-year osteopathic medical students across 41 campuses and reported strong internal consistency, with a Cronbach's alpha near 0.82, and a stable three-part structure that reappeared under confirmatory analysis. That work also produced national norm tables and confirmed a pattern seen earlier, that women tend to score modestly higher than men.

The scale has since been translated into dozens of languages and carried into studies worldwide, which is a large part of why it dominates the literature. So when a paper calls the JSE "validated," the meaning is concrete: it is reliable, its internal structure survives testing, it separates groups the way theory predicts, and it tracks alongside related measures. As an instrument for the construct it was built to measure, it is a good one.

The outcome question, and where the evidence thins#

The more ambitious claim is that measured empathy tracks how patients actually fare, and it leans mainly on two studies. In a 2011 Academic Medicine report, Hojat and colleagues connected higher physician JSE scores to better control of hemoglobin A1c and LDL cholesterol among 891 patients with diabetes cared for by 29 family physicians. The 2012 Del Canale study in Parma, Italy, looked at more than 20,000 patients with diabetes under 242 primary care physicians and found that higher physician empathy scores went with fewer acute metabolic complications severe enough to send someone to the hospital.

Both are large and carefully executed, and they point the same way. They are also correlational, confined to a single disease, and unable to show that empathy caused the difference rather than traveling with something else, such as thoroughness or continuity of care. And the direction is not unanimous. A 2019 cross-sectional analysis in the Journal of General Internal Medicine ran a comparable test and found no association between physician JSE scores and diabetes laboratory outcomes. Because the samples and settings differed, these studies are not clean replications of one another. The honest summary is that empathy plausibly matters, the evidence is mixed rather than closed, and a self-reported score is at best a distant proxy for whatever clinical mechanism is doing the work.

The ceiling is built in#

The deepest limit is not a flaw to be patched; it is structural. The JSE asks clinicians to rate their own attitudes, so it reports disposition, not conduct. Asking people to score themselves on a socially prized trait invites social-desirability bias, and most of us overrate our own perspective-taking. A high number is a statement about how a clinician sees their own orientation, not a transcript of the encounter.

That is exactly why careful researchers pair it with instruments that look through a different window. Patient-rated tools such as the CARE measure ask the person on the receiving end how understood they felt, and standardized-patient encounters score behavior an observer can actually see. Self-rated and patient-rated empathy do not always agree. Other scales encode other definitions entirely: the widely used Interpersonal Reactivity Index, for one, includes an affective component the JSE intentionally leaves out. So "measuring empathy" always smuggles in a decision about what empathy is, and instruments that answer that question differently are not interchangeable.

Reading a score without overreading it#

Treat a JSE result as evidence about an attitude within a stated framework, never as a verdict on a clinician's care or a patient's experience. It shines on group-level questions, such as whether empathy scores drift downward across training or whether a new curriculum nudges them upward, and it misleads the moment it is used as a personal report card. This is educational writing about measurement rather than clinical advice, and the same discipline applies to any figure that promises to quantify a human quality: the score is only ever as wide as the definition behind it. The Jefferson Scale answers its own question honestly. The skill is remembering which question that was.

Sources and further reading

  1. Jefferson Scale of Empathy nationwide measurement study, Hojat et al., Adv Health Sci Educ (2018)
  2. Hojat et al., Physicians' Empathy and Clinical Outcomes for Diabetic Patients, Academic Medicine (2011)
  3. Del Canale et al., Physician Empathy and Disease Complications, Academic Medicine (2012)
  4. Physician Empathy Is Not Associated with Laboratory Outcomes in Diabetes, J Gen Intern Med (2019)

Questions and answers

Does a higher Jefferson Scale score mean a doctor is more empathetic?

It means they report a stronger cognitive orientation toward understanding patients, as the scale defines it. Because the rating is self-generated and covers disposition rather than observed behavior, a higher number does not certify that any given patient will feel more understood.

Why do studies of empathy and outcomes disagree?

The two influential diabetes studies were large and positive, but a later study found no association. They differed in samples, settings, and design, and all were correlational, so they show a plausible link rather than proof of cause, and they do not cleanly replicate each other.

Is self-report a reliable way to measure empathy?

It is reliable in the technical sense that scores are consistent, but self-report of a valued trait is prone to social-desirability bias. That is why researchers often add patient-rated tools like the CARE measure or standardized-patient encounters that capture behavior an outside observer can see.