Evidence explainer

Mental and behavioral health

Mental Health Screening Tools: PHQ-9 and GAD-7 Explained

What a PHQ-9 or GAD-7 number actually tells you, and why a positive screen opens a conversation rather than ending one.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. What mental health screening tools actually are
  3. What the PHQ-9 measures and how its score is read
  4. What the GAD-7 measures and how its score is read
  5. Sensitivity, specificity, and why a cutoff is a tradeoff
  6. Why clinicians use screeners: detection, severity, and tracking change
  7. Why a score is a starting point, not a diagnosis
  8. How to think about your own screening result
  9. Reading your score in context

Key points#

What mental health screening tools actually are#

Mental health screening tools are short, standardized questionnaires that ask how often specific symptoms have troubled you over a set period, usually the past two weeks. A high score does not mean you have a diagnosis. It means your reported symptoms are frequent enough that a clinician should look more closely. The PHQ-9 (for depression) and the GAD-7 (for anxiety) are the two most widely used examples, and both were built to flag people for evaluation, not to confirm a condition.

So what does the number actually tell you? It sorts recent symptom burden into rough bands and points the conversation in a direction. A diagnosis is a separate step: it comes from a clinical assessment that weighs your history, how long symptoms have lasted, how much they affect daily life, and other possible explanations. A screen is meant to be cast wide, so it catches people who might otherwise slip past, including those who came in for a sore knee and never mentioned their mood. The screen points; the evaluation decides.

The PHQ-9 and GAD-7 are two of the most widely used examples in general medicine. Both grew out of a primary care instrument built to detect common mental disorders during ordinary office visits, and both are freely available, brief, and easy to score. Because they are self-report, a person often fills them out in the waiting room, which is part of why they fit so naturally into a busy clinic day.

What the PHQ-9 measures and how its score is read#

The PHQ-9 has nine items. Each one asks how often a particular symptom has bothered you over the past two weeks, scored from 0 (not at all) to 3 (nearly every day). Add them up and the total runs from 0 to 27. Because the nine items map onto the recognized symptoms of depression, the questionnaire works as both a screen and a rough severity measure at once.

Clinicians usually read the total against a few familiar thresholds. Scores of 5, 10, 15, and 20 mark the entry points for mild, moderate, moderately severe, and severe symptoms. These bands are not sharp biological lines; they are a convenient way to translate a number into a sense of how heavy the symptom load is right now.

Where does the widely cited cutoff of 10 come from? The 2001 validation study by Kroenke and colleagues compared PHQ-9 scores against a structured diagnostic interview and found that a score of 10 or higher identified major depression with roughly 88 percent sensitivity and 88 percent specificity. At that cutoff the tool caught most people who truly had major depression while correctly clearing most people who did not.

One item deserves special mention. The ninth question asks about thoughts that you would be better off dead or of hurting yourself. Any endorsement of that item, even the mildest, is a signal for direct clinical follow-up rather than something to be averaged into a total and set aside.

What the GAD-7 measures and how its score is read#

The GAD-7 is built the same way, just shorter and aimed at anxiety. Seven items, each scored 0 to 3 over the past two weeks, give a range of 0 to 21. The cut points most often used are 5, 10, and 15 for mild, moderate, and severe anxiety.

The 2006 development study by Spitzer and colleagues reported good reliability and validity, and identified a cut point with about 89 percent sensitivity and 82 percent specificity for generalized anxiety disorder. Those numbers describe a tool that flags most true cases while accepting that a fair share of positives will need a second look.

There is a wrinkle worth understanding. The GAD-7 was designed for generalized anxiety, but it also performs reasonably well as a screen for panic disorder, social anxiety, and post-traumatic stress. That breadth explains why a positive GAD-7 leads to more questions rather than a specific label: a raised score says anxiety of some kind may be present, but not which kind. The 2010 review by Kroenke and colleagues is a useful reference on how these scales behave across settings.

Sensitivity, specificity, and why a cutoff is a tradeoff#

Two words carry most of the weight behind every screening number. Sensitivity is the ability to catch true cases: of everyone who really has the condition, what fraction does the test flag? Specificity is the mirror image: of everyone who does not have the condition, what fraction does the test correctly clear? A perfect test would score 100 percent on both. Real tests never do.

Move the cutoff and you trade one for the other. Lower the threshold and you catch more true cases (higher sensitivity) at the cost of flagging more people who turn out to be fine (lower specificity). Raise it and the reverse happens. No setting makes both problems vanish, which is why choosing a cutoff is a judgment about which error you would rather make.

This is the heart of why a positive screen is not a diagnosis. A positive result raises the probability that a condition is present, sometimes substantially, but it does not confirm it. False positives are expected, built into the design, and the reason confirmation matters. It is also why the same score can carry different weight in different populations: performance depends on how common the condition is in the group being screened. A high score means more in a clinic where depression is common than in a group where it is rare, because a rarer condition produces relatively more false alarms. The 2010 systematic review lays this out in more detail.

Why clinicians use screeners: detection, severity, and tracking change#

So why bother with a questionnaire at all? Three practical reasons.

The first is detection. Depression and anxiety are common, and both are easy to miss during a visit centered on something else. A person may not raise the subject, and a busy clinician may not think to. A standardized screen gives everyone the same chance to be noticed.

The second is grading severity. A score of 6 and a score of 22 both count as positive, but they point toward very different conversations. Putting a rough number on symptom burden helps a clinician and patient decide together how much and how quickly to act.

The third is tracking change. Because the tool is standardized, the same person can complete it again in a month and compare the two numbers. A PHQ-9 that falls from 18 to 9 is concrete evidence that something is moving in the right direction, in a way that "I think I feel a bit better" is not.

These uses are why brief tools fit a primary care visit, where time is short and the agenda is full, and they reflect where the evidence points. The US Preventive Services Task Force recommends screening adults for depression and for anxiety disorders, including during and after pregnancy, in both cases as a B recommendation (moderate net benefit). That recommendation comes with a condition: screening should happen where systems are in place to support accurate diagnosis and follow-up, because a flag with nothing behind it helps no one.

Why a score is a starting point, not a diagnosis#

A diagnosis is more than a number crossing a line. It rests on a clinical evaluation that considers how long symptoms have lasted, how much they interfere with work and relationships, whether a medical condition or medication could be producing them, what role substances might play, and what is happening in a person's life. Above all it includes the person's own account, which a questionnaire can never fully capture.

Physical illness complicates the picture in a specific way. Several depression and anxiety symptoms (fatigue, poor sleep, changes in appetite, difficulty concentrating) overlap with the effects of thyroid disease, anemia, chronic pain, and many other conditions. A score can rise for reasons that have little to do with a mood or anxiety disorder, one more argument for reading the number in context rather than at face value. Language and cultural framing matter too. How a question is phrased, and how a person interprets it, can shift an answer, as can the ordinary variation of a single hard day. None of this makes the tools unreliable; it makes them instruments that need a skilled hand.

This is where structured screening is genuinely useful in generalist primary care. Evidence-based instruments give a clinician a consistent starting point and a shared language with the patient, without pretending to replace the clinical judgment that follows. They can make a first conversation more precise while keeping the decision where it belongs, with the person and their clinician.

How to think about your own screening result#

If you have taken one of these questionnaires, here is how to hold the result. A number is information, not a label. A PHQ-9 of 14 does not define you, and it does not predict your future; it describes how the past two weeks have felt, in a way that can be discussed and acted on.

Bring the result, and any relevant context, to a clinician. Mention the physical symptoms, the recent stresses, the sleep, the medications: all the things a bare score leaves out. Be honest on the self-harm item in particular. There is no benefit to underreporting it, and there is real value in flagging it, because help exists and works.

Expect that follow-up can take several forms: more questions to sharpen the picture, a period of watchful waiting with a repeat screen, a referral, or the start of treatment. Any of these is a reasonable path; a screen simply opens the door.

One firm exception to the calm framing above: if you are in immediate danger or crisis, do not wait for an appointment. Contact your local emergency number or a crisis line right away. Screening tools are designed to start care, not to stand in for a clinician's judgment or for urgent help when it is needed.

Reading your score in context#

Mental health screening tools like the PHQ-9 and GAD-7 do one job well: they turn a set of two-week symptoms into a number that helps a clinician decide what to ask next. That number is a beginning. It flags, it grades, and it can be repeated to track change, but it does not diagnose, and it is not meant to. If your score comes back high, the useful response is not alarm and not dismissal, but a conversation with someone trained to read it in the full context of your life.

Sources and further reading

  1. Kroenke K, Spitzer RL, Williams JB. The PHQ-9 (J Gen Intern Med, 2001)
  2. Spitzer RL, et al. A brief measure for assessing GAD: the GAD-7 (Arch Intern Med, 2006)
  3. Kroenke K, et al. PHQ Somatic, Anxiety, and Depressive Symptom Scales: a systematic review (Gen Hosp Psychiatry, 2010)
  4. USPSTF. Screening for Depression and Suicide Risk in Adults (JAMA, 2023)
  5. USPSTF. Screening for Anxiety Disorders in Adults (JAMA, 2023)
  6. National Institute of Mental Health. Depression and Anxiety health topics

Questions and answers

Does a high PHQ-9 or GAD-7 score mean I have depression or anxiety?

No. A high score means your reported symptoms are frequent enough that a clinician should take a closer look. The validation studies designed these tools to flag people for evaluation, not to confirm a diagnosis. A diagnosis comes from a clinical assessment that weighs your history, how long symptoms have lasted, how much they affect daily life, and other possible explanations.

What do the PHQ-9 and GAD-7 actually measure?

Each asks how often you have experienced specific symptoms over the past two weeks, with every item scored from 0 (not at all) to 3 (nearly every day). The PHQ-9 has nine items (range 0 to 27) focused on depression symptoms, and the GAD-7 has seven items (range 0 to 21) focused on anxiety symptoms. The total is a snapshot of recent symptom burden, not a measure of your character or your future.

Why do clinicians use short questionnaires instead of just talking to me?

They use both. A brief, standardized questionnaire helps detect conditions that are common yet easy to miss, puts a rough number on severity, and can be repeated to track whether things are improving. It makes the conversation more focused, and the US Preventive Services Task Force recommends screening adults for depression and for anxiety disorders when follow up is available. The questionnaire supports the conversation with your clinician; it does not replace it.

Can a screening tool be wrong?

Yes, and that is expected. Every cutoff trades off catching true cases (sensitivity) against correctly clearing people who do not have the condition (specificity), so some results are false positives or false negatives. Accuracy also depends on how common the condition is in the group being screened. That is exactly why a positive screen leads to a fuller evaluation rather than an immediate label.

What should I do if my score is high, especially on the self-harm question?

Share the result with a clinician, who can put it in context and discuss next steps. If you had any thoughts of harming yourself, treat that as a reason to reach out promptly rather than wait. If you are in immediate danger or crisis, contact your local emergency number or a crisis line right away. Help is available, and a positive screen is a starting point for care.

Are these tools valid for everyone?

They are widely validated in primary care populations, but no single cutoff fits every person or setting. Physical illnesses that share symptoms (such as fatigue or sleep problems), language, cultural context, pregnancy, and older age can all affect how a score should be interpreted. This is another reason a trained clinician reads the result in context rather than acting on the number alone.