A suicide-risk score can combine history, symptoms, and current circumstances into a category such as low, medium, or high. Such categories may correlate with outcomes across groups. They cannot determine whether one person will attempt or die by suicide, and they should not be used alone to decide who receives treatment, observation, admission, or discharge.
The central problem is not that risk factors are meaningless. Prior suicidal behavior, current intent, and severe agitation matter. So do access to lethal means, intoxication, and psychosis. So do acute loss, pain, and many other factors. The problem is converting population association into an individual forecast for a rare, dynamic event.
Large reviews show the mismatch. In one longitudinal meta-analysis of psychiatric populations, high-risk groups had higher relative odds of suicide, yet about 95 percent of people classified high risk did not die by suicide and 44 percent of suicides occurred in the lower-risk group. A strong group association can coexist with poor individual prediction.
The base-rate problem#
Positive predictive value asks: among people classified positive or high risk, how many later have the outcome? Sensitivity and specificity are only part of that answer. How common the outcome is in the population you are testing matters just as much.
Suicide death is rare even in psychiatric settings. When there are many more people who will not die by suicide than people who will, a modest false-positive rate produces a large number of false positives. Improving sensitivity often increases that number further.
This does not mean a false-positive person has no distress or no need for help. "False positive" refers only to the specific predicted outcome within the study window. Someone may still need urgent treatment for severe depression, psychosis, or substance withdrawal. The need can come from abuse, homelessness, pain, or another crisis.
What the 2016 meta-analysis found#
Large and colleagues pooled 37 studies containing 53 patient samples, more than 315,000 psychiatric patients, and 3,114 suicide deaths. The high-risk group had pooled odds of suicide 4.84 times those of the lower-risk group. That relative difference was statistically substantial.
Predictive performance told a less reassuring story. Pooled sensitivity was 56 percent, so 44 percent of suicides occurred in the lower-risk group. Specificity was 79 percent. Positive predictive value was 5.5 percent, which means about 94.5 percent of those classified high risk did not die by suicide during follow-up.
The studies were highly heterogeneous. Populations, tools, and thresholds differed. So did follow-up periods and clinical settings, and the I-squared statistic was about 93 percent. A pooled number therefore does not tell you the performance of every instrument in every clinic. The review also did not show improvement in predictive strength over several decades: adding more risk factors has not overcome the rare-event and dynamic-state problems enough to create a dependable individual forecast.
What the 2017 risk-scale review added#
Carter and colleagues reviewed clinical instruments intended to predict suicide or self-harm. For suicide, the pooled positive predictive value of high-risk categorization was again around 5.5 percent. Performance varied by outcome and instrument, but no scale provided a reliable basis for allocating care by a simple cutoff.
Different outcomes should not be collapsed. Predicting nonfatal self-harm, suicide attempt, and suicide death involves different base rates and meanings. A tool may be more sensitive for one and less useful for another. Marketing a scale as a general "suicide predictor" hides those distinctions from you. The reviews themselves evaluated published studies, which can be affected by selection, missing follow-up, changes in care after assessment, and differences in how deaths are classified; those limitations add uncertainty, and they do not rescue deterministic prediction.
Sensitivity and specificity do not answer the care question#
Sensitivity is the proportion of people with the later outcome who were classified high risk. Specificity is the proportion without the outcome who were classified lower risk. Neither says what intervention should follow a category.
Even a tool with improved sensitivity could be clinically unhelpful if it labels almost everyone high risk. A more specific threshold could reduce false positives while missing more people who later die. There is no harmless cutoff because both error types matter. The care decision also depends on whether an intervention benefits only the high-risk group, and many supportive actions, such as safety planning, follow-up, treatment of depression or substance use, and help with practical crises, may be appropriate based on need rather than on a predicted probability.
Risk changes faster than many scores#
Suicidal crises can intensify or recede over hours. Intoxication, withdrawal, or interpersonal conflict can change the situation. So can legal events, sleep loss, or pain. So can access to a firearm, a new diagnosis, or loss of housing. A score recorded yesterday may not describe today.
Protective conditions also change. A person may contact a supportive relative, agree to secure lethal means, begin treatment, or move to a safer setting. These changes are not guarantees, but they alter the immediate plan.
Static factors such as a remote attempt or diagnosis matter for background risk. They cannot specify the timing of a future act. Dynamic assessment and repeated contact are more appropriate than treating a category as a permanent trait.
Why low risk can be a dangerous label#
A lower-risk classification can end inquiry. Staff may reduce observation, shorten assessment, omit follow-up, or reassure family members beyond what evidence supports. Because a large share of suicides occurs outside the high-risk category, "low" cannot mean safe.
The label can also make it harder for a person to ask for help later. They may believe their distress is not severe enough or that services will reject them. A needs-based plan avoids promising that the future has been predicted. Documentation can state the current findings and limits instead: whether thoughts, intent, planning, means, psychosis, intoxication, recent behavior, and supports are present; what changed; what actions were taken; and when reassessment will occur.
Why high risk can also cause harm#
A high-risk label may lead to proportionate lifesaving action when immediate danger is present. Used as an automatic score outcome, however, it can drive coercive care, stigma, loss of autonomy, and resource use without considering the person's actual needs or the least restrictive safe option.
False positives are especially important when a category determines admission or discharge. A tool cannot account fully for treatment response, housing, or family context. It cannot account for capacity, local alternatives, or a collaborative safety plan. This is not an argument against hospital care. It is an argument for making that decision from a full assessment and current safety needs, not a numerical threshold.
What NICE recommends instead#
The NICE guideline on self-harm says not to use risk assessment tools or scales to predict future suicide or repetition of self-harm. It also says not to use them to determine who should receive treatment or who should be discharged, and not to use global low-medium-high stratification for those purposes.
NICE recommends focusing assessment on the person's needs and on supporting immediate and long-term psychological and physical safety. Mental health professionals should develop a risk formulation as part of psychosocial assessment.
Formulation is not a hidden score. It organizes the person's history, current state, and precipitating and maintaining factors. It organizes available supports, foreseeable changes, and planned responses. It should remain provisional and be revised when circumstances change.
Tools can structure questions without becoming verdicts#
Standardized instruments may help ensure that important questions are asked, support communication, or monitor symptoms over time. Their use should match their validation and purpose. A screen for current ideation is not a prediction model, and a symptom measure is not a discharge rule.
Scores should be documented alongside the answers and clinical context. A single total can hide a critical item, such as current intent, or overemphasize a historical factor. Clinicians need to respond to the content, not merely the sum. Automated models built on electronic health records face the same conceptual limits, plus risks from missing data, coding bias, changing practice, and unequal performance; higher statistical discrimination does not by itself establish safe clinical use.
What a needs-based assessment includes#
Assessment commonly considers current thoughts, intent, plan, access to means, recent preparatory behavior, past attempts, self-harm, mental state, substance use, physical illness, pain, trauma, abuse, acute loss, sleep, and treatment changes. It also asks about support, responsibilities, coping, reasons for living, and willingness and ability to use a safety plan.
The conversation should identify modifiable needs. These may include treating agitation or psychosis, managing withdrawal or pain, securing lethal means, arranging a safe place, involving a trusted person with consent when possible, scheduling rapid follow-up, and connecting to ongoing care.
A collaborative safety plan names warning signs, internal coping strategies, people and places for distraction or support, professionals and crisis services, and steps to make the environment safer. It is not a contract promising no self-harm.
Immediate safety steps#
Current intent, a plan with available means, or a recent attempt requires urgent professional help. So does escalating preparation, severe intoxication, psychosis with dangerous commands, or inability to maintain immediate safety. Call emergency services or go to an emergency department when danger is immediate. Do not leave a person at imminent risk alone.
If you are in the United States, call or text 988 for the Suicide & Crisis Lifeline, and call 911 for an immediate life-threatening emergency. A trusted person can help reduce access to firearms, medications, or other lethal means while help is arranged. If you are somewhere else, contact your local emergency number or crisis service.
Sources and further reading
Questions and answers
Does a relative odds ratio near five mean prediction is accurate?
No. It shows a group association. In the same meta-analysis, positive predictive value was about 5.5 percent and sensitivity about 56 percent, leaving many false positives and false negatives.
Is a person labeled low risk safe to discharge?
The label cannot answer that. Discharge requires a full current assessment, a feasible safety and follow-up plan, and judgment about needs and the least restrictive safe setting.
Should risk tools be abandoned completely?
They may help structure questions or monitor defined symptoms when used for a validated purpose. They should not be treated as individual forecasts or automatic care-allocation rules.
What is risk formulation?
It is a reasoned, revisable account of the person's history, current state, triggers, supports, likely changes, and safety plan. It communicates uncertainty and action rather than assigning a permanent tier.
What should someone do if suicidal thoughts are present now?
Seek immediate help. In the United States, call or text 988; call 911 or go to an emergency department if danger is immediate. Do not rely on an online score to decide whether the situation is serious enough.