In a blinded drug trial, identical capsules can conceal allocation from participants and clinicians. A psychotherapy trial rarely has that option. A participant knows whether sessions are occurring. A therapist knows whether the session follows cognitive therapy, supportive counseling, problem-solving, or another model.
The absence of full masking creates several pathways for bias. Expectations may change self-reported symptoms. Therapists may communicate more confidence in one treatment. Participants can seek extra care when disappointed with assignment. Assessors may interpret ambiguous responses differently if they learn the group.
Randomization still protects against baseline confounding, and several trial roles can remain masked; the design challenge is to specify who knew what, choose a comparator that answers the intended question, preserve assessor masking, measure delivery, and analyze all randomized participants appropriately. "Double blind" is too vague to do that work.
Blinding is a set of roles, not one label#
Trial reports once used terms such as single blind or double blind without naming who was masked. Those phrases are especially unhelpful for complex interventions, and a useful report tells you separately about the participants, the therapists, the supervisors, the outcome assessors, the data managers, the statisticians, and the adjudicators.
In psychotherapy, participants know the content they receive, even if they are not told the investigators' preferred hypothesis. Therapists must know which manual or approach to deliver. Supervisors often know assignment because they review fidelity.
An interviewer who conducts a standardized symptom rating may remain masked. A statistician may analyze groups labeled A and B until the main analysis is complete. That partial masking does not recreate a placebo pill, but it blocks avoidable bias at important points.
Allocation concealment still matters#
Blinding happens after assignment. Allocation concealment protects the randomization process before assignment. If a recruiter can predict the next group, conscious or unconscious enrollment choices can create baseline imbalance.
A central randomization system or comparable secure procedure can prevent foreknowledge. Sequence generation, concealment, enrollment, and assignment should be described separately. Psychotherapy's visible treatment does not excuse weak randomization. Baseline prognostic balance remains the foundation for a causal comparison, and concealment is one of the most controllable protections in the design.
Expectations can become part of the observed effect#
People who believe they received the favored treatment may report more hope, practice skills more often, or interpret symptoms differently; people assigned to a less desired group may withdraw, seek care elsewhere, or expect little improvement.
Some of that pathway may be part of how psychotherapy works. Credibility, therapeutic rationale, and willingness to use techniques are not necessarily nuisances, and the problem arises when a trial claims a specific technique caused the entire difference even though contact, expectancy, or disappointment differed substantially. Measuring treatment credibility and outcome expectancy after rationale delivery can help interpret this pathway. Such measures are not a simple adjustment variable, because treatment can itself change expectations.
The comparator defines the claim#
A waitlist asks whether offering the intervention now is better than delaying it. It does not control for clinician time, hope, structure, assessment, or human support. Large differences against waiting can reflect both the therapy's specific components and those broader effects.
No-treatment controls answer a similar additive question but may differ ethically and behaviorally. Participants can obtain outside care, so "no study treatment" is not always no treatment. An attention control matches time and contact while avoiding selected active ingredients. Designing a credible inert psychological control is difficult because listening, validation, goal setting, and monitoring may themselves help.
Usual care is clinically relevant but heterogeneous#
Treatment as usual asks whether adding or replacing care improves outcomes in the real system. Usual care can range from no service to medication, counseling, crisis support, or another structured therapy. It may differ across sites and change over time.
Trials should measure what usual care participants actually receive. A label without service data leaves the comparison undefined. Access outside the study should be recorded in every group.
Usual care often produces a smaller contrast than a waitlist because it contains effective elements. That does not make the trial unfair. It answers a harder and often more useful policy question.
Active comparators test relative benefit#
Comparing two credible psychotherapies can balance session time, attention, expectations, and treatment intensity more closely. The estimate asks whether one approach outperforms another, not whether therapy helps relative to waiting.
The comparison can still be biased if therapists are more skilled in one method, participants prefer one, or one treatment is delivered with greater fidelity; a flexible supportive comparator may be less standardized than the intervention.
Noninferiority designs require special care. A poorly delivered active comparator can make treatments look similar even when the new one would not match competent care, so the margin, assay sensitivity, adherence, and analysis populations need strong justification.
Therapist allegiance can shape delivery#
A therapist who believes strongly in one model may convey enthusiasm, persist through difficulty, and use the manual skillfully, while the same therapist may deliver a comparison approach mechanically or with less confidence. Research teams can also design measures and interpretations around a favored theory.
Allegiance is a nonfinancial intellectual interest, not proof of misconduct. It should be made visible. Trials can balance therapists across conditions when competence permits, use separate therapist pools with equivalent training, and measure beliefs and competence. They can standardize supervision and include investigators with varied perspectives. No arrangement removes every allegiance effect, but transparent roles and balanced quality control let you judge which way it probably pushed.
Treatment fidelity needs more than a manual#
A manual states intended content, dose, sequence, and allowed adaptation. It does not prove that sessions followed it. Fidelity assessment can sample recordings and use raters trained to score adherence and competence.
Raters should ideally be unaware of outcomes. Full masking to treatment may be impossible because techniques identify the approach. Reliability should be reported, along with how sessions were selected and whether ratings covered both groups.
TIDieR asks authors to describe why, what, who delivered, and how. It asks where, when, and how much. It asks about tailoring, modifications, and actual fidelity. That level of detail is what lets you know what was actually compared, and whether another service could reproduce it.
Therapists create clustered data#
Outcomes for participants treated by the same therapist may be correlated because of skill, warmth, or style. They may also be correlated because of caseload or local context. Treating every participant as fully unrelated can understate uncertainty.
Trials should consider therapist effects when randomizing, sizing, and analyzing the study. Multilevel models or robust methods may be appropriate. The number of therapists and distribution of participants per therapist affect what can be estimated.
If each therapist delivers only one condition, therapist and treatment effects can be difficult to separate. If therapists deliver both, contamination and unequal competence become concerns. Design should make the tradeoff explicit.
Self-report is both essential and vulnerable#
Depression, anxiety, pain, intrusive thoughts, and well-being are partly subjective. The participant is the only source for some outcomes. Calling self-report biased does not make it dispensable.
Awareness of assignment can affect reporting. So trials can pair validated self-report with masked clinician ratings, functioning, or school or work participation. They can also pair it with health-service use or other relevant measures. No objective measure is automatically superior if it poorly captures what matters.
Outcome timing should be prespecified. Repeated questionnaires and several scales create multiplicity. A primary measure chosen after looking at results can turn ordinary variation into a positive story.
Assessor masking needs active protection#
Masked assessors should work separately from therapists, avoid reading treatment notes, and remind participants not to reveal assignment, and if disclosure occurs, the trial should document it and, when possible, assign future interviews to another assessor.
Asking assessors to guess assignment can show whether masking was credible, but interpretation is tricky. Correct guesses may arise from genuine treatment response rather than accidental disclosure. So the protection you should be looking for is procedural: separate systems, scripted contacts, recorded breaches, and an analysis that asks whether unmasking could have moved the subjective ratings.
Attrition is rarely random#
Participants may stop because they improve, worsen, or dislike the assigned treatment. They may stop because they face logistical barriers or seek another service. If dropout differs by group and relates to outcome, complete-case analysis can be biased.
Follow-up should continue after treatment stops whenever consent permits. The primary analysis should match the effect of assignment and use justified methods for missing outcomes. Sensitivity analyses can show how conclusions change under plausible departures from missing-data assumptions. Reports need group-specific reasons and timing, not only an overall completion percentage. Treatment attendance and research follow-up are separate: a person can stop sessions yet still provide essential outcome data.
Harms deserve planned measurement#
Psychotherapy can coincide with symptom worsening, distress, or conflict. It can coincide with dependency, stigma, delayed access to other care, or crisis. Some events reflect the underlying condition rather than treatment, but that is also true in drug trials.
Trials should define how adverse events, serious events, and deterioration are collected. They should do the same for suicidality where relevant and treatment-emergent problems. Open-ended reporting alone detects different information from systematic questions. CONSORT Harms 2022 emphasizes prespecified definitions, collection methods, denominators, and absolute results for harms. A statement that no harms occurred tells you nothing until you know whether anyone went looking.
Longer follow-up changes interpretation#
A post-treatment symptom difference may fade, grow, or be offset by relapse. Skills-based therapies often claim durable benefit, so follow-up should match that theory when feasible.
After the treatment period, participants may receive other therapy or medication. Those later services are outcomes of the assigned strategy as well as possible modifiers of long-term symptoms, and they should be measured rather than ignored, because durability, retreatment, crisis care, and functioning are what separate a temporary change on a questionnaire from a sustained clinical benefit.
How to read a psychotherapy trial#
First identify the question the comparator actually created. Then work through the design: how allocation was concealed, how therapists were trained and where their allegiance sat, whether fidelity was rated, who stayed masked, what kind of outcome was measured, how much data went missing, what care people found outside the study, and how many participants each therapist saw.
Look for a preregistered primary endpoint and enough precision to support the conclusion, and keep superiority to waiting separate in your mind from superiority to competent care. Check for deterioration as well as benefit, and for how long anyone was followed.
The inability to mask two roles is a limitation, not an automatic verdict. A credible trial shows where awareness could act and builds protections around the roles, measurements, and decisions that can still be controlled.
Sources and further reading
- CONSORT-SPI 2018 extension for social and psychological intervention trials
- Cochrane Handbook chapter on risk of bias in randomized trials
- TIDieR checklist for complete intervention description
- CONSORT Harms 2022 reporting guideline
- Review of waitlist control conditions in anxiety-disorder research
- Systematic appraisal of researcher allegiance in psychotherapy trials
Questions and answers
Can a psychotherapy trial ever be fully blinded?
Usually not in the drug-trial sense. Participants and therapists tend to know the treatment, while assessors and some analysis roles can remain masked.
Is a waitlist an adequate control group?
It is adequate for asking whether immediate treatment outperforms waiting. It cannot isolate effects of attention, credibility, structure, or contact.
Why not use only objective outcomes?
Many mental-health outcomes are experiences known primarily to the participant. Objective measures can complement valid self-report but may not replace it.
What is therapist allegiance?
It is a therapist's or researcher's preference for a treatment model. It can influence delivery and interpretation, so it should be measured, balanced where possible, and disclosed.
Does lack of blinding make randomization useless?
No. Randomization still balances prognosis on average. Awareness can create post-randomization bias, which requires comparator, measurement, follow-up, and analysis safeguards.