Living guide

How to Read a Clinical Trial

A practical guide to understanding what a trial proves, suggests, and cannot establish.

Scan Understand Investigate
Understand · why it's complicated

A reported result is never just the drug.

Seven influences run alongside the pharmacology in every psychedelic trial. Separating what the drug did from what belief, support, and later care did is the central interpretive problem in this field.

The drug itselfThe pharmacological effect being tested.
ExpectancyWhat the participant believed would happen.
The psychedelic experienceSubjective intensity, which is hard to standardize.
Psychological supportHours of preparation and integration therapy.
Therapist / site effectsSkill and rapport vary by provider and location.
Outcome assessmentWho scored the result, and how independently.
Later treatmentWhatever happened after the primary endpoint was measured.
↓ converge intoObserved result
Understand · how we got here

Methodology evolution

Six milestones. Click one for detail, an example, and a source.

Understand · the guide

Twelve things to check, one at a time

Grouped by where in a trial's life each issue shows up. Each entry: what to look for, what it can tell you, what it cannot.

Scan
Trial design
01

Blinding

Did participants and raters really stay unaware of assignment?
What to look for
Whether the trial measured if participants correctly guessed their assignment, not just whether it used a placebo.
What it can tell you
Whether the study attempted blinding and what it found when it checked.
What it cannot tell you
Whether blinding actually held — many psychedelic trials that assess blinding report patients correctly guessing their assignment above chance.
Expert view

“Functional unblinding does not bias the observed treatment response, but it potentially biases the specific treatment effect, which is a component of the treatment response.” Szigeti adds that regulators such as the FDA evaluate new drugs mainly by that narrower treatment effect — which is exactly what functional unblinding puts at risk, not a patient’s overall reported improvement.

02

Expectancy

How much of the result is belief rather than pharmacology?
What to look for
Whether the trial measured what participants expected to happen before dosing, and whether that expectation was matched between arms.
What it can tell you
Whether belief in receiving the active drug correlated with the reported outcome.
What it cannot tell you
How much of the total effect is expectancy versus pharmacology — most trials aren’t designed to separate the two.
Expert view

“With the emergence of psychedelics, the problem is so obvious that we can no longer ignore it.” Szigeti says psychiatry spent decades treating blinding as if it neutralized expectancy effects, even though many trials of conventional psychiatric drugs aren’t truly blind either. By his account, the field still has no agreed way to measure expectancy or to fold functional unblinding into trial analysis — a gap he expects will become a research priority in the coming years.

03

Control arms

What was the treatment actually compared against?
What to look for
What the comparator actually was — inert placebo, low-dose active comparator, or waitlist.
What it can tell you
Whether the drug outperformed that specific comparator under these conditions.
What it cannot tell you
Whether it would outperform a different, more rigorous comparator, or standard antidepressant care.
Expert view

“There is no single solution, instead we should embrace a pluralistic approach, acknowledging the pros and cons of various methods.” Szigeti is separately raising funds for what he describes as the first psychedelic Zelen-style trial — a proposed, still-exploratory design that is not yet running. Patients would be treated openly but not told what the alternative study arms are, which he says mostly removes the feeling of being randomized into a "better" or "worse" arm, provided the control condition itself is made compelling enough that patients don’t feel short-changed.

Expert view

Heifets separately points to dose-response as a design feature that strengthens interpretation: the most interpretable trials, he says, include "a dose-response with a meaningful range" — meaning "low-dose but psychoactive doses," not just placebo versus a single full dose.

04

Psychological support

How much therapy came bundled with the dose?
What to look for
How many hours of preparation and integration support came bundled with the dose, and whether that was standardized across the trial.
What it can tell you
The effect of drug plus a specific, described support protocol as a package.
What it cannot tell you
The effect of the drug alone, or whether the same support without the drug would do comparably well.
Expert view

“The quality and quantity of supportive treatment before and after the treatment session. These factors are very likely to influence the long term efficacy, but are going to be very difficult for the FDA to evaluate, and difficult to get reimbursement for.” He adds that despite those obstacles, this factor is "nonetheless critically important" in his view, and suggests the VA health system — not bound by the same reimbursement constraints — as a natural place to test it.

Understand
During and after
05

Rater independence

Who scored the outcome, and how independent were they?
What to look for
Whether the person scoring the outcome was blinded and separate from the person delivering the treatment.
What it can tell you
Whether a scoring bias is plausible given who assessed the outcome.
What it cannot tell you
Whether that bias actually occurred — independence reduces risk, it doesn’t rule it out.
06

Endpoints

Was the primary outcome set before the trial started?
What to look for
Whether the primary endpoint was registered before the trial began, and whether it changed partway through.
What it can tell you
Whether the drug met the specific, predefined measure of benefit.
What it cannot tell you
Whether that measure captures what patients or clinicians consider meaningful improvement.
07

Pain / distress outcomes

Was acute distress during the session itself tracked?
What to look for
Whether distress or discomfort during the session itself was tracked and reported, not just the treatment outcome.
What it can tell you
The rate and severity of acute psychological distress during dosing.
What it cannot tell you
The long-term psychological risk profile — most trials aren’t powered or timed to detect it.
08

Durability

Did the effect last beyond the primary endpoint?
What to look for
How many follow-up time points were measured after the primary endpoint, and how many participants remained in the study.
What it can tell you
Whether benefit was still present at the last measured time point, for those who remained.
What it cannot tell you
Whether benefit would persist indefinitely, or in patients who left the study early.
Expert view

“Long term follow up while still blinded to treatment arm - does not rule out placebo response, but makes this interpretation less likely and certainly less meaningful. If an intervention produces a year long remission, the question about blinding and placebo seems irrelevant.”

Investigate
Interpreting the result
09

Retreatment

Did a "durable" result actually require a second dose?
What to look for
Whether the protocol allowed or measured additional dosing sessions, and how that was counted in the durability result.
What it can tell you
Whether a single dose or a defined dosing schedule produced the reported effect.
What it cannot tell you
Whether that effect holds without retreatment — a "durable" result with an unreported second dose is a different claim.
10

Missing data

Were dropouts excluded in a way that flatters the result?
What to look for
The dropout rate by arm, and whether the analysis used intention-to-treat or completers-only.
What it can tell you
The result for participants who completed the trial as designed.
What it cannot tell you
The result for the full enrolled population if dropout was uneven between arms.
11

Safety attribution

Were adverse events tracked long enough, and attributed carefully?
What to look for
How adverse events were defined, who adjudicated them, and over what time window.
What it can tell you
The rate of pre-specified adverse events during the observed window.
What it cannot tell you
Rare or delayed effects outside that window, or effects the study wasn’t designed to capture.
12

Generalizability

Who was excluded, and how narrow is the trial population?
What to look for
The exclusion criteria — psychiatric comorbidities, medication washouts, cardiac history — and how narrow the enrolled population became.
What it can tell you
The result for a population that met every inclusion criterion.
What it cannot tell you
The result for the broader population of patients who would actually seek this treatment.
Expert view

“[...] the patient population should be representative of the disease population - particularly with respect to psychedelic use history.”

Investigate · side by side

How four real trials were designed

These four trials show how comparator choice, blinding, psychological support, rater independence, follow-up, and retreatment can materially change what a result means.

Compass COMP360, Ph2b Comparator 25 mg vs. 10 mg vs. 1 mg (active low-dose) Blinding Rater-blinded
Comparator
25 mg vs. 10 mg vs. 1 mg (active low-dose)
Blinding
Rater-blinded
Support
Therapist-supported dosing session
Raters
Independent, remote
Durability
12 weeks — effect narrowing by then
Retreatment
No
Active low-dose comparatorRater-blindedIndependent ratersUnblinding not measured
Design detail

n=233 across 22 sites in 10 countries. No guess-the-dose or functional-unblinding check was built into the protocol at any point — not pre-registered, not collected after the fact.

Missing-data handling

Primary endpoint analyzed with a mixed-effects model for repeated measures (MMRM), not last-observation-carried-forward; LOCF was used only for secondary continuous measures.

Key limitation

An outside academic reviewer (University of Edinburgh) noted after publication that participants were never asked whether they could guess their treatment arm — a gap the trial’s own design left unaddressed.

Source: ClinicalTrials.gov — NCT03775200 →
Lykos MDMA-AT, Ph3 Comparator Inert placebo Blinding Functionally unblinded
Comparator
Inert placebo
Blinding
Functionally unblinded
Support
3 prep + 3 dosing + 9 integration sessions
Raters
Independent pool, centralized
Durability
18 weeks (primary)
Retreatment
No formal rule
Inert placeboFunctionally unblindedIndependent rater poolNo retreatment rule
Design detail

This is MAPP2 (n=104, 13 sites) — the larger, more diverse of two nearly identical Phase 3 trials. A companion trial, MAPP1 (NCT03537014, n=90, 15 sites, severe-PTSD-only), used the same design and fed the same FDA review.

Missing-data handling

Primary analysis used a de jure estimand via MMRM with no imputation; a de facto estimand and a tipping-point sensitivity analysis under multiple imputation were run separately.

Key limitation

FDA’s own review named functional unblinding and the resulting expectation bias as the central factor limiting how confidently the efficacy results can be read, citing that about 90% of the drug arm and 75% of the placebo arm correctly guessed their assignment.

Source: ClinicalTrials.gov — NCT04077437 →
Usona psilocybin, Ph2 Comparator 100 mg niacin (active placebo) Blinding Rater-blinded
Comparator
100 mg niacin (active placebo)
Blinding
Rater-blinded
Support
6–8 hrs prep, 7–10 hr dosing, 4 hrs integration
Raters
Independent, by phone
Durability
6 weeks (Day 43)
Retreatment
No
Active comparatorRater-blindedIndependent ratersUnblinding not assessed
Design detail

n=104 across 11 U.S. sites. Niacin was chosen specifically to mimic psilocybin’s physical sensations and help preserve blinding.

Missing-data handling

Intent-to-treat population analyzed via MMRM with no imputation. Dropout was uneven — 2% on psilocybin vs. 21% on niacin.

Key limitation

The trial’s own published discussion states plainly that blinding success was never formally assessed, and notes the uneven dropout may itself be a sign unblinding occurred.

Source: ClinicalTrials.gov — NCT03866174 →
MindMed MM120, Ph2b Comparator 25 / 50 / 100 / 200 µg vs. placebo Blinding Rater-blinded
Comparator
25 / 50 / 100 / 200 µg vs. placebo
Blinding
Rater-blinded
Support
None — safety monitoring only, no psychotherapy
Raters
Independent, by phone
Durability
12 weeks
Retreatment
No
Dose-response designRater-blindedNo psychotherapyUnblinding measured
Design detail

n=198 across 22 U.S. sites. Registered and reported as Phase 2b for generalized anxiety disorder — this corrects an earlier draft that mislabeled it Phase 3.

Missing-data handling

Primary endpoint used multiple imputation (20 iterations); a separate placebo-based imputation handled data missing due to prohibited-medication use.

Key limitation

The one trial of the four that directly measured guess-accuracy: about 85% of the active-dose group and over 65% of the placebo group correctly identified their assignment, though the trial’s own authors argue the lack of effect at the two lowest doses weighs against unblinding driving the result.

Source: ClinicalTrials.gov — NCT05407064 →
Ongoing · maintained

What is changing now

First published September 13, 2026. Future revisions will summarize material methodology and regulatory changes here.

Evolving

Active comparator strategies are evolving

More Phase 3 programs are replacing inert placebo with an active low-dose arm to address functional unblinding.
Increasing

Retreatment is increasingly explicit in protocols

Optional or scheduled second doses are being written into trial design rather than handled informally after the primary endpoint.
Unresolved

Expectancy measurement remains unsettled

No sponsor has yet published a standardized instrument for measuring participant expectancy before unblinding occurs.
Regulatory scrutiny

FDA scrutiny continues around trial conduct

Advisory committee reviews continue to focus on rater independence and adverse-event reporting timelines, following the Lykos review.
Scan · common mistakes

What headlines get wrong

Headline claim

"Double blind" ≠ blinding worked

What to check
Check whether the trial measured and reported guess accuracy — many psychedelic trials that assess blinding report patients correctly guessing their assignment above chance.
Headline claim

"Six-month benefit" ≠ no retreatment occurred

What to check
Read the protocol for optional or scheduled second doses before treating a durability claim as a single-dose result.
Headline claim

"Statistically significant" ≠ clinically meaningful

What to check
A significant p-value describes confidence the effect isn’t zero — not that the effect size matters to a patient’s daily life.
Verify · if you remember only five things

The five questions to remember

1What was it compared with?
2Could participants tell what they received?
3Who measured the result?
4Did the effect last?
5Who does the result actually apply to?
Reference

Expert contributors

  • Balázs Szigeti, PhDUCSF Translational Psychedelic Research Program / Imperial College London
  • Boris Heifets, MD, PhDStanford University School of Medicine
Reference

Sources & guide history

Primary / methodology
  • FDA, Psychedelic Drugs: Considerations for Clinical Investigations (final guidance, July 2026)
  • FDA Psychopharmacologic Drugs Advisory Committee briefing documents, MDMA-AT review
Trial records
  • ClinicalTrials.gov registry filings — see the Regulatory tracker
  • Sponsor Phase 2/3 protocol filings
First published September 13, 2026.