Oregon’s licensed psilocybin service centers served 5,935 clients in 2025, and Colorado has opened a similar program. Three datasets published this year describe what happened to some of the people who use these services. Read side by side, they agree on one point and diverge on another, and the divergence follows the method.
All three find that serious events during or just after a dosing session are uncommon. The one study that went back to clients and asked them, for three months, found more: emergency department visits the facilitator never recorded, a small group whose depression got worse, and a few people reporting new thoughts of dying or suicide.
None of this shows that state psilocybin services are safe, or that they are unsafe. It shows that the answer depends on who is asked and for how long.
Three datasets, three ways of counting
The first is the state’s own. Feliciano Yu and colleagues, writing in Frontiers in Psychiatry in May, analyzed Oregon’s public reporting dashboard for calendar year 2025: 5,935 clients across 5,375 sessions. Service centers reported 13 behavioral and 15 medical adverse events, or 2.42 and 2.79 per 1,000 sessions. The paper describes an adverse reaction in this system as one requiring emergency services or contact with a medical provider during a session. The data are aggregate totals submitted by each center. They cannot follow an individual after the session, and the state collects no symptom measures.
The second comes from a software company. Scott Thompson and colleagues at the University of Colorado, working with Althea staff, analyzed records from Althea, a platform that facilitators in both states use to document sessions. Their preprint, posted to medRxiv in August and not yet peer reviewed, covers 2,363 dosing records entered by 253 facilitators between December 2024 and March 2026. Eighty-two percent were in Oregon. Facilitators logged adverse events on the dosing day and at a later integration session. Three of the authors co-founded Althea, the company paid for the data collection, and the university held equity in it during the study.
The third is a research cohort. Todd Korthuis and colleagues at Oregon Health and Science University enrolled 346 consenting clients at 24 of Oregon’s 26 active service centers and surveyed them at one week, one month and three months. Ninety percent were still responding at three months. Their study, published in JAMA Network Open in August and covered here at the time, collected reports from facilitators and, separately, from the clients themselves.
A fourth, smaller dataset is not examined here: a preprint posted in February that surveyed 88 clients of Oregon services for 30 days. It also asked clients directly, and it has not been peer reviewed.
Where they agree
Events recorded around the session are rare in all three.
Oregon’s dashboard shows 28 adverse events in 5,375 sessions. Thompson’s records list mostly nausea, headache and anxiety, described as mild and transient, plus five events the authors call more serious: two emergency calls, one request for emergency help for emotional distress, one client who later said she had a heart attack on dosing day, and one fall with no reported injury. The paper says there was no follow-up on any of them and describes them as not clearly related to treatment. In the Korthuis cohort, facilitators reported a single serious reaction during a session, a client taken to hospital after violent behavior.
Where the picture changes
Korthuis counted four serious adverse reactions, not one. The other three came from clients’ own follow-up surveys. Two reported no problem that the facilitator had recorded during the session, and the third had only mild stomach discomfort noted. All three later reported seeking emergency care they connected to the experience, along with new medication and worsening anxiety or depression. Two reported new thoughts of death, suicide or self-harm.
The authors draw the conclusion themselves. Facilitators reported only one of the four serious reactions, they write, which suggests that state reporting rules focused on the day of services may underestimate the true rate.
The follow-up surveys also recorded smaller changes that an incident report would never contain. At three months, 14 of 312 respondents (4.5 percent) reported persistent, worsening depression, up from 9 at one week. Six (1.9 percent) reported new thoughts of dying or suicide, up from three. Seven (2.3 percent) agreed that the experience had been harmful.
Those figures sit inside a favorable average. The cohort’s mean depression score fell from 9.3 to 5.2 on the PHQ-9, and the share with moderate to severe symptoms fell from 42.2 percent to 16.5 percent. Most people improved and a few got worse. An average cannot show both.
Why the rates cannot be set side by side
It is tempting to line up the numbers: about 5 events per 1,000 sessions in the state data, 5 serious events in 2,363 records on the Althea platform, 4 in 346 clients in the cohort. They do not measure the same thing.
The state counts events that required emergency or medical contact during a session, as the Yu paper describes it; the Korthuis authors describe the mandate as covering events within three days of services. Thompson’s facilitators entered free-text descriptions under state categories that, the authors found, were applied inconsistently. Korthuis counted reactions reported by either party over three months, including care a client sought weeks later and attributed to the session. A client’s attribution is not a clinical judgment, and with four events the estimate is imprecise.
What can be said is narrower. In the one dataset that contains both channels for the same people, the client channel found three serious reactions that the facilitator channel did not.
The suicidality finding in the larger dataset
Thompson’s abstract reports “evidence of possible risk of increased suicidality.” The full paper shows how limited that evidence is.
The authors looked at one item on the PHQ-9 depression questionnaire, which asks about thoughts of being better off dead or of self-harm. They counted any answer of “several days” or more as risk. Among 103 people who reported no such thoughts before their session and answered again two weeks later, 11 reported some. In eight of those 11, the total depression score had also risen or stayed the same. Among 28 who had reported such thoughts beforehand and answered again, 22 reported them less often and 6 reported no change.
There is no comparison group, so the paper cannot say whether 11 of 103 is more than would occur over any two weeks among people who had screened positive for depression. The item captures passive thoughts as well as intent. No case was reviewed for cause, and the paper reports no attempts or deaths. The authors’ own wording is that they found de novo suicidal thinking in “several” participants in a small sample, and they call for monitoring to be “considerably strengthened.”
That is a monitoring signal. It is not evidence that the services cause suicidal thinking, and it is consistent with what the Korthuis cohort shows: improvement on average, deterioration in a few.
The same paper reports that 44 clients disclosed thoughts of suicide at intake and that all received psilocybin, although it says Oregon does not allow service to clients who screen positive for suicidal ideation. It does not break the 44 down by state. Another 48 reported using antipsychotic medication, and all were dosed, in what the authors call “apparent contravention” of both states’ regulations. The authors add that some of them may not have been taking the drugs at the time.
What the outcome numbers rest on
Both outcome datasets find that clients reported lower depression and anxiety scores after a session. Neither has a control group, so neither can attribute the change to psilocybin as distinct from preparation, support, expectation or time.
The Thompson figures need particular care. The paper reports a 49 percent improvement in PHQ-9 scores and 51 percent in GAD-7 anxiety scores at two weeks. Those results come from 138 and 181 people, not 2,363. The full questionnaires were optional and were offered only to clients who screened positive on a two-item version, and about 38 percent of those eligible answered both times. People selected for high scores tend to score lower on retest whatever happens in between, and the authors acknowledge that selective dropout may bias the results toward the positive.
The large number in that study is a count of dosing and safety records. On symptom outcomes, the 346-person cohort is the bigger and better-followed dataset.
Thompson is also the only one of the three to examine dose and subjective intensity. Improvement did not differ significantly between clients who received 30 mg or less and those who received more, and scores on a mystical-experience questionnaire correlated only weakly with symptom change. Clients who responded did report higher mystical-experience scores than those who did not. These are small subgroups in one dataset, and they do not establish that higher doses add nothing.
What the cohort cannot settle
The gap between the cohort and the records may not be all method. The Korthuis cohort is about 5 percent of Oregon’s clients in that period, and people who agree to three months of surveys may differ from those who do not. Nearly two-thirds of the cohort said they came to treat mental health symptoms, which may put them at higher risk than the wider client base. Some worsening depression would be expected in any group of that kind over three months, with or without a session. And the authors themselves judge the services “generally safe.”
The two channels also disagreed in the other direction. Facilitators recorded nausea or vomiting in seven further clients, and six of those seven reported no challenges or adverse effects in any survey. Neither record is complete on its own.
Each of those points limits how far the cohort’s numbers can be generalized. None of them explains why the facilitator’s record and the client’s account of the same session differ.
What the three datasets establish together
The limit on evidence from state psilocybin programs is increasingly the way it is collected: incident reports filed by centers, notes entered by facilitators, optional questionnaires, and whether anyone contacts the client afterward. A program whose surveillance ends within days of the session will record fewer harms than one that asks again at three months. That requires no bad faith from anyone. It follows from where the reporting window closes.
All three sets of authors say as much. Yu’s group notes that low event counts may reflect incomplete reporting in a non-medical framework. Thompson’s calls the adverse-event data limited by lack of detail and follow-up. Korthuis recommends that states adopting these programs assess client safety over the months after a session.
Colorado’s results are not reported separately in any of the three. None of the three follows clients past three months. And the most complete picture so far came from a research network funded in part by federal grants, not from the program’s own reporting rules.