Ten days after a sham-controlled trial found no measurable benefit from personalizing accelerated theta-burst stimulation for treatment-resistant depression, a French group posted a revised preprint pointing the other way for a different condition. In MULTIMODHAL, low-frequency rTMS aimed at each patient’s hallucination-related brain activity, located with a functional MRI scan taken while the patient was hallucinating, reduced the severity of drug-resistant auditory hallucinations in schizophrenia more than the same stimulation placed at a standard scalp landmark. The study has not been peer reviewed.

It is tempting to read the pair as a scorecard on personalized neuromodulation. The trials do not support that. They differ in indication, in what was personalized, in the stimulation protocol, and above all in what the personalized arm was compared against. MULTIMODHAL shows that where the coil goes can matter for hallucinations when the alternative is an imprecise scalp target. The depression trial asked whether personalized stimulation beats no real stimulation at all.

What MULTIMODHAL tested

The trial randomized 73 adults with schizophrenia and auditory-verbal hallucinations persisting despite stable antipsychotic treatment, 36 to fMRI-guided rTMS and 37 to standard targeting. Every participant first had a symptom-capture fMRI scan to map the brain networks active during their hallucinations, and only patients who reported at least one hallucination during the scan went on in the trial. Both groups then received the same active protocol: ten sessions of 1,200 pulses at 1 Hz, twice daily over five days, set at 100 percent of resting motor threshold and delivered at roughly 90 percent on average. In the guided arm, neuronavigation placed the coil over the individual target. In the standard arm, it went to T3P3, a point on the left temporo-parietal scalp located with an elastic measurement cap, the conventional placement for this use.

To keep patients blind, the standard arm wore the cap and the guided arm had a neuronavigator mounted on the coil. Only the TMS operator knew the allocation, and that operator did no clinical ratings. The authors report a Bang blinding index of 0.16, below the 0.2 threshold they treat as successful blinding.

The trial ran for a long time at two centers, CHU Lille and GHU Paris, recruiting from June 2011 through February 2022. It was funded by a French Ministry of Health hospital research grant. The corresponding author discloses invitations to meetings and expert boards from Janssen, Lundbeck, Otsuka, and Rovi during the trial period, which he describes as unrelated to the work.

The result, and the endpoint change behind it

On the Auditory Hallucination Rating Scale one month after treatment, the guided group improved by 6.1 points on average and the standard group by 0.7. The baseline-adjusted difference was 5.4 points (95% CI 1.9 to 8.9) in the guided arm’s favor, and the result held under multiple imputation and in the per-protocol population. Using a threshold of at least a 25 percent reduction, 45 percent of guided patients responded against 18 percent of standard patients, a number needed to treat of 3.5, with a wide confidence interval of 2.0 to 12.4. Hallucination-specific secondary measures, including the PANSS positive subscale, moved in the same direction. Broader measures did not: PANSS total, global functioning, and clinical global impression showed no significant difference between arms.

The rating scale that carries the headline was not the trial’s registered primary outcome. The ClinicalTrials.gov record, NCT01373866, still lists patient-rated visual analogue scales for hallucination severity and frequency as primary, and the sample size was calculated on them. The authors say they switched to the AHRS before unblinding because about 90 percent of enrolled patients turned out to have predominantly auditory hallucinations, for which the AHRS is the more specific instrument. The two versions of the preprint describe the timing slightly differently. The January version places the decision after recruitment was complete; the September version describes a mid-trial adaptation that became apparent during screening and enrollment.

The switch does not appear to have rescued a failed result. The registered measures, now reported as key secondary outcomes, also favored guided stimulation at one month: the intensity scale by 1.6 points on a 10-point scale (95% CI 0.7 to 2.5), which meets the 1.5-point difference the trial was powered to detect, and the frequency scale by a smaller margin, a median difference of half a point. A reader should still know the headline number comes from an endpoint adopted after the trial was under way, not the one registered, and that the public registry does not reflect the change.

What “sustained” rests on

The preprint reports that the advantage held at three and six months, with adjusted differences of 5.3 and 5.5 points, and faded to non-significance at 12 months. The numbers behind those estimates shrink quickly. The protocol let patients leave once their hallucinations returned to baseline severity, and 44 did so after three months. AHRS data were available for 23 and 24 patients at three months, 21 and 20 at six, and 10 and 14 at twelve. Because the patients most likely to leave were those whose symptoms had come back, the remaining sample is not a random subset of either arm. The authors’ own multiple-imputation analysis produced smaller effects at three, six, and twelve months, and the September version adds that the attrition “warrants caution” beyond three months.

The one-month result is the trial’s evidence. The durability claim is a signal worth testing, not an established finding.

What changed between the two versions

The September 26 revision, which follows the January 15 original, does not change the main estimates. It adds “multicenter phase 3” to the title, restructures the abstract, and changes the abstract’s patient count from 70 to 73, the number randomized in the results. It fixes sign errors in several confidence intervals, including the three-month interval, which the first version printed as crossing zero. It replaces the earlier chi-square blinding check with the Bang index, adds a sensitivity analysis testing whether the eleven-year recruitment period biased the result, states that participants who were randomized but withdrew before stimulation were replaced, and adds the attrition caveat.

It also drops a comparison. The January version argued that its effect sizes against an active comparator were comparable to those of a 2024 sham-controlled trial of resting-state-guided rTMS for hallucinations. That sentence is gone.

The phase label deserves a note of its own. The authors call MULTIMODHAL a phase 3 trial. The registry, as is standard for device and procedure studies, lists its phase as not applicable. The label is the investigators’ description of where they place the trial in development, not a regulatory designation, and it describes a 73-patient study at two sites. The registry also lists 85 enrolled against the 73 randomized in the paper, a gap the preprint’s replacement procedure may partly explain but does not reconcile explicitly.

Why this is not the aiTBS trial reversed

The depression trial, published by a UC San Diego group in the Journal of Affective Disorders on September 16, randomized 51 patients with treatment-resistant depression to accelerated intermittent theta-burst stimulation individualized by resting-state connectivity, the same individualized by connectivity and EEG frequency, or sham. Neither active arm separated from sham on the depression scale at one week or on an anhedonia measure. Several differences separate it from MULTIMODHAL, and each would matter on its own.

The comparator is the largest. The aiTBS trial asked whether personalized stimulation beats sham. MULTIMODHAL asked whether guided stimulation beats a different active placement, one that in this trial barely moved symptoms at all. A win over a weak active comparator shows that placement matters. It does not show how much guided rTMS helps compared with no active stimulation, since the guided arm’s 6.1-point improvement includes whatever placebo and regression effects any treated group experiences. The active-comparator design has an advantage: both arms received the same attention, sensations, and visible technology, so expectancy cannot easily explain the gap between them.

The signal used for targeting also differs. MULTIMODHAL located targets from brain activity recorded during the symptom itself. The aiTBS trial used resting-state connectivity, a measure of how brain regions co-fluctuate at rest. MULTIMODHAL’s own secondary analysis found that baseline resting-state connectivity at the stimulation site did not distinguish responders from non-responders, while overlap between the modeled electric field and the hallucination network did. That analysis is post hoc and pooled across arms, and it concerns hallucinations, not depression, so it cannot explain why the aiTBS trial was negative. It does suggest that “personalized targeting” covers methods different enough that one trial’s result says little about another’s.

The protocols and indications are not comparable either: twice-daily 1-Hz rTMS intended to reduce excitability in temporo-parietal and related cortex, for hallucinations, against accelerated intermittent theta-burst stimulation given ten times a day for five days, for depression. Both trials are small. A 51-patient, three-arm trial can miss a modest effect, and a 73-patient trial with a one-month primary endpoint can overstate one.

Where the hallucination evidence stands

Earlier sham-controlled trials of neuronavigated rTMS for hallucinations are split, according to the MULTIMODHAL authors’ review: two published trials, which located targets from language-task activation and from a motor-response symptom-capture method, found no advantage over sham, and a 2024 trial using resting-state guidance reported benefit. MULTIMODHAL adds what its authors describe as the first randomized trial comparing symptom-capture targeting directly against the standard placement, with a positive one-month result on both its revised and its registered endpoints.

What it does not yet have is peer review, a sham arm, an independent replication, or a durability estimate that survives its own attrition. The authors say their group has since validated an automated fMRI decoder meant to make target identification feasible outside specialist centers, and that larger multicenter replications are needed.

For now the evidence supports a narrow conclusion that neither trial alone would: aiming stimulation at the circuit active during the symptom improved on a crude scalp target for drug-resistant hallucinations, while individualizing a depression protocol by resting-state and EEG measures did not beat sham. Those are different claims, and the field has early evidence for one of them.