The National Alliance on Mental Illness named John Torous, director of digital psychiatry at Beth Israel Deaconess Medical Center and associate professor at Harvard Medical School, as special advisor to NAMI on artificial intelligence and mental health. The announcement reads, on its face, like a routine advisory appointment. It represents something more specific: the formalization of a partnership that has already produced a live, operational evaluation platform for AI mental health tools, in a market category where no equivalent regulatory certification currently exists, and where the safety stakes are no longer theoretical.

What already exists, before this appointment

NAMI and Torous’s Division of Digital Psychiatry announced their benchmarking initiative in December, building on MindBench.ai, an open platform described in a peer-reviewed paper published that November in NPP, Digital Psychiatry and Neuroscience, that profiles and benchmarks large language models and LLM-based mental health tools across three specific categories: safety and crisis response, accuracy of information, and cultural relevance. This is not a proposed framework awaiting development. It is a functioning evaluation platform, built on Torous’s earlier work creating mindapps.org, the largest existing database of mental health apps, and the American Psychiatric Association’s own app evaluation model, giving this project a credibility base that predates NAMI’s involvement by years.

Why the timing makes this more than a research curiosity

Torous provided expert testimony to the House Energy and Commerce Subcommittee’s November 2025 hearing on AI chatbots, one of the first instances of direct congressional attention to this specific category. That hearing did not produce legislation, and none appears imminent. In the absence of a regulatory framework, litigation has started filling the gap: lawsuits have alleged that AI chatbots encouraged suicide or reinforced delusional thinking in vulnerable users. And the scale of use is already substantial rather than speculative: a NAMI and Ipsos survey found 12 percent of adults say they are likely to use an AI chatbot for mental health support within six months. A category this large, this unregulated, and already generating litigation is exactly the condition under which a credible, non-governmental evaluation standard tends to matter most, filling a gap regulators have not yet addressed.

Why NAMI specifically is a meaningful actor here, not simply a visible one

NAMI’s own public materials are notably unambiguous about what this partnership is not: its FAQ states plainly that NAMI does not endorse AI for mental health treatment, for any age group. That is a real distinction from a company-sponsored safety seal or an industry self-regulatory body, since NAMI has no commercial stake in any AI product succeeding and is on record declining to endorse the category at all. Combined with Torous’s specific technical credibility and his direct experience testifying before Congress, this positions the partnership less as an advocacy statement and more as a considered attempt at independent, patient-centered evaluation infrastructure, built by people with no obvious incentive to grade generously.

What this could mean for the broader digital mental health market

This desk has tracked similar dynamics before in adjacent categories, professional accreditation and evaluation frameworks that end up functioning as de facto market standards well before, or entirely instead of, formal government regulation catches up. A credible, independently developed benchmark for AI mental health tools carries real weight for exactly the audiences this desk covers: investors assessing which AI mental health products carry real safety credibility versus which are making unverified claims, health systems and insurers deciding which tools to recommend or cover, and companies building in this space who now have a named standard to design toward rather than operating in a complete evaluation vacuum. Whether MindBench.ai becomes the standard, or one of several competing frameworks, its existence changes the conversation from whether independent evaluation is possible to which evaluation framework companies and consumers should trust.

The caveats

MindBench.ai is a new platform, described in a single published paper from late 2025, and this desk has not independently reviewed its specific methodology or tested its results against a range of commercial AI mental health products. A benchmark built by one research group, however credible, is not the same as a binding regulatory standard, and nothing about this partnership creates legal consequences for a company whose AI tool scores poorly. NAMI’s own caution, that it does not endorse AI for mental health treatment at all, should be read as a real limit on how far this initiative goes: it is building an evaluation tool, not certifying any product as safe or effective.

The frame

An advisor appointment is rarely worth close attention on its own. This one is, because it marks the point at which an evaluation framework for a large, unregulated, and legitimately risky category of consumer mental health technology moved from an academic research project into a formalized, ongoing institutional partnership with the country’s largest patient advocacy organization behind it. Congress has held one hearing and passed no law. Courts are handling individual lawsuits one at a time. In that gap, NAMI and Torous are building something concrete: a live, published, named framework for telling users, clinicians, and companies which AI mental health tools are behaving safely, and which are not. That is a different kind of story than most AI-and-mental-health coverage this year, less about what a chatbot claims to do, and more about who is building the infrastructure to check.