Back to Field Notes
SAFETY5 min readMelanie Nimmo

AI safety is more than crisis detection

On the clinical risk that only becomes visible when you read the whole conversation.


Crisis detection is the most obvious safety problem in mental health AI. If someone expresses an acute risk of harm during a conversation, the system needs to recognize it and get that conversation in front of a clinician. That is table stakes. The harder problem is what to do with the conversations that are clearly concerning when you read them from beginning to end, but never contain a single moment that looks like a crisis.

We kept seeing why those cases need their own safety layer. A person might mention that they have stopped going to work, then later describe barely eating, and later still talk about feeling detached from reality. Any one statement can be ambiguous. Together, they can form a clinically meaningful pattern. Real-time detection is designed to notice acute signals as they emerge. It is not the best tool for reconstructing the arc of a 20 or 30-minute conversation after the fact.

The difference between acute risk and clinical complexity

Thoughtful has a real-time crisis layer that runs during active sessions and looks for three broad categories of risk: risk to self, risk to others, and risk from others. When a signal is ambiguous, the system can ask a clarifying question before it escalates. When risk is confirmed, the transcript is routed for human clinical review. The model surfaces the signal; a clinician decides what happens next.

After the session ends, a separate high-acuity evaluation reads the completed transcript holistically. It evaluates the conversation across eight defined clinical domains: psychosis or distorted reality perception; severe compulsive or addictive behaviors; major functional impairment; non-suicidal self-harm; severe trauma symptoms; high-risk current problems; eating disorders with medical risk; and mandated reporter concerns.

Those categories are not eight versions of crisis. They cover presentations that may require human clinical involvement even when there is no immediate risk to life. Non-suicidal self-harm is a useful example. It belongs in high acuity, not automatically in the acute crisis pathway, because the clinical question is different. The system needs to recognize that the conversation may have moved beyond what AI should support on its own, without treating every such disclosure as an emergency.

Why a second pass matters

The distinction sounds simple once you say it out loud, but it changes how you build the system. A live detector has to work with partial information. It has to make sense of what has happened so far, without introducing noticeable latency, and it has to distinguish a real risk signal from a quotation, a third-person reference, a joke, or an ambiguous statement. That is a very specific job.

A post-session evaluator gets a different input: the whole record. It can ask a different question. Instead of 'is something dangerous happening right now?', it can ask 'does the overall clinical picture suggest needs that exceed the appropriate scope of AI-assisted support?' That is where gradual disclosure becomes visible. A symptom mentioned early in the session can be interpreted alongside a functional detail 15 minutes later and a pattern that only makes sense once the conversation is complete.

This is also why we do not treat high-acuity detection as keyword matching. The evaluation has to assess context and pattern. Severe functional impairment is not a phrase a person reliably types. Trauma severity does not announce itself with one canonical sentence. A system that only searches for expected language will miss the cases that are clinically coherent but linguistically messy — which is how real conversations tend to be.

Detection is the beginning of the workflow

The most important constraint in both layers is the same: detection is a trigger, not a clinical decision. A flagged high-acuity transcript goes to a human clinician for review. The clinician reads the conversation and decides whether any further action is warranted. The AI does not independently initiate clinical outreach, and a high-acuity flag should not be interpreted as an automatic phone call to the member.

That nuance matters. In a crisis workflow, the clinician may decide that a safety check, resources, a referral, safety-plan development, or emergency coordination is appropriate. In high acuity, the right next step may be different. Often the clinically sensible action is to orient the member toward an existing provider or appropriate resources rather than insert a new clinician into the relationship. Human review is consistent; the response is contextual.

The timing is different too. Detection and routing can happen in real time during an active session, but that does not mean a clinician instantly contacts the member. In the current crisis pathway, a clinician reviews the flagged transcript and determines whether outreach is warranted within the defined response window. High-acuity review happens after the session and follows its own clinical-review process. We keep those concepts separate because 'real-time detection' and 'real-time clinical response' are not the same claim.

Safety includes the boring failure modes

There is a less visible part of this architecture that I think matters just as much as the model behavior: what happens when the safety system itself fails. If a completed transcript cannot be processed by the high-acuity evaluator, it should not disappear silently. It needs to surface for manual review. If a primary alerting mechanism fails, there needs to be another escalation path. A safety system is only as strong as what happens on the bad day when one of its dependencies does not cooperate.

That is the broader principle behind the two layers. The goal is not to make one model smart enough to identify every clinically meaningful situation in every moment. It is to design multiple mechanisms around the ways risk actually appears, give each one a narrow job, and reserve consequential decisions for clinicians.

The best mental health AI will not be the system that can keep every conversation going. It will be the system that can recognize when continuing the conversation is no longer the most responsible thing to optimize for.

The Dispatch

Get the next field note
in your inbox.

A short note from the team whenever we publish. No sales sequences.

Unsubscribe anytime · No spam, ever

thoughtful· A Spring Health Incubation
© MMXXVI · All Research Rights Reserved