'Human in the loop' has become standard language in healthcare AI, but the phrase can describe very different systems. A clinician might help write prompts, review model outputs, audit quality, or monitor a dashboard after the fact. All of that is useful. None of it necessarily means a clinician becomes involved when a person using the product actually needs human judgment.
That is the distinction we care about. When I hear a product described as human in the loop, my first question is not whether clinicians work at the company. It is where the human sits in the actual decision path. Do they review what the AI already did, or are there decisions the AI is not allowed to make without them?
There is a loop around the AI, and a loop around the person
Clinical oversight around the AI is essential. Clinicians help define the use case, establish safety criteria, review flagged conversations, and evaluate whether the product is operating inside its intended boundaries. That work shapes the system before a member ever sees it.
But there is another loop that matters just as much: the one that begins when the product detects that a member may need something different from what the AI can provide. In our crisis pathway, an LLM-based detector monitors active conversations for risk to self, risk to others, or risk from others. When a signal is confirmed, the transcript is routed to clinical review. A clinician reads the conversation in full and decides whether outreach is warranted and what form it should take.
That clinician has real options. Depending on the circumstances, they may conduct a phone safety check, provide resources or referrals, develop a safety plan, encourage a more appropriate level of care, or coordinate emergency services or a wellness check when imminent risk is identified. The AI can surface the situation. It does not choose among those interventions.
A flag is not an automatic intervention
Our high-acuity pathway makes this distinction even more important. After a session, a separate evaluator reviews the completed transcript for clinical presentations that may warrant human involvement even when there is no acute crisis. Every flagged transcript receives clinical review. But routine proactive outreach is not the default response to every high-acuity flag, and it would be misleading to describe the workflow that way.
Sometimes the right decision is outreach. Sometimes it is to provide resources. Sometimes it is to encourage the member to bring the issue to a provider they already have. The point of putting a clinician in the loop is precisely that the answer should depend on the context, rather than on a model converting a category into a fixed action.
This is also why we separate real-time detection from clinical response. A system can identify and route a risk signal while a conversation is active without claiming that a clinician is instantly contacting the user. The detection layer and the human response process have different time horizons, and collapsing them into one statement sounds impressive while making the workflow less accurate.
Human involvement has a privacy boundary
There is a second misconception we have had to design around: that putting humans into the safety pathway means a member's assigned therapist should automatically see their AI conversations. We do not assume that. The private, judgment-free quality of an AI conversation is part of what makes the modality useful for some people, and our research has been clear that members do not want an assigned provider reading those conversations without consent.
Clinical safety review and the provider relationship are deliberately connected without being the same thing. Safety alerts are managed through a triage system of rotating clinicians, while an assigned provider is notified that a safety flag occurred and given enough context to support continuity of care. What they do not automatically receive is the member's full AI conversation. The member retains control over how much more they want to share at their next appointment, creating a through-line for coordinated care without collapsing the privacy boundary around the AI interaction.
What the blended model is actually testing
This boundary is especially relevant as we begin studying a more blended model of care. The Blend pairs an AI-first experience with periodic check-ins from licensed providers who can guide direction, evaluate progress, and provide clinical support when needed. The premise is not that a provider reads everything the AI sees. It is that AI and human care can exist in the same ecosystem, with a more deliberate handoff between them.
We are still very early in that work. We do not have participant outcomes to report, and we do not yet have evidence that tells us how much of the work is ultimately best handled by AI versus providers. Those are questions the study is designed to examine. What we can describe today is the design intent: if a person needs human care, the product should have somewhere meaningful to route them rather than treating every departure from AI as a failure.
That changes the incentive structure of the product. An AI-only experience is naturally optimized to keep as much activity as possible inside the AI. A care ecosystem can optimize for a different question: what is the right level of care for this person now? Sometimes that will be continued AI support. Sometimes it will not.
A more useful standard
For buyers evaluating mental health AI, 'human in the loop' is not a sufficient answer. Ask who the human is, what causes the system to involve them, what information they see, what decisions they control, how quickly the process is designed to move, and what privacy boundaries still apply. Most importantly, ask what the AI is structurally prevented from deciding on its own.
The human should not be there only to improve the AI. In the moments that matter most, the human needs to be there to apply judgment to the person's situation.