The empathy layer: what we learned building Interview with trauma-informed clinicians
What should a voice agent do when the caller starts crying? We spent six months with three clinicians and 140 mock intakes finding out. The answers were mostly about pacing and consent.
What should a voice agent do when the caller starts crying?
It is not a hypothetical. AI Interview handles intake for legal-aid clinics through Open Bar, and a meaningful share of those calls concern eviction, domestic abuse, asylum and dismissal. The first version of Interview, before general availability, handled distress the way a well-meaning form does: it waited briefly, said something kind, and asked the next question. The clinicians we later worked with had a word for this. The word was “interrogation”.
This note describes the work that followed, the method we used to test it, what changed, and what we still do not know.
Who we worked with, and what the layer is
Three clinicians, over roughly six months spanning Interview’s general availability in May 2025: one who works with survivors of domestic abuse, one who works with asylum applicants, and one in a hospital trauma service. Two intake supervisors from legal-aid clinics joined for the later sessions. We do not name them, at their request and ours.
The clinicians named the result “the empathy layer” and we kept the name despite being uneasy with it. It is not a personality and it is not warmth bolted onto the voice. It is a set of constraints that sit between Interview’s dialogue planner and its speech output, governing pacing, sequencing, consent and vocabulary. The planner decides what information is still needed. The layer decides whether now is the moment to ask for it, and how.
Method
We borrowed the standardised-patient method from medical education. Trained actors were given detailed scenario briefs and asked to play callers consistently across many sessions, including the emotional arc of the call. Six scenarios: an eviction notice with a short response deadline, unpaid wages after dismissal, an application for a protective order, an asylum intake, a consumer debt dispute, and a workplace injury. Three languages: English, French and Spanish.
140 mock intakes in total, 70 with the pre-layer version and 70 with the post-layer version, balanced across scenarios, languages and actors. Each call was scored on three things: the share of facts from the scenario key that reached the intake brief; whether the call was completed or abandoned; and the actor’s own rating of distress, on a five-point scale, at three points in the call. Separately, one clinician reviewed forty transcripts without being told which version produced them and recorded a preference.
What changed
Most of the changes are small and unglamorous.
Silence. In segments the planner marks as sensitive, Interview waits 2.4 seconds of silence before speaking, up from 0.8. Callers who are collecting themselves are no longer talked over.
Consent before the hard part. Interview now asks before moving into a sensitive topic: whether it is all right to ask about what happened, now, or whether the caller would rather come back to it. Callers who defer are asked again later, once, and if they defer again the brief notes the gap for the lawyer.
No “why”. Questions beginning with “why” were removed from sensitive segments. “Why didn’t you leave” became “what happened next”. The clinicians were unanimous on this and the transcripts bear them out.
The caller’s words. If the caller says “assault”, Interview says “assault”. It does not reclassify events into neutral terms the caller did not choose.
No reassurance about outcomes. Interview does not say that things will be fine or that the caller has a strong case. It says what it is doing and what happens next.
An exit. At any point the caller can ask for a person, and the layer makes that offer itself if distress markers persist. The brief so far is handed over rather than discarded.
Limitation-period maths still runs. In a distressed call the deadline is reported to the lawyer in the brief, not read out to the caller.
Results
Facts captured rose from 71% of the scenario key to 83%. Completion rose from 64% to 86%. Actor-rated distress at the end of the call fell from a mean of 3.4 to 2.3. The blind clinician review preferred the post-layer transcript in 36 of 40 pairs.
The cost is time. The median post-layer call is four minutes longer. Clinics have told us the trade is acceptable; a longer call that reaches the end gathers more than a shorter one that does not.
Caveats
Actors are not survivors. Standardised patients are the best available proxy and remain a proxy. A distress rating given by an actor is partly a rating of their own performance.
Three languages out of 38. The layer’s phrasing has been reviewed by native-speaking clinicians in six languages; in the remaining 32 it is translated and unreviewed. We would rather say that than imply otherwise.
Three clinicians and two supervisors. They were generous and expert and they are five people.
We measured the call. We did not measure what happened to the caller afterwards, and the thing that matters most, whether a person in difficulty got help sooner, is outside what this study can see.
Clinics in the Open Bar network now run the post-layer version by default. What they tell us is anecdotal, and it is consistent with the numbers, which is the most we can claim.