Asking in English, answering from German, Mandarin and Arabic precedents
We built a 1,200-query benchmark to measure how much Recall loses when the question and the precedent are in different languages. German loses little. Mandarin loses some. Arabic loses more, and we think we know why.
Does a lawyer who asks Recall a question in English get the same answer she would have got if she had asked in the language of the precedent? That is the question this benchmark was built to answer, for three language pairs that account for most of the cross-language queries we see: English to German, English to Mandarin and English to Arabic.
The short version is that German is nearly transparent, Mandarin loses a measurable amount on certain clause types, and Arabic loses more than we are comfortable with. The longer version follows, with method first.
Setup
We assembled a test corpus of 9,400 clauses drawn from anonymised contracts contributed by firms under our research agreements, in roughly equal thirds of German, Simplified Mandarin and Modern Standard Arabic. Each clause was labelled by a bilingual lawyer with its type (limitation of liability, governing law, termination for convenience, change of control, and twelve others) and with up to three semantic tags describing its position: for instance, “cap at twelve months’ fees”, “mutual”, “carve-back for confidentiality”.
We then wrote 1,200 queries in English, 400 per target language, of the kind associates actually type: “termination for convenience with less than 60 days’ notice”, “change of control that excludes intra-group reorganisation”. Each query was matched by the labelling lawyers to its gold set of relevant clauses in the target-language corpus. Gold sets ranged from 3 to 41 clauses, median 11.
For comparison, each English query was translated into the target language by a second bilingual lawyer, independently of the labelling, to produce a monolingual query set. The difference between cross-language and monolingual performance on the same gold set is the cost of the language gap.
The metric is recall at ten: the share of gold clauses that appear in the top ten results. We also report mean reciprocal rank for the first correct result, because an associate stops reading when she finds one good precedent.
The system under test is Recall 2.7 as deployed in production, with no benchmark-specific tuning.
Results
Recall at ten, monolingual against cross-language, with the gap:
- German: 0.84 monolingual, 0.82 cross-language. Gap of two points.
- Mandarin: 0.81 monolingual, 0.74 cross-language. Gap of seven points.
- Arabic: 0.77 monolingual, 0.63 cross-language. Gap of fourteen points.
Mean reciprocal rank told the same story in a gentler register. For German the first correct result appeared, on average, at the same position regardless of query language. For Mandarin it slipped from position 1.4 to 1.9. For Arabic it slipped from 1.6 to 2.8, which in practice means the associate reads two or three results before finding a good one instead of one or two.
Breaking down by clause type, the Mandarin gap was concentrated in three types: limitation of liability, indemnity, and representations and warranties. These are the clause types where Mandarin legal drafting uses structures that do not map cleanly onto English terms, and where a single English word such as “indemnify” corresponds to several distinct Mandarin formulations depending on the document’s origin. Governing law, term and notice clauses showed almost no gap.
The Arabic gap was broader. Every clause type lost at least six points, and the losses were largest on clauses with numerical thresholds and on clauses from documents originating in different Arabic-speaking jurisdictions, whose drafting conventions differ more from one another than we had accounted for.
Why Arabic is harder
Three causes, in order of estimated contribution.
Morphology. Arabic roots generate a wide family of related forms, and legal drafting uses several of them for concepts that English expresses with one term. Our embedding model handles this well for common vocabulary and less well for the specialised legal register. The monolingual Arabic result being the lowest of the three suggests the corpus itself is harder, not just the crossing.
Document extraction. A disproportionate share of the Arabic corpus arrived as scanned or image-based PDFs, and right-to-left text in mixed-direction documents, where party names or amounts appear in Latin script, is extracted less reliably. We rebuilt right-to-left handling between Recall 2.6 and 2.7, and the improvement is visible against the earlier release, but extraction errors remain a source of missed matches. Where we re-ran the Arabic queries against only the clauses that were born-digital, the cross-language gap narrowed to nine points.
Term alignment. The bilingual term tables that Recall uses to anchor legal concepts across languages are deepest for German, which has had the longest production use, and shallowest for Arabic. We can see the effect directly: queries containing a term that has an entry in the alignment table lose less than queries that do not.
Caveats
The corpus is not a random sample of anything; it reflects which firms participated and which document types they contributed. The gold sets were produced by one lawyer per clause and not independently verified, so some of what we count as misses may be labelling disagreements. The queries were written by our team, not harvested from production, and may be cleaner than what associates type. Recall at ten is a convenient metric but says nothing about whether the clauses retrieved were the most useful ones in the set. And the benchmark measures retrieval, not the trace view’s reasoning over what was retrieved, which may compensate for or compound the gaps.
What we are testing
We are testing an expanded Arabic term alignment table built with practitioners in three jurisdictions, and a change to how numerical thresholds are normalised before embedding, which addresses the largest single source of Arabic misses. We are also testing whether showing the associate the translated query, so she can correct it before retrieval runs, closes part of the gap more cheaply than any model change would. Results from those tests will be reported when we have them.
For now, an associate querying Arabic precedents from English should read further down the list than she would for German, and should know that the system is less sure than it looks.