Research 4 min read

The parity guarantee: how we check that a Mandarin and an English draft say the same thing

A bilingual contract that differs between its two language versions is a dispute waiting for a tribunal. Here is how we test for parity, what the test catches, and what it still misses.

2025 · 04 · 08·admin

Does the Mandarin version of this clause say what the English version says?

For a bilingual contract, that question is not academic. Where both versions are expressed to be equally authentic, a divergence between them is an ambiguity that a tribunal will have to resolve. Where one version prevails, the other side may have agreed to something they did not read. Either way, the time to find the divergence is before signature.

We call the standard we hold Draft to the parity guarantee: a clause generated in two languages must have the same legal effect in both. This note describes how we test that, on what sample, and where the test is weaker than we would like.

The method

Parity is not translation accuracy. A translation can be fluent and faithful to the words while changing the effect, and it can depart from the words considerably while preserving it. So we do not score translations. We score effects.

The test has three layers.

Layer one: structural alignment. Each clause in version A is paired with its counterpart in version B. Obligations, conditions, exceptions, defined terms and cross-references are extracted from each and compared as structures. A condition present in one and absent in the other is a structural failure. A defined term used in one and undefined in the other is a structural failure. This layer is mechanical and catches around 60 percent of the divergences we eventually find.

Layer two: back-rendering. Version B is rendered back into the language of version A by a separate model that has not seen version A. A third process compares the back-rendering to the original at the level of legal propositions: who must do what, by when, subject to what, with what consequence. Differences are listed. This layer is where most of the subtle divergences surface.

Layer three: bilingual lawyer review. A sample of clause pairs, stratified by clause type and by whether layers one and two flagged anything, goes to a panel of reviewers qualified in both the relevant languages and at least one of the relevant jurisdictions. Each reviewer answers one question per pair: would these two clauses produce the same outcome before a competent tribunal? Yes, no or uncertain. Reviewers do not see each other’s answers.

The sample

The current benchmark covers English paired with Mandarin, German, French, Spanish, Portuguese, Japanese and Arabic. The English–Mandarin pair is the deepest, with 2,400 clause pairs across 180 documents, drawn from anonymised templates contributed by participating firms and from our own synthetic contracts. The other pairs range from 900 to 1,600 clause pairs.

Clause types are weighted towards those where divergence is most consequential: payment, termination, limitation of liability, indemnity, governing law and dispute resolution, conditions precedent. Boilerplate is included but under-weighted.

The reviewer panel for English–Mandarin is six people. Inter-reviewer agreement on the yes/no question is 91 percent; the disagreements are concentrated in clauses where the two legal systems have no clean equivalent, which we discuss below.

What we found

Across the English–Mandarin benchmark, Draft’s current output passes all three layers on 96.8 percent of clause pairs. The failures are instructive.

Just under half are modal strength. English legal drafting distinguishes “shall”, “will”, “may” and “must” in ways that are conventional rather than strictly logical, and Mandarin renders obligation through different constructions. A “shall” that functions as a future tense in English can be rendered as a firm obligation, and a “may” with a contextual restriction can lose the restriction. We have tuned for this and the rate is falling, but it remains the largest category.

Around a quarter are scope of exceptions. “Except as otherwise provided in this Agreement” and similar phrases attach to different things depending on placement, and placement conventions differ between the languages. The effect is that an exception narrows or widens.

The remainder are a mixture: numerals and dates (ordinal and inclusive-counting conventions differ), defined-term drift where a term is defined once and then paraphrased, and a small number of cases where the back-rendering layer produced a false alarm that the panel cleared.

Where the test is weaker than we would like

Concepts without equivalents. Some legal concepts do not translate because the receiving legal system does not have them. Consideration, in the common-law sense, has no clean counterpart in a civil-law system. Certain PRC regulatory concepts have no English equivalent. In these cases, the panel’s answer is frequently “uncertain”, and we report these clause pairs separately rather than counting them as passes or failures. Across English–Mandarin, they are about 2 percent of the sample.

Synthetic contracts. Around a third of the benchmark is synthetic. Synthetic contracts are cleaner than real ones. Real bilingual contracts contain inconsistencies that pre-date any model, and Draft’s parity on those is harder to assess because the baseline is not itself at parity.

Jurisdiction, not just language. A Mandarin draft for a Singapore-governed contract and one for a PRC-governed contract should differ. Our benchmark currently treats the language pair as the unit and controls for jurisdiction only partially. We are restructuring it so that the unit is language pair plus governing law.

Drift after editing. The guarantee covers what Draft generates. Once a lawyer edits one version by hand, parity is no longer assured unless the other version is regenerated or Review is run across both. We are testing a bilingual Review mode that flags post-edit divergence, and have no date for it.

What the guarantee means in practice

When Draft produces a bilingual clause, it has passed layers one and two before you see it. Layer three is a benchmark, not a per-clause check; no human has reviewed your specific clause unless you have. The 96.8 percent is a rate across a sample, not a promise about a document.

If one version prevails under the contract, read the prevailing version. If both are authentic, read both, and read the exceptions first.

See it on a contract you have already reviewed.

Send us a draft your team has already redlined and we will show you what ZAAN catches, and what it misses.