How we evaluate clause-level reasoning without leaking the contract
Zero retention means we cannot read the contracts our model reviews. So how do we know it is improving? A method built on synthetic drafts, in-tenant scoring and a panel of lawyers who disagree.
How do you measure a model on documents you are not allowed to see?
The honest answer for most of 2024 was: badly. We had a benchmark of synthetic contracts and we had accept rates from the Word add-in, and the gap between the two was where the real question lived. This note describes the harness we built to narrow that gap, the sample it runs on, and the things it still cannot tell us.
The constraint
Zero retention is not a policy; it is an architecture. Contract text enters a tenant’s region, is reviewed in memory by ZAAN-7B Counsel, and is gone. The firm’s flags and redlines persist only inside that tenant, encrypted with the tenant’s key. Nobody at ZAAN can read a customer’s contract, and we do not want a research exception to that, because an exception is a hole.
So any evaluation has to satisfy two conditions. It must say something about real contracts, because synthetic ones are too clean. And it must move no contract text, and nothing from which contract text could be reconstructed, out of the tenant.
Three layers
The harness has three layers, each more realistic than the last and each giving us less access.
The first is the open benchmark: 1,840 synthetic contracts across eleven types (SPA, MSA, DPA, NDA, ISDA schedules, licence agreements, employment contracts, leases, loan agreements, shareholder agreements and supply contracts), each written by a drafter on our legal team against a specification, then degraded on purpose. Degradation means inserting the issues a reviewer should catch: a liability cap that excludes nothing, a survival clause that forgets confidentiality, a governing-law clause pointing one way and a jurisdiction clause the other. Every inserted issue is labelled by severity and by the playbook rule it offends. We can read these contracts, run anything we like, and publish the numbers.
The second is the panel set: 212 real contracts contributed under written consent by nine firms, with all parties, amounts and dates replaced by a legal secretary before we received them. These are messier than the synthetic set and closer to what Review actually sees, but they are a small sample, skewed to the practice areas of the firms that volunteered.
The third is in-tenant scoring. A firm that opts in runs the harness inside its own tenant, against its own matters, with its own playbook. The harness scores the model’s flags against the firm’s reviewers’ final decisions and sends back a scoreboard: counts by severity, by clause type, by language, by outcome. Nothing else. No text, no clause, no identifiers. Twenty-three firms participate. Their scoreboards are the closest thing we have to ground truth.
What “correct” means
Lawyers disagree. Before we could score anything we needed to know by how much.
Five lawyers independently labelled 300 contracts from the open benchmark. Pairwise agreement on whether a flag should exist at all was high. Agreement on severity was lower: a Cohen’s kappa of 0.81 on High versus not-High, and 0.64 on Medium versus Note. That second number matters. A model that places a flag at Medium where two of five lawyers would say Note is not wrong; it is within the range of the panel.
So the harness scores severity with tolerance. A flag counts as correct if its severity matches the panel majority or sits one tier away in a direction at least one panellist chose. Exact-match severity is reported separately and is always lower.
What we found, briefly
On the open benchmark the current model catches 96% of inserted High issues and 89% of Medium. It raises a flag that no panellist would have raised on 4% of clauses. On the panel set those numbers are 93%, 84% and 7%. In-tenant, across the twenty-three participating firms, the median scoreboard shows 91% of High flags accepted or accepted with edit, and a dismissal rate (flags closed without edit) of 9%, with a wide spread between firms.
The gap between synthetic and real is what we expected and is the reason the harness has three layers. The spread between firms is more interesting and is mostly explained by playbook maturity, which we will write about separately.
Caveats
Synthetic contracts are cleaner than real ones. We insert issues deliberately; real issues arrive by accident, in worse prose, in documents carrying tracked changes from three parties.
The panel set is small and skewed. Nine firms, mostly Swiss, German and Nordic, mostly corporate and commercial. It says little about litigation documents or English-law employment contracts.
In-tenant scoring measures agreement with the firm’s reviewers, not truth. If a firm’s reviewers consistently miss a class of issue, the scoreboard rewards the model for missing it too. We can see this only when one firm’s scoreboard looks unlike the others, and even then we can only ask.
Non-English coverage is thin. The open benchmark is 70% English; the panel set is 80% English. The in-tenant scoreboards break out by language, which is how we know the model is weaker in Arabic, but the counts are small.
The method tells us where the model is improving and where it is not. It does not tell us what any one contract said, and that is the point.