ZAAN-7B Counsel: a small model trained end-to-end for clause-level reasoning
Our in-house model: what it was trained on, what it was not, how it was evaluated against the models we previously routed to, and where it still falls short.
Seven billion parameters is small by the standards of general-purpose language models. It is roughly the size at which a model can run on a single accelerator inside each of our four regions, answer a clause-level question inside Review’s 500-millisecond inline budget, and be loaded onto a sealed laptop. Those three constraints, more than any ambition about capability, set the size of ZAAN-7B Counsel, which has been serving Review’s inline pass for all accounts since the first week of October and is being introduced into Draft and Recall over the coming weeks.
This note is about how it was built and tested. A companion piece from Engineering covers why we chose this path at all.
What it was trained on
Three sources, in descending order of volume.
Public and licensed legal text. Legislation, published judgments and awards, regulatory guidance and model forms from jurisdictions covering our 38 languages, together with precedent libraries and commentary we licensed from publishers under terms that permit model training. This is the bulk of pre-training.
Authored playbooks and annotated clauses. Our legal team, with contract lawyers engaged for the purpose, wrote around 2,100 playbook positions across 40 clause families and annotated 180,000 clause instances drawn from the public and licensed corpora with clause type, position relative to a playbook, severity, and a redline. This is the supervised layer that teaches the model what a flag is.
Synthetic negotiation traces. Sequences of drafts, counter-drafts and acceptances generated from the annotated clauses, then reviewed by lawyers for plausibility. About 60% of generated traces survived review. These teach fallback behaviour: what the next position after a rejection looks like.
What it was not trained on: customer data. No document, clause, query, redline or rejection from any ZAAN tenant was used in pre-training, fine-tuning or evaluation. This is the no-train guarantee, and the training data manifest is available to any account under NDA for verification. Evaluation on customer material happened only in the sense that firms ran the candidate model on their own matters, in their own tenants, and told us what they thought.
How it was evaluated
We maintain an internal benchmark, Clause Bench, of fourteen tasks that correspond to things Review, Draft and Recall actually do: clause classification, playbook position matching, severity assignment, redline generation, defined-term consistency, cross-reference resolution, survival mapping, governing-law extraction, limitation-period computation, bilingual parity checking, precedent retrieval ranking, summarisation of long clauses, unusual-jurisdiction handling, and a refusal task (declining to flag when the playbook has no position). Each task has between 400 and 3,000 held-out items, hand-labelled by at least two lawyers with disagreements adjudicated.
ZAAN-7B Counsel was compared against the general-purpose models we were routing to before October, under the same prompts, retrieval and citation constraints. We do not name them; the comparison is to our own previous production path, not to any vendor’s best configuration.
Results. On eleven of fourteen tasks the small model was within two points of the previous production path. On three it was better by a meaningful margin: defined-term consistency (+9 points), cross-reference resolution (+7) and survival mapping (+6). These are tasks where having seen a great many contracts matters more than general reasoning. On two it was worse: long-clause summarisation (−5) and unusual-jurisdiction handling (−8), the latter mostly on jurisdictions thinly represented in the training corpus.
Latency at p95 on the inline path fell from 240 ms to 170 ms, which is the figure that made the switch possible rather than merely interesting.
How it is deployed
One model version per region, pinned and signed. Every flag records the model version that produced it, so a trace from March can be reproduced in September. Updates roll by region with a two-week overlap during which both versions run on a shadow basis and disagreements are reviewed by hand before the new version takes the inline path.
For the two tasks where the small model is weaker, Review routes to a larger model in the background pass. The routing decision is logged and visible in the trace view, and the citation requirement is identical whichever model answered.
Caveats
- Clause Bench is ours. It was built by the same team that built the model, and however carefully held out, it reflects our view of what matters. Firms evaluating the model should run their own documents.
- The 38-language coverage is uneven. Performance on the twelve languages with the largest training representation is close to English. On the long tail, it is measurably lower, and the per-language confidence shown in Draft and Recall reflects this.
- Shadow-run disagreement review is human and slow. It catches regressions on the clause types reviewers encounter most. Rare clause types get less scrutiny.
- Pre-training data cut-off is mid-2025. Legislative change after that date is handled by retrieval, not by the model’s own knowledge.
The model is small because the job required it to be. Whether it is good enough is a question each firm should answer on its own matters, with the trace view open.