Perspective 4 min read

Why we bet on a small model for contracts

Latency, residency, reproducibility and the sealed laptop: the four constraints that made a 7-billion-parameter model the right size for clause-level work, and what we gave up.

2025 · 10 · 30·admin

Why would a company whose product is judgement about contracts build a model with a fraction of the parameters of the largest available ones?

We have been asked this a good deal over the past fortnight, since ZAAN-7B Counsel took over Review’s inline pass. The question usually assumes that bigger is better and that we must have accepted worse answers for some other benefit. The second half is right. The first half is right less often than people expect, for this particular kind of work.

Four constraints that chose the size

Latency. Review’s inline flag has a budget of 500 milliseconds at the 95th percentile, from keystroke to margin. The inference share of that is around 200 milliseconds. A model that cannot answer a clause-level question in that time is not an inline model, whatever else it can do. We measured the largest models we had access to; none met the budget with the retrieval and citation steps included. A 7-billion-parameter model does, with room to spare.

Residency. Our tenants’ data stays in Switzerland, the EU, the US or Singapore, and we do not make cross-region calls. That means inference has to run inside each region. Running the largest models in four regions, at the availability a legal team expects, was possible in two of them and marginal in the others. A small model runs on a single accelerator anywhere. This is the constraint that made the decision for Singapore before any other argument was heard.

Reproducibility. Every flag Review raises cites a playbook position and records the model version that produced it. For that citation to mean anything a year later, the model has to be the same model. A pinned, signed, self-hosted model is. A model served by a third party, updated on their schedule, is not, however good it is on the day. We had two incidents in 2024 where a provider’s silent update changed severity assignments on a clause family overnight. Both were caught by shadow runs. Neither should have been possible.

The sealed laptop. A number of our accounts, and more of the ones we are talking to, need to run review on machines with no network connection at all. That is only realistic with a model that fits on the machine. We are testing air-gapped deployment now.

None of these is a capability argument. Together they are a shape argument: for clause-level work, delivered inline, in region, reproducibly, a small model is the right shape, and the only question is whether it can be made good enough.

Why it can be

Clause-level reasoning is narrow. The number of things a limitation-of-liability clause can do is large but finite. The vocabulary is stable. The structures repeat. A model that has seen hundreds of thousands of annotated instances of a clause family does not need general-purpose breadth to recognise the hundred-thousand-and-first; it needs depth in the family. That depth is cheaper in parameters than breadth is.

Our evaluation bore this out in the places we expected: defined-term consistency, cross-reference resolution and survival mapping all improved when we moved from a general model to a specialised one. These are tasks of pattern and bookkeeping across a long document, and having seen many documents matters more than reasoning hard about any one.

What we gave up

Breadth, mostly. A small specialised model is worse at summarising a long, discursive clause fluently, and worse on jurisdictions it saw little of in training. Our benchmark shows this plainly and we published the figures last week.

So we route. The inline pass uses the small model. The background pass, which has seconds rather than milliseconds, can call a larger model for the tasks where the small one is weaker, still in region, still with the same citation requirement, and the trace view says which model answered. We would prefer not to need this. For now we do, and the honest thing is to show it.

We also gave up a story. “Powered by the largest model available” is a sentence buyers understand. “Sized for the budget and the region” takes a paragraph. We are comfortable with that trade.

What we would say to a sceptical partner

Ask what the model needs to do, how fast, and where. If the answer is clause-level, inside half a second, inside your jurisdiction, with a citation you can reproduce next year, then the size of the model is set by those requirements, not by what is largest. A bigger model that answers late, from the wrong country, with a version you cannot pin, is not more capable for your purposes. It is less.

Small was not the ambition. It was what the constraints left standing once we took them seriously.

See it on a contract you have already reviewed.

Send us a draft your team has already redlined and we will show you what ZAAN catches, and what it misses.