Perspective 3 min read

2025, told quietly: what shipped, what didn’t, what we learned

Ten version numbers, three general availabilities, one model. A short account of the year, with the slips left in.

2025 · 12 · 30·admin

Ten version numbers, three general availabilities, one model. That is the year in a sentence. The rest of this note is the same year told in a little more detail, with the slips left in.

What shipped

Review went from 3.6 in January to 4.1 in November. The January release replaced the single risk score with three severity tiers, High, Medium and Note, that a firm can tune per practice area. April rebuilt the Word add-in on the native comments API, which removed most of the “my redlines vanished” tickets. July’s 4.0 was the release that mattered most to us: Review began learning from rejected redlines inside a matter, so that a partner who strikes the same suggestion twice does not see it a third time. November added governing-law and jurisdiction mismatch detection, after we counted how often an English-law agreement carried a New York arbitration clause nobody had intended.

Draft moved from 3.4 to 3.8. Voice modelling in February, trained on a firm’s last two hundred closed deals, is still the feature partners mention first. The deal-point form in June let boilerplate write itself from twenty answers. December’s 3.8 delivers a single English draft in any of 38 languages, with the parity checks we described in April.

Recall went from 2.3 to 2.6: ethical walls enforced at query time, then the trace view in June, then citations that survive translation in September.

AI Interview reached general availability in May. Trial & Arbitration Simulation followed in September, with 22 venues modelled. In October we announced ZAAN-7B Counsel, the small in-house model that now does most clause-level reasoning across all three surfaces.

Open Bar, our free tier for legal-aid clinics and public defenders, spent the autumn in Lagos. We wrote about that earlier this month and will not repeat it here, except to say it changed our roadmap more than any customer meeting did.

What didn’t

Recall 2.7 was meant for December. It will ship in late January. The new index is faster than we hoped and the migration tooling was slower than we hoped; we chose not to ask firms to re-index over the holidays.

Two-way sync with document management systems did not ship. One-way import works. The write-back path is in a pilot with four firms, and we are not yet satisfied with how it handles a document that two people edit in the same hour.

Slack and Teams triage, which several firms asked for, is still internal. We are testing it. It is not ready.

Air-gap mode, which lets a sealed laptop run Review and Draft with no network at all, spent the year as an engineering project and is now running on a handful of machines at six firms. More on that in January.

What we learned

Rejections teach more than acceptances. The single most useful data stream we have is the redline a partner declines, and the one they decline with an edit. Review 4.0 was built on that, and the research team spent much of the second half of the year building a taxonomy of why redlines get declined. Early finding: the most common reason is not “wrong” but “right, but not worth the conversation with the other side”.

Latency is a product decision, not an infrastructure one. We set a 500-millisecond budget for per-clause review in August and held to it. Reviewers who waited more than a second started reading ahead of the model and then ignored it.

Small models were the right bet for us. ZAAN-7B Counsel is cheaper to run inside a tenant’s region, fits on a laptop, and on our clause-level benchmark sits within a point or two of models many times its size. It is worse at open-ended drafting prose, which is why Draft still uses a larger model for the first pass and the small one for checking.

Pro-bono software has different constraints. In Lagos the limiting factor was never model quality. It was the bandwidth of a shared connection, the number of paralegals per lawyer, and whether the thing worked on a phone.

Numbers, with the usual caveats

Across our active accounts: 340 firms and legal teams in 28 countries; 2.4 million contracts reviewed since the first version of Review; 38 languages supported in Draft and Interview. These are our own counts. They include trial accounts that ran for a month and left. They do not include Open Bar, which we count separately and which, at 4,200 paralegals in June, is now larger by headcount than the paying side.

We did not ship everything we said we would. We shipped most of it, and the things that slipped slipped for reasons we can explain. That will do for a year.

See it on a contract you have already reviewed.

Send us a draft your team has already redlined and we will show you what ZAAN catches, and what it misses.