Release 4 min read

Review 4.2: faster, quieter, and better at knowing when to say nothing

Review 4.2 cuts per-clause latency, suppresses flags a firm has repeatedly dismissed, and adds an explicit 'not assessed' state. Fewer flags is the intent. What changed, and what to watch.

2026 · 03 · 10·admin

Review 4.2 rolls out to all tenants this week. It is a smaller release than 4.0 and a more opinionated one. The headline is that Review will now flag less, on purpose, and will tell you when it has chosen not to assess a clause rather than guessing.

Summary

  • Per-clause latency on the hosted service: median down from 410 to 240 milliseconds, 95th percentile from 480 to 290. The budget we set in August 2025 was 500. We now have headroom.
  • Quiet mode. A Note that reviewers in a practice area have dismissed three or more times for the same clause pattern under the same playbook rule is suppressed and moved to a collapsed tray. This builds on the per-matter rejection learning introduced in 4.0 and extends it across matters within a practice area.
  • Abstention. When the model’s confidence falls below threshold, or a clause sits outside what the model is calibrated for, Review now returns “Not assessed” with a one-line reason instead of a low-confidence Note.
  • Governing-law detection from 4.1 now also checks the arbitration seat against governing law and forum, and flags three-way mismatches.
  • Comment anchoring inside Word tables is more reliable. Memory use in the add-in on long documents is down by about a third.

Why quieter

Our research team spent the last quarter of 2025 measuring what a dismissed flag costs. The direct cost is small: about forty seconds of a reviewer’s time. The indirect cost is larger. After a run of dismissed flags, reviewers accept the next few flags faster and edit them less, which is the behaviour you would expect from someone who has started skimming. A full write-up will follow.

The conclusion for the product was straightforward. A flag that will be dismissed is not neutral. It spends trust that the next flag, which might matter, will need.

Quiet mode acts on the clearest case: the same Note, under the same rule, on the same clause pattern, dismissed three times by reviewers in the same practice area. Suppression is per practice area and per rule, not per user. A partner’s dismissals suppress for the associates too, which is intended. Administrators can disable quiet mode per practice area or clear the suppression list.

High-severity flags are never suppressed. Medium flags are suppressed only after five dismissals and only if no reviewer has accepted the same flag in the same period.

Abstention

Until now Review assessed every clause it was given. When confidence was low, that usually produced a Note, and the Note was usually dismissed. 4.2 replaces the guess with a statement.

“Not assessed” carries one of six reasons: the clause exceeds the length the model is calibrated for; the text came from a scan with low OCR confidence; no playbook rule covers this clause type; the clause is in a language the playbook does not have rules for; the clause references a schedule or annex that was not supplied; or the model’s confidence was below threshold without a more specific cause.

In our benchmark the abstention rate is 3% of clauses on English-language documents, 11% on documents in other languages, and 19% on scanned documents. Those numbers are honest reflections of where the model is weaker, and we would rather show them than hide them inside Notes.

A “Not assessed” entry appears in the flag list and can be filtered. It is not a severity tier. It means a human should read the clause without help.

What to watch for

  • Fewer flags does not mean fewer issues. In the first week, open the suppressed tray and check what is in it. If something there should not be, clear it and the rule returns.
  • Suppression crosses matters. A clause pattern dismissed on three small NDAs will be suppressed on the next SPA in the same practice area. If a practice area handles very different document types, consider splitting it in the playbook.
  • The three-way governing-law check will produce new Medium flags on existing documents that 4.1 passed. Most are genuine. Arbitration clauses with a seat chosen for convenience and governing law chosen for substance are common and sometimes deliberate; the flag asks, it does not insist.
  • Faster can feel abrupt. Redlines now arrive before most reviewers have finished scrolling to the clause. Nothing has changed in what is proposed, only when.

Known limitations

  • Abstention reasons are templated. There are six. A clause can fail for two reasons at once and will show only the first.
  • Quiet mode does not yet learn from acceptances with edit. A flag that is always accepted but always rewritten is a different problem, and we are looking at it.
  • Table anchoring is improved for regular tables and still unreliable for merged cells.
  • Air-gap bundles receive 4.2 with the next bundle build, not automatically.

The intent of this release is that Review interrupts less and is more honest when it does. If it is interrupting less than it should, the suppressed tray and the abstention filter will show you where.

See it on a contract you have already reviewed.

Send us a draft your team has already redlined and we will show you what ZAAN catches, and what it misses.