A 500-millisecond budget: why review latency is a product decision
Inline review either keeps pace with drafting or it interrupts it. We set a hard budget of half a second at the 95th percentile, and most of our architecture follows from it.
An associate is in Word. She has just pasted a counterparty’s indemnity clause into the draft and is reading it. If a severity flag appears in the margin while she is still reading, the flag is part of the reading. If it appears after she has moved on to the next clause, it is an interruption, and she will go back, or she will not.
We measured where the line falls with 40 reviewers across six firms in late 2024, by delaying flags artificially and watching what happened. Below about 600 milliseconds, reviewers behaved as though the flag had always been there. Between one and two seconds, they looked up, and about a third went back. Beyond two seconds, most had moved on and a quarter never returned to the flag at all. The study was small and the reviewers knew they were being observed, so we treat the thresholds as indicative. But they were consistent enough to act on.
We set the budget at 500 milliseconds, measured from the keystroke or paste that finishes a clause to the flag rendering in the Word margin, at the 95th percentile, in every region. Not the median. The 95th.
What half a second buys
A budget is only useful if it is spent deliberately. Ours, as of Review 4.0, is allocated roughly like this:
- Network, round trip, client to regional endpoint: 60–90 ms. We do not route across regions. A Singapore tenant talks to Singapore. That is a data-residency commitment first, but it is also the only way to keep this number under control.
- Clause segmentation and playbook retrieval: 70–110 ms. Finding the clause boundary, classifying the clause type, pulling the relevant playbook entries and the matter’s settled state.
- Model inference: 180–240 ms. The first-pass model reads the clause against the retrieved positions and produces the flag, severity and a candidate redline.
- Citation assembly and rendering: 50–80 ms. Attaching the playbook citation, writing the comment through the native Word comments API, painting the margin.
Add it up and there is almost no slack. A slow DMS lookup, a cold cache, or a 90-page draft that forces re-segmentation will blow the budget. That is why the budget is a product decision rather than an engineering target: it forces choices about what Review is allowed to do inline.
What we gave up
The inline pass uses a small model. It has to. A larger model would produce marginally better redlines at two to four times the inference time, and the reviewer would have moved on. So the inline pass is handled by a model sized for the budget, and a second, slower pass runs in the background over the whole document and surfaces anything the first pass missed as a batched update when the reviewer next pauses. We call this two-pass review internally. It is less elegant than one good pass and considerably more useful.
Flags sometimes appear before redlines. If the candidate redline is not ready inside the budget, the flag ships without it and the redline follows. Reviewers told us a flag with a redline arriving a second later is fine. A flag arriving a second late is not.
Trace view is lazy. The full reasoning trace, which shows which playbook entries were considered and why one was chosen, is assembled on request, not inline. Nobody reads a trace at 500 ms.
Some checks are not inline at all. Cross-document consistency, defined-term audits across a bundle, and governing-law checks against schedules are document-level and run in the background pass only. A margin flag that depends on page 60 cannot appear on page 3 in half a second, and we stopped pretending it could.
How we hold ourselves to it
Every region reports p50, p95 and p99 inline latency per tenant. The p95 is the number on the wall. In the first half of 2025 it sat between 410 and 470 ms in CH, EU and US, and between 440 and 510 ms in SG, where it crossed the budget on nine days, mostly during a storage migration in April. We reported those nine days to the affected tenants.
Any change that adds more than 20 ms to the inline path at p95 needs sign-off from the product owner, not just the engineering lead. That rule has killed features. It killed, for instance, an inline comparison of the counterparty’s standard terms against the draft, which was useful and cost 140 ms. It now runs in the background pass.
Why this is not just an engineering story
Latency determines what kind of product Review is. At 500 ms it is a second reader at the reviewer’s elbow. At two seconds it is a checklist the reviewer runs afterwards. Those are different products with different adoption curves, and the second one loses to the reviewer’s own habits most of the time.
It also shapes the model strategy. A tool that must answer inside half a second, in four regions, without cross-border calls, cannot depend on the largest available model somewhere else. That constraint has pushed us towards smaller, specialised models that run close to the tenant, and we will say more about that in the autumn.
The practical version: if you are evaluating any inline review tool, ask for p95 latency by region, not for the demo on the vendor’s laptop.