Research 4 min read

How we measure first-pass review time without flattering ourselves

Time saved is the easiest number to inflate. Here is how we define first-pass review time, who we measure, what we exclude and where the measurement still misleads.

2025 · 02 · 18·admin

Every vendor in this field quotes a time saving. Most of the numbers are not wrong, exactly. They are measured in a way that makes them look better than they are. This note describes how we measure first-pass review time across our accounts, so that when we quote a figure you can decide for yourself how much of it to believe.

The definition

First-pass review time is the elapsed working time between a reviewer opening a draft for the first time and the reviewer marking the draft as ready for internal sign-off or return to the counterparty, whichever comes first.

Each part of that sentence is doing work.

Elapsed working time, not wall-clock time. A draft opened on Friday afternoon and finished on Monday morning is not a sixty-hour review. We count only intervals in which the document was in the foreground in Word or in the Review pane, and we cap any single interval at twenty minutes of inactivity before pausing the clock.

Opening for the first time. Re-opens after sign-off, to deal with the counterparty’s response, are a second pass and are measured separately.

Ready for sign-off or return. We do not measure to execution. Execution depends on the other side.

Who we measure

The baseline is the harder half of the problem. To say that something got faster, you need to know how fast it was before.

We use two baselines and report both.

Within-firm, before and after. For each firm, we take the twelve weeks before Review was switched on and compare them with weeks five to sixteen after. Weeks one to four are excluded because they are learning weeks and they flatter nobody. This baseline requires the firm to have been tracking time on review tasks before onboarding, which about two-thirds of our accounts were, through their time-recording system. For the other third, we have no before and we do not pretend to.

Within-firm, concurrent. Some firms roll out by practice group. For a period, the commercial team is using Review and the real-estate team is not. Where the two groups review comparable documents (NDAs are the usual case), we compare them over the same weeks. This controls for seasonality and for anything else that happened to the firm that quarter.

As of this month, the before-and-after sample covers 61 firms and roughly 38,000 first-pass reviews. The concurrent sample is smaller: 14 firms, about 6,000 reviews.

What we exclude

  • Documents under four pages. Short NDAs and side letters have review times dominated by opening and saving the file. Including them improves the percentage saving and tells you nothing.
  • Documents over 200 pages. Too few to analyse separately and highly variable.
  • Reviews where the reviewer accepted every proposed redline without edit. We treat these as a signal that the reviewer did not review, and they are reported as a separate rate (it is low, under 3 percent, but not zero).
  • The first four weeks after onboarding, as above.
  • Any firm that changed its playbook substantially mid-window, since that changes the task.

What we found, with the caveats attached

Across the before-and-after sample, the median first-pass review time fell by 41 percent. The interquartile range of that saving across firms is 29 to 54 percent. In the concurrent sample the median saving is 36 percent, with a range of 24 to 47.

The concurrent number is lower. We think it is also closer to the truth, because the before-and-after design cannot fully separate the effect of Review from the effect of a firm having just decided to pay attention to review efficiency. Firms that onboard are firms that have been thinking about the problem.

Savings are larger for mid-length documents (15 to 60 pages) and for document types where the firm’s playbook is mature. They are smallest for bespoke agreements with little precedent. On a first-of-its-kind joint-venture agreement, the saving is close to zero, and we would be suspicious of anyone who told you otherwise.

Where the measurement still misleads

Foreground time is a proxy. A reviewer who reads a flag, switches to email to ask a colleague and comes back has paused our clock while still, in a real sense, reviewing. We undercount. Conversely, a reviewer who leaves the document open while taking a call is overcounted until the twenty-minute cap. We do not know the net direction.

Quality is not in this number. A faster review that misses more is not a better review. We measure quality separately, through partner acceptance of redlines and through a sampled blind re-review, and we report it separately. The two should always be read together.

Survivorship. Firms that found Review unhelpful in the first quarter tend to reduce usage, and a reviewer who reverts to reading the whole document unaided while the pane is open still counts as a Review session. This biases the saving upward by an amount we cannot measure precisely. We have tried to bound it by looking at sessions with no interaction with the pane; they are around 7 percent of the sample and show no meaningful saving, which is what you would expect.

Self-selection of document type. Reviewers may route easier documents through Review and keep harder ones for themselves. We see some evidence of this in the first two months and less afterwards.

For all of that, there are things we will not do. We will not quote the 41 percent without the 36. We will not quote either without the range. We will not describe the saving as “up to 54 percent”, which is a true statement about one firm’s upper quartile and a misleading one about everyone else.

When you see a time-saving figure from us, it will come with a definition, a sample size and a sentence about what it does not capture. If it does not, ask.

See it on a contract you have already reviewed.

Send us a draft your team has already redlined and we will show you what ZAAN catches, and what it misses.