Research 4 min read

94% after one month: what happens to accept rate as a playbook matures

Across 112 accounts onboarded since Review 4.0, median accept rate on proposed redlines rose from 71% in week one to 94% in week four. Here is how we count, what drives the curve and what it does not tell you.

2026 · 08 · 18·admin

Ninety-four per cent. That is the median share of Review’s proposed redlines that reviewers accepted without modification in the fourth week after onboarding, across the 112 accounts that joined between July 2025 and March 2026. In the first week the figure was 71 per cent. This note is about the twenty-three points in between: how we measure them, what moves them and why the headline number should be read with some care.

How we count

A proposed redline can end in one of four ways inside a matter. It can be accepted as drafted. It can be accepted after the reviewer edits it. It can be rejected, with or without a reason. Or it can be left untouched when the document is closed, which we treat as a soft rejection.

Accept rate, as we report it, is the first category only, divided by all four. Accepted-after-edit is reported separately because it means something different: the issue was real but the fix was not quite right. Conflating the two flatters the system.

The denominator is every redline proposed on a document that the reviewer actually opened and closed. Documents opened and abandoned are excluded. Flags at Note severity, which carry no redline, are excluded. The unit of analysis is the account, not the redline, so a large firm generating thousands of redlines does not dominate the median.

The 112 accounts were selected because they onboarded after Review 4.0 shipped in July 2025. That release introduced in-matter learning from rejections, which is the mechanism this note is really about. Accounts onboarded earlier show a flatter curve and are excluded to keep the comparison clean.

The curve

The median by week, with the interquartile range in brackets:

  • Week 1: 71 per cent (62 to 78)
  • Week 2: 82 per cent (75 to 87)
  • Week 3: 89 per cent (84 to 92)
  • Week 4: 94 per cent (90 to 96)
  • Week 8: 95 per cent (92 to 97)
  • Week 12: 95 per cent (93 to 97)

Two things stand out. The curve is steep for three weeks and then flat. And the spread narrows as it rises: the interquartile range is sixteen points in week one and six by week four. Accounts converge. Whatever is happening in those three weeks happens to nearly everyone.

Accepted-after-edit follows a mirror pattern, falling from a median of 18 per cent in week one to 4 per cent by week four. Hard rejections fall from 9 per cent to 2 per cent. Soft rejections stay roughly constant at around 1 per cent, which suggests they are a property of reviewer behaviour rather than system quality.

What drives the rise

We can attribute the rise to three mechanisms, in descending order of effect, based on accounts where we were able to isolate each.

The first is playbook correction. A new account’s playbook has been written by lawyers who know their positions but have not yet seen how a model reads them. The first week exposes ambiguities: a fallback that was meant to apply only above a deal-size threshold but was encoded without one, a red line that was really a preference. Firms fix these in the playbook editor, usually in a burst during days three to ten. In accounts where we could measure it, playbook edits accounted for roughly half the improvement.

The second is in-matter learning. When a reviewer rejects a redline and gives a reason, Review 4.0 and later adjust the proposal within that matter and, with the firm’s permission, within that practice group. This is not retraining; it is a weighting layer over the playbook. It accounts for perhaps a third of the improvement, concentrated in weeks two and three.

The third is reviewer calibration, the polite term for reviewers learning what the system is good at and ceasing to argue with it where it is consistently right. We can see this in the rejection reasons: “disagree with playbook” rejections fall sharply even on redlines that have not changed. This is real improvement in the team’s experience, but it is not improvement in the system, and it is the component we are least comfortable counting.

What accept rate does not tell you

A high accept rate is consistent with at least two bad outcomes.

The first is that the system has learnt to propose only the redlines it knows will be accepted and has gone quiet on the hard ones. We watch for this by tracking flag volume per thousand clauses alongside accept rate. Across the 112 accounts, flag volume fell by 11 per cent over the first month, which is consistent with playbook tightening and not with the system suppressing difficult flags. But the check is necessary, and Review 4.2’s “know when to say nothing” changes made it more necessary, not less.

The second is that reviewers have stopped reading. An accept rate of 99 per cent in a team that closes documents in four minutes is a warning sign, not a success. We surface time-on-document alongside accept rate in the admin view for this reason, and we have flagged this pattern to three firms.

Caveats

This is an internal observation across our own accounts, not a controlled study. The accounts that disengaged in the first month, seven of the original 119, are excluded, which flatters the curve by an amount we cannot measure precisely. Accept rate is measured per account, and accounts differ in document mix; an account reviewing mostly NDAs will converge faster than one reviewing bespoke SPAs. And week four is not the end: the plateau at 95 per cent is a plateau in acceptance, not in quality, and the two are related but not the same.

The number worth watching is not the 94 but the three weeks it takes to get there. That is the time a firm should budget for a playbook to become a playbook.

See it on a contract you have already reviewed.

Send us a draft your team has already redlined and we will show you what ZAAN catches, and what it misses.