Research 4 min read

How we evaluate redline quality: accept rate is not enough

Accept rate rewards timid redlines. Here is the four-part measure we use instead, the sample behind it, and the places where it still misleads us.

2025 · 07 · 24·admin

Eighty-eight per cent. That is the share of Review’s proposed redlines that reviewers accepted across our active accounts in the first half of 2025. It is a number we could put on a slide. We try not to, because on its own it tells you less than it appears to.

The problem with accept rate

Accept rate measures whether a reviewer clicked yes. It does not measure whether the redline was any good.

Consider two systems. System A proposes small, safe changes: fix a cross-reference, tighten a definition, add “reasonable” before “endeavours”. Reviewers accept nearly all of them. System B proposes the change the playbook actually calls for: replace a six-month cap with twelve, add a carve-out for data breach. Reviewers accept perhaps three in four, because the fourth needs a conversation with the client first.

System A has the higher accept rate. System B is the one you want in the room.

Accept rate, optimised in isolation, pulls a model towards System A. We noticed this in late 2024 when an internal fine-tuning run improved accept rate by four points and reviewers at two pilot firms told us, independently, that Review had become “polite”. That was the end of accept rate as a target.

What we measure instead

Since January 2025 we have scored redlines on four axes. None is sufficient alone.

1. Accept rate, split. We still record it, but we separate accepted as proposed from accepted then edited. An edited acceptance is a partial failure: the reviewer agreed with the direction and disliked the wording. Across the first-half sample, 71% of acceptances were as proposed and 17% were edited. The edited share is the number we watch.

2. Edit distance on edited acceptances. When the reviewer changes the wording, how much? We compute a token-level edit distance between what Review proposed and what the reviewer kept. Small distances (a word or two) suggest a style problem. Large distances suggest Review got the direction right but the drafting wrong, which is a more serious fault in a tool meant to draft.

3. Survival to signature. Did the accepted redline, or something recognisably descended from it, appear in the executed agreement? This is the hardest to collect and the most informative. A redline that is accepted internally and then traded away in the first counterparty turn may still have been right, but a pattern of non-survival on a clause type tells us the proposals are starting from an unrealistic position.

4. Blind panel score. A rotating panel of nine practising lawyers, drawn from three firms that have agreed to take part, scores a monthly sample of redlines on a four-point scale without seeing whether the reviewer accepted them. They see the clause, the playbook position and the proposal. Nothing else.

The sample

The figures above come from 11,400 redlines proposed between January and June 2025 across 38 accounts that consented to evaluation use. Consent here means that anonymised clause-level data, stripped of party names, amounts and identifiers at the tenant boundary, is visible to our research team for this purpose. The remaining accounts are not in the sample, consistent with our zero-retention defaults.

Survival-to-signature data exists for 3,100 of those redlines, where the firm also connected the executed version through the DMS integration. The panel scored 720 redlines over six monthly rounds. Inter-rater agreement (Krippendorff’s alpha) was 0.68, which is acceptable for a subjective four-point scale and lower than we would like.

Where this still misleads us

Consent bias. The 38 consenting accounts are not a random draw. They skew towards mid-sized European firms with established knowledge-management functions. We do not know how Review performs on redlines at firms that opted out, and we should not pretend the sample speaks for them.

Survival is confounded. A redline can be excellent and still not survive, because the client chose to concede it. We cannot see the client’s reasoning. We treat survival as a trend indicator by clause type, never as a verdict on an individual proposal.

Panels are slow and they drift. Nine lawyers scoring 120 redlines a month cannot keep pace with a model that proposes thousands a day. The panel is a calibration tool for the other three metrics, not a replacement. We also observed the panel becoming stricter over the six rounds; the mean score fell by 0.2 with no corresponding fall in the other measures. Familiarity, perhaps.

Edit distance punishes house style. A firm that habitually rewrites “shall” to “will” generates edit distance that has nothing to do with Review’s judgement. We now exclude a short list of per-firm style substitutions before computing it. The list is maintained by hand and is certainly incomplete.

None of this measures what Review failed to propose. A redline it should have raised and did not generates no accept, no edit, no survival and no panel score. We address recall separately, through seeded documents with known issues, and will write that up in its own note.

What changed because of it

Two things. The fine-tuning objective since February has been a weighted combination of as-proposed acceptance and panel score, with survival used as a monthly check rather than a target. And the fallback ladder in Review 4.0, which offers the next position when a reviewer declines a proposal as too aggressive, came directly from watching the edit-distance distribution: a cluster of large edits that were, on inspection, reviewers manually writing the fallback themselves.

If a vendor gives you one number for redline quality, ask what it would look like if the system simply proposed less.

See it on a contract you have already reviewed.

Send us a draft your team has already redlined and we will show you what ZAAN catches, and what it misses.