The cost of a false positive in contract review
A dismissed flag costs a reviewer forty-one seconds. That is the small number. The larger one is what happens to the next three flags.
Forty-one seconds.
That is the median time a reviewer spends on a flag they then dismiss, measured from opening it to closing it, across 212,000 flag interactions in the fourth quarter of 2025. On its own the number is not alarming. Dismissed flags were 11% of all flags in the sample, so the direct cost works out at roughly 5% of total review time. If that were the whole story, false positives would be a nuisance and nothing more.
It is not the whole story. This note describes what we measured, what we found about the flags that follow a dismissal, and the limits of the method.
What we measured
Thirty-eight firms opted into interaction telemetry in the Word add-in for the quarter. Telemetry records events, not content: a flag was shown, a flag was opened, a decision was made (accept, accept with edit, dismiss), the time of each, and the severity tier and playbook rule identifier of the flag. No clause text, no document identifiers, no user identity beyond a per-session token. The data stays in the firm’s region and we received aggregates.
The sample covers 212,000 flag interactions across roughly 9,400 documents. The firms are a mix of sizes and practice areas, with the usual skew towards Swiss, German, Nordic and UK corporate work. They are also, by definition, the firms engaged enough to opt in.
Two costs
The direct cost is the forty-one seconds. For comparison, a flag accepted without edit took a median of 23 seconds; a flag accepted with edit took 68. Dismissal sits in between, which makes sense: the reviewer reads the clause, reads the suggestion, decides it does not apply, and moves on.
The indirect cost shows up in what happens next. We looked at the three flags following each decision and compared their handling with the baseline.
After a single dismissal, the next flag was accepted without edit 14% faster than baseline, and the share of accept-with-edit decisions fell by 9%. After two consecutive dismissals, the median time to accept the next flag without edit dropped to 15 seconds, and stayed below baseline for the following three flags. After three consecutive dismissals, the pattern was stronger again, and in a minority of sessions the reviewer began accepting flags in under eight seconds, which on a clause of any length is not reading.
We cannot see attention. We can see latency, and latency behaves exactly as it would if reviewers were deciding, after a few wasted interruptions, that the tool was not worth their full attention for a while. The effect faded after four or five flags, and reset entirely after a flag that was accepted with edit, which we read as the reviewer re-engaging because something was worth engaging with.
The effect held across all thirty-eight firms, across practice areas, and across severity tiers of the following flags. A dismissed Note made reviewers skim the next High.
Which flags get dismissed
Notes accounted for 71% of dismissals, Medium for 26%, High for 3%. By playbook rule type, the leaders were style and formatting rules, rules phrased as “consider adding”, and defined-term consistency rules of low consequence. By clause, dismissals clustered in boilerplate: notices, counterparts, entire-agreement and further-assurance clauses.
In other words, the flags most likely to be dismissed are the flags least likely to matter, and they spend the reviewer’s attention on behalf of the flags that do.
A dismissal is not always a false positive. Some are “right, but not worth raising with the other side”, a category we have written about before. From the model’s point of view the distinction is real; from the reviewer’s attention budget it is not.
What changed as a result
Review 4.2, released this month, suppresses Notes that reviewers in a practice area have dismissed three or more times under the same rule on the same clause pattern. It also abstains, with a stated reason, where it previously produced a low-confidence Note. Both decisions came from this data.
We have also changed the guidance we give firms writing playbooks. Fewer “consider” rules. Fewer style rules at Note level. If a rule exists because someone once wanted it, and it is dismissed nine times in ten, it is costing more than it returns.
Caveats
Opt-in firms are the engaged ones; a less engaged firm might skim sooner, or might never have engaged enough to skim.
Latency is a proxy for attention. A fast acceptance may be a confident one. We think the consistency of the pattern across firms makes the skimming interpretation the likelier one, but we cannot prove it from timestamps.
We cannot see outcomes. The question that matters is whether skimmed flags led to missed issues in signed documents, and nothing in this data can answer it. We would need the document after signature and we do not have it.
The quarter includes December, which is not a typical month in corporate practice.
The practical conclusion is simple enough to state in one line: a flag that will be dismissed is not free, and the price is paid by the flag after it.