Perspective 4 min read

Why every flag needs a citation

A flag without a reason cannot be checked, cannot be disagreed with and cannot be defended later. The audit argument for citations is stronger than the usability argument, and it is the one that matters.

2025 · 04 · 29·admin

Suppose a review tool tells you that clause 14.2 is a problem, and nothing else. What do you do?

You read clause 14.2. You decide for yourself whether it is a problem. The tool has saved you nothing except, perhaps, the decision of where to look first. If you disagree with it, you have no way of knowing whether the tool knows something you do not or has simply misread the clause. If you agree with it and the clause later turns out to have been fine, you cannot explain to anyone why you changed it.

This is what we mean when we say a flag without a citation is not a flag. It is a hunch with a user interface.

The usability argument, briefly

The usability case is well known and we will not labour it. Reviewers trust flags they can check. In our own measurements, when we tested a view that hid the triggering rule and matched text behind an extra click, accept rates fell and the time spent per flag rose, because reviewers were re-deriving the reason from scratch. Showing the reason is faster. That is sufficient on its own to justify doing it.

But it is not the strongest argument, and if it were the only one we would be vulnerable to a tool that was simply so accurate that people stopped checking.

The audit argument

A law firm’s work product is subject to scrutiny after the fact in a way that most professional output is not. A client asks why a position was conceded. A regulator asks how a document was reviewed. An insurer asks, after a claim, what the firm’s process was. A court asks whether advice was reasonable. In each case, the question is not only “was the decision right” but “can you show how it was made”.

A review process that produces decisions without reasons cannot answer that question. It does not matter how good the decisions were. If the firm’s answer is “the tool said so”, the firm has delegated judgement it was not entitled to delegate, and it will be judged accordingly.

A citation changes this. When a flag says “limitation of liability cap is below the playbook minimum of 12 months’ fees (rule LL-03, last reviewed by the commercial practice head in January 2025); matched text: ‘the Supplier’s aggregate liability shall not exceed the fees paid in the six months preceding the claim'”, the firm can answer every part of the question. Here is the firm’s position. Here is who set it and when. Here is the text that engaged it. Here is what the reviewer decided to do about it and why. The tool’s role in that chain is visible and bounded: it applied a rule the firm wrote to text the firm can read.

A citation is the difference between a tool that assists a lawyer’s judgement and a tool that replaces it without saying so.

What a citation has to contain

Not all citations are equal. We hold ours to four elements.

The rule. The specific playbook entry that triggered the flag, by identifier, with its current text. Not “liability concerns” but the entry itself, so that the reviewer can see whether the rule is right as well as whether the match is.

The provenance of the rule. Who last edited it and when. A rule last touched three years ago by someone who has left is still a rule, but the reviewer should know that before relying on it.

The matched text. The exact words in the document that engaged the rule, anchored to their position. Paraphrase is not enough; the reviewer needs to see what the model saw.

The reasoning step, where there is one. Some flags are direct: the rule says X, the text says not-X. Others involve inference: the cap is expressed in a currency amount, the fees are elsewhere in the document, and the model has computed that the cap is roughly five months’ fees. Where there is a computation or a cross-reference, the flag shows it, so that an error in the inference is as checkable as an error in the rule.

A flag with all four can be audited by someone who was not there. A flag with fewer cannot, and the gap will be found at the worst time.

The objection, and why we think it is wrong

The objection is that this is slow, that it clutters the interface, and that a sufficiently accurate model should be trusted the way a sufficiently experienced associate is trusted, without a footnote on every sentence.

The comparison to the associate is instructive, and it cuts the other way. An experienced associate is trusted because she can be asked. Her reasoning is available on demand, she can be cross-examined on it and she bears professional responsibility for it. A model that cannot explain itself has none of those properties. Trust without the ability to interrogate is not trust in the professional sense; it is reliance, and reliance on an opaque process is what the audit question is designed to expose.

As for clutter: the reason can be visually quiet. It does not need to shout. It needs to be there.

Where this leaves us

Every flag in Review carries its rule, the rule’s provenance, the matched text and any inference. Every clause Draft composes carries a citation to the precedent or playbook entry it drew on. Every Recall result shows the permission path that allowed it to appear. We built it this way because the people using these tools will one day be asked to show their working, and the tools should make that easy rather than impossible.

If a tool cannot tell you why, it has not told you anything.

See it on a contract you have already reviewed.

Send us a draft your team has already redlined and we will show you what ZAAN catches, and what it misses.