Recall 2.7 ships: ten years of precedent, in three seconds
Recall 2.7 rebuilds the index from the clause up. Queries across a decade of matters return in under three seconds, with the trace view showing each hop. A re-index is required.
Recall 2.7 is available today for all tenants. It was due in December. It is late because the migration tooling was not ready and we did not want firms re-indexing over the holidays. This note covers what changed, why, what to watch for, and what still does not work well.
Summary
- New clause-level index. Median query time across a ten-year matter history of around 400,000 documents is 2.8 seconds, down from 11 seconds on 2.6. The 95th percentile is 6.1 seconds.
- Time-aware ranking. Ask how the firm handled MAC carve-outs for pandemics before 2020 and the results respect the date, not just the words.
- Trace view shows each hop. Where 2.5 showed which documents answered a question, 2.7 shows the path: query, interpreted terms, clause matches, precedent attached, and where the ethical wall was checked.
- Ethical walls are now enforced at index time as well as at query time. Behaviour for users is unchanged; the second check is defence in depth.
- A full re-index is required. See below.
What changed and why
The 2.6 index was built on documents. A matter’s SPA was one entry, with clause boundaries inferred at query time. That worked for small histories and degraded steadily above 100,000 documents, a line several firms crossed last year.
2.7 indexes clauses. Each document is segmented at ingest; each segment carries its heading path, its language, its matter metadata and its date; and the ranking model scores segments rather than files. A document becomes a container for its clauses, which is closer to how a lawyer thinks about it anyway.
The practical effect is that “indemnity cap in software MSAs with German counterparties” returns clauses, each with the surrounding document a click away, rather than forty documents that each mention the words somewhere.
Time-aware ranking exists because precedent ages. A warranty package from 2016 is not wrong, but it should not outrank one from last March unless you asked for it. 2.7 reads date intent in the question and otherwise applies a gentle decay. The decay can be switched off per query.
The trace view was introduced in 2.5 to show which sources answered a question. Firms used it more than we expected, mostly to check that nothing from the wrong side of an ethical wall had leaked in. 2.7 makes that check explicit. Each trace now lists the wall evaluations that ran, the user’s clearance at the time, and the segments excluded as a result, shown as counts rather than content. If a result set is small, the trace tells you why. The trace is exportable as plain text; several firms attach it to the matter file.
Migration
Every tenant must re-index. The job runs in the background and the old index serves queries until the new one is complete, so there is no downtime, but there is a window in which results reflect the old behaviour. Re-indexing takes roughly four hours per 100,000 documents in our testing; large histories should expect a day or two. Administrators can watch progress in the console and pause it.
During migration, the trace view shows which index answered a query. Do not compare result quality across the two in that window; they differ by design.
Scanned documents are re-run through OCR as part of the re-index. If a firm’s OCR quality was poor, 2.7 will surface that as clauses with low confidence rather than hiding them. Expect a few surprises.
What to watch for
- Ranking has changed. Users who had learned the quirks of 2.6 will find that familiar queries return different top results. Most are better. Some will need rephrasing.
- Saved queries are preserved but re-run against the new index. Check them.
- Clause segmentation is imperfect on documents without headings. A long letter agreement may be split oddly. The segments are still searchable; they are just less tidy.
- Query logs now include interpreted terms. If a firm exports logs to its own systems, the schema has an additional field.
- The Word add-in, Draft’s precedent lookup and Review’s citation links all use the new index automatically. Nothing is needed on the firm’s side beyond the re-index.
Known limitations
- Cross-language ranking is uneven. English and the major European languages rank well against each other. Mandarin and Arabic segments rank well within their language and less well across it. We are measuring this and will publish what we find.
- Time-aware ranking relies on document dates from metadata. Undated documents are treated as old, which is usually right and occasionally not.
- The index slice used by air-gap mode is still built in the 2.6 format. Sealed laptops will see 2.6 behaviour until the next bundle.
- Very long queries, above roughly 300 words, are truncated. Recall is a search tool. If the question is a memo, write the memo.
If a query that used to work no longer does, the trace will usually show why. If it does not, send us the trace.