Ethical walls in Recall: enforcement at index time and query time
A wall enforced only at query time leaks through counts, snippets and ranking. Recall enforces twice, and this guide explains why, what each layer does and what administrators should check.
If a wall is enforced at query time, why enforce it again at index time? The question comes up in almost every security review we sit through. The short answer is that query-time filtering alone leaks, in ways that are subtle and, in a legal context, unacceptable. The longer answer is this guide.
What a wall has to prevent
An ethical wall separates the lawyers acting for one client from information about another, usually because the firm acts on both sides of something, or has done, or has taken on a lateral hire who did. The wall is a professional obligation, not a convenience. If it fails, the consequences range from an embarrassing conversation to disqualification from a matter.
“Fails”, for a search system, means more than returning a document the user should not see. It includes a result count that reveals a document exists. It includes a snippet from a permitted document that quotes a walled one. It includes a ranking that shifts because a walled document influenced the model’s sense of what is relevant. And it includes a trace view that shows a reasoning step passing through a document the user cannot open.
Recall 2.3, in January 2025, closed the first of these. The releases since have closed the rest, and the architecture that does it has two layers.
Index time: partitioning before anything is learnt
When a matter is indexed, its documents are embedded into a vector store that is partitioned by access group. The partition key is derived from the access control list the DMS attaches to the matter, which is the same list that governs who can open the documents directly. A matter visible to group A and group B is indexed into a partition readable by A and B. A matter walled to group C alone sits in a partition only C can reach.
Partitions are physically separate. Embeddings from one partition are never used to build or tune the index structure of another. Approximate nearest-neighbour indices learn their structure from the data they contain; if walled documents shaped the index an unwalled user queries, they would influence rankings even when filtered from the results. Partitioning at index time removes that path.
The same applies to the term and citation tables alongside the embeddings: a clause reference that appears only in walled documents never reaches the autocomplete of a user outside the wall.
Query time: filter before rank, not after
At query time, the user’s identity is resolved against the firm’s identity provider, and their group memberships determine which partitions the query may touch. The query is executed only against those partitions. There is no step where results are retrieved broadly and then filtered; the filter determines where retrieval happens.
This ordering is the difference between “we removed the walled results” and “the walled results were never candidates”. Filtering afterwards leaks: a count of 47 before filtering and 44 after tells the user three documents exist, and top-k retrieval that returns fewer than k results signals that something was removed.
The trace view respects the same boundary. Each reasoning hop is itself a retrieval scoped to the user’s partitions. A trace never passes through a document the user cannot open, and if the answer would have been better with a walled precedent, the user does not find out.
When walls change
Walls move. A lateral hire arrives, a conflict is identified mid-matter, a wall is lifted when a matter closes. Each changes which partitions a user may touch, and the two layers respond at different speeds.
The query-time layer responds as soon as the identity provider does: a user removed from a group loses that group’s partitions on their next query. This is sufficient for lifting a wall or restricting a user.
The index-time layer is slower when the wall itself moves: when a matter that was indexed into a shared partition must now be walled. The matter’s documents have to be removed from the shared partition and re-indexed into a new one. Until that completes, Recall applies a provisional query-time exclusion on the matter identifier so that no user outside the new wall can retrieve from it, and the admin console shows the matter as “wall pending re-index”. Re-indexing a typical matter takes minutes, a very large one up to an hour, and the provisional exclusion holds throughout.
Administrator checklist and known limitations
- Confirm that the DMS access control lists are the source of truth for matter visibility and that no matter relies on a convention the DMS does not encode.
- Run the partition audit from the admin console after initial indexing. It samples matters, resolves their access groups and confirms each sits in the expected partition.
- Test a wall from both sides. Log in as a user inside the wall and one outside, run the same query and compare results, counts and autocomplete suggestions.
- After a lateral hire or a new conflict, apply the wall in the DMS first, then verify the “wall pending re-index” state appears and clears.
- Review the quarterly wall report, which lists every matter whose partition changed and who authorised it.
Three limitations to know about. Walls are enforced at the matter level, not the document level. A single document that should be walled within an otherwise open matter must be moved to its own matter in the DMS before Recall can treat it separately. Group memberships are read from the identity provider on each query, but a cached session can persist for up to five minutes after a removal, so urgent removals should also be made in the DMS. And none of this protects against a user who is inside the wall and should not be; that is a question for the firm’s conflicts process, not for the index.
Two layers, each sufficient for the case it is built for, and together sufficient for the case where one of them is late.