Notes for autumn 2026: what we are testing, carefully
Two betas open, five things in testing, three things we are deliberately not doing. A plain account of where the work is this autumn, with no promises attached.
We do not publish a roadmap. We have explained elsewhere why we prefer to ship quietly and describe things after they exist. But once a year, around now, enough firms ask the same question in the same fortnight that it is simpler to answer it once. The question is: what are you working on. This is the answer, as of the first week of October, with the usual caveat that “testing” means testing and not “shipping in the spring”.
What is in beta
Two things, both described in their own release notes last month.
Triage for Slack and Teams, available to tenants on Review 4.2 since 17 September. Early use is concentrated in a small number of dedicated channels, as we recommended, and the early finding is that the card is used less for routing documents to the right lawyer, which is what we designed it for, and more for telling the salesperson who dropped the file that the answer is not coming in the next five minutes. We will take that.
Salesforce deal context in Draft templates, available since 24 September. The most common administrator question so far is about the legal-entity field, which confirms our suspicion that most CRMs do not have one, or have one nobody maintains. We are considering whether Draft should offer to look the entity up from the counterparty’s jurisdiction’s company register, and we have not decided.
Both betas are expected to run into early 2027. Neither will move to general availability until the firms using them tell us the limitations we listed have stopped mattering.
What we are testing
Arabic retrieval. Our cross-language benchmark showed a fourteen-point gap between English queries over Arabic precedents and the same queries written in Arabic. We are testing an expanded term alignment table, built with practitioners in three jurisdictions, and a normalisation step for numerical thresholds. We are also testing whether simply showing the associate the translated query before it runs closes more of the gap than either model change. Early signs favour the cheap option, which is usually how these things go.
Voice-aware redlines in Review. Nine per cent of reasoned rejections are a reviewer rewriting an acceptable redline in the partner’s preferred voice. Draft has modelled a firm’s voice since 3.4; Review has not, because a redline is a surgical change and voice seemed secondary. The rejection data says otherwise. We are testing redlines that draw on the same voice model Draft uses, scoped per practice group, and we are watching for the failure mode where a redline sounds right and says something slightly different.
Session linking in Simulation. The tribunal currently has no memory across split sessions, which means an objection on cumulative evidence on day three cannot refer to what was admitted on day one. We are testing linked sessions that share a record. The engineering is not difficult; the question we are still working through is how the scorecard should treat a multi-day performance.
Field-level confidence in Interview briefs. A computed deadline is only as good as the inputs behind it. We are testing a marker on each field in the intake brief that shows whether the value came from a direct answer, an inference or a default, so that the reviewing lawyer knows which dates to check first. Clinics using Open Bar are the first to see this, because their briefs are read by the fewest people with the least time.
A narrow Salesforce write-back. One activity record, “draft generated”, with a link. Nothing else. We said as much in the beta notes and the position has not changed.
What we are not doing
We think it is as useful to say what is not being built, so here are three things firms have asked for that we have decided against for now.
Autonomous review. A mode in which Review applies its redlines and returns a finished document without a lawyer reading the flags. We are asked for this a few times a quarter. The answer is no, for the reasons we gave when describing why we sometimes decline to onboard a firm. Review proposes. Someone with a practising certificate disposes.
A general-purpose chat interface over the matter file. Recall answers questions about precedents and shows its trace. It does not summarise the matter, draft the client email or speculate about strategy, and we are not building a surface that does. The failure modes of an open-ended assistant inside a privileged file are not ones we can audit, and audit is the whole argument.
More venues in Simulation, for now. Twenty-two are modelled. We have requests for a dozen more. Each new venue requires published awards or judgments in sufficient volume to model a tribunal’s behaviour with any confidence, and for most of the requested venues that volume does not exist. We would rather model twenty-two venues honestly than forty thinly. We will add venues when the material supports it, and we will say which material.
How to read these notes
Everything under “testing” might not ship. Some of it will ship in a form different from the one described here. Nothing here has a date, and if someone from our side gives you one, they are guessing.
What we can say is that the shape of the work has not changed. Review, Draft and Recall get a little more careful each release. Interview and Simulation get a little better at knowing what they do not know. The security architecture does not move. The free tier stays free.
If you are a firm deciding whether to wait for something on this list before starting, do not. The things that are shipped are the things that work.