Ruling on objections, with reasoning: how Simulation keeps a record you can argue with
A simulated tribunal that rules without explaining is a coin toss with a gavel. We describe how Simulation records each ruling, how we checked 600 of them against practitioners, and where it still disagrees.
Across 1,840 simulated examinations run in our benchmark environment between January and May 2026, the tribunal ruled on 23,114 objections. Every one of those rulings is stored with its reasoning. This note is about why that matters, how we checked whether the reasoning holds up, and what the checks told us.
Why a ruling needs a record
When Trial & Arbitration Simulation became generally available in September 2025, the tribunal already ruled on objections. What it did not do well was explain itself. A ruling of “overruled” with no grounds is not useful to an advocate preparing for a real hearing, because she cannot tell whether the tribunal was applying the rule correctly, applying a venue-specific convention, or simply being lenient.
A ruling that comes with reasoning is something she can argue with. She can see that the tribunal treated her question as leading because of its form, decide that in the venue she is actually going to the chair would let it pass, and move on. Or she can see that the tribunal was right and fix the question.
So since the Simulation release in early 2026, each ruling record contains six fields: the objection type as raised, the grounds the objecting counsel cited, the ruling, a short paragraph of reasoning, the procedural rule or venue convention the tribunal relied on, and a confidence marker. The record is attached to the transcript at the line where the objection was raised and surfaces in the post-session scorecard under “weakest pleading” where relevant.
How we checked them
Explanations are only valuable if they are usually right. To test this, we drew a stratified sample of 600 rulings from the 23,114, weighted so that each of the 22 modelled venues contributed at least 20 and so that overruled and sustained rulings were roughly balanced.
We then asked six practitioners, three litigators and three arbitration counsel, each with more than ten years of hearing experience, to rate each ruling blind. They saw the transcript excerpt, the objection and the ruling with its reasoning, but not the venue label or the confidence marker. For each ruling they answered three questions: would you have ruled the same way; is the stated reasoning correct as a matter of procedure; and is the cited rule the right one for this venue, once the venue was revealed in a second pass.
Each ruling was rated by two practitioners. Where the two disagreed, a third resolved it. We paid for their time and did not tell them which rulings we were most confident about.
What the panel found
On the first question, whether they would have ruled the same way, the panel agreed with the tribunal on 87 per cent of rulings. Agreement was highest for form objections, leading and compound at 93 and 91 per cent respectively, and lowest for relevance at 74 per cent. That gap is unsurprising. Relevance is a judgement about where the case is going, and a simulated tribunal has a narrower view of that than a human chair who has read the pleadings three times.
On the second question, whether the reasoning was procedurally correct, the panel marked 91 per cent of explanations as correct and a further 5 per cent as correct but incomplete. The remaining 4 per cent contained an error. The most common error was citing the general rule where a venue-specific exception applied, which leads to the third question.
On venue, the panel agreed that the cited rule was the right one in 89 per cent of cases. The disagreements clustered in arbitral venues, where the tribunal tended to cite the institutional rules on evidence when experienced counsel would expect the chair to apply a looser, more pragmatic standard and cite nothing at all. We have adjusted the reasoning templates for arbitral venues so that the record now distinguishes between “the rules say” and “tribunals in this venue typically”.
The tribunal is a stricter reader of the rules than most chairs. That is useful for preparation and misleading if you forget it.
That comment, from one of the arbitration counsel on the panel, is the single most accurate description of the system we have heard.
Where the tribunal still disagrees with practitioners
Three patterns remain after the adjustments.
First, hearsay in arbitration. The tribunal sustains hearsay objections more often than the panel would, because the modelled venues apply the rules on written evidence faithfully, while real arbitral tribunals routinely admit hearsay and weigh it later. We now surface this in the reasoning but have not changed the ruling behaviour, because the stricter standard is more useful for preparation.
Second, speaking objections. Where objecting counsel makes a short speech instead of stating grounds, human chairs often rule on the underlying point. The tribunal tends to rule on the speech as framed. We think this is a defensible choice but acknowledge it is a modelling choice, not a legal one.
Third, cumulative evidence. The tribunal has no memory of how many times a point has been made across a multi-day hearing unless the session is run as one continuous simulation. Split sessions lose this context. We are testing a session-linking feature to address it.
Caveats
Six raters is a small panel. Agreement between the practitioners themselves, before adjudication, was 84 per cent on the first question, so a ceiling in the high eighties for tribunal agreement is roughly what one should expect. The sample is drawn from our benchmark environment, not from client sessions, which we do not retain. And the 22 modelled venues are weighted towards common-law trial courts and the major arbitral seats, so the figures say less about civil-law procedure than we would like.
A ruling you can argue with is more useful than a ruling you must accept. That was the design goal, and the record is now good enough that we are comfortable showing you where it fails.