Research 4 min read

Modelling a tribunal from published awards: what we can and cannot infer about a chair

Simulation models venues, not named arbitrators. This note sets out what a corpus of published awards actually supports, where selection bias defeats it, and why we stopped short.

2025 · 09 · 18·admin

The question we are asked most often about Simulation, usually in the first ten minutes of a demonstration, is some version of: “Can it tell me how this particular arbitrator will rule?”

The short answer is no, and we have built it so that it cannot. The longer answer is about what a corpus of published awards is actually evidence of, and it is worth setting out because the same reasoning applies to anyone claiming otherwise.

The corpus

Simulation’s venue models were built from published arbitral awards and court judgments available in public and licensed repositories, together with institutional rules, procedural orders where published, and practitioner commentary that we licensed. For the 22 venues in general availability, the corpus runs to roughly 14,000 awards and judgments, unevenly distributed: the English Commercial Court and the major European arbitral seats are well represented; some of the Asian seats rest on a few hundred documents.

Each document was decomposed into procedural events (applications, objections, rulings), the treatment of categories of evidence (documentary, witness, expert), the structure of the reasoning, the cost allocation and its stated basis, and the outcome by head of claim. Decomposition was done by a model and checked by hand on a stratified sample of 600; agreement was 91% on procedural events and 83% on the reasoning-structure categories, which are more interpretive.

What the corpus supports

Procedural conventions. How a venue handles document production, whether witness statements stand as evidence-in-chief, how strictly hearsay-type objections are treated, whether the tribunal intervenes in cross-examination. These are stable across many awards and across chairs, and they are what Simulation’s venue models mostly encode.

Reasoning style. Some venues produce awards that reason from the contract text outward; others from the commercial purpose inward. Some address every pleaded point; others dispose of the case on the narrowest ground available. This is visible in published awards and reasonably stable within a venue.

Treatment of expert evidence. Venues differ measurably in how often expert evidence is preferred over documentary evidence on quantum, and in whether tribunals appoint their own experts. This shapes how Simulation’s tribunal weighs a quantum case.

Cost allocation tendencies. Whether costs follow the event, how often they are apportioned by issue, and how often conduct adjusts them. Published awards state the basis for cost orders more reliably than they state the basis for anything else.

All four are venue-level properties. A team that knows them can prepare better for a hearing in that venue regardless of who sits.

What the corpus does not support

Outcome propensity for a named individual. There are three reasons, each sufficient on its own.

First, selection. Published awards are a small, non-random fraction of all awards. In most commercial arbitration, publication requires the parties’ consent, and parties consent more readily to awards that embarrass nobody. The published record for any individual arbitrator is a sample chosen by the losing parties’ willingness to be seen losing. Nothing estimated from it generalises.

Second, attribution. A three-member tribunal produces one award. The reasoning may be the chair’s, a co-arbitrator’s, or a negotiated compromise, and dissents are rare. Treating the award as evidence of the chair’s own views is an assumption, not an inference.

Third, base rates. Even if the sample were unbiased and attribution were clean, most arbitrators have a published record of a few dozen awards at most, across unrelated subject matter. A “rules for claimants 58% of the time” figure from forty awards has a confidence interval wide enough to be meaningless, and it would be presented, inevitably, without one.

We tested this directly during development. We built per-arbitrator outcome models for a set of individuals with unusually large published records, held out a fifth of each record, and asked the models to predict outcomes on the held-out awards. Across the set, prediction accuracy was indistinguishable from predicting the venue base rate. The individual signal, if it exists, was not recoverable from the published record. We did not ship the feature, and we removed the per-arbitrator pipeline from the codebase so that it would not quietly return.

What we did instead

Simulation’s tribunal is a venue persona with adjustable procedural strictness and reasoning style, chosen by the user. If a team knows from experience that their chair is interventionist, they can set that. The setting is a hypothesis the team is responsible for, labelled as such in the scorecard, not a claim Simulation makes about a real person. No arbitrator or judge is named anywhere in the product.

This costs us something. A named-arbitrator feature would demonstrate well, and some products describe one. We think it would be selling a confidence that the evidence cannot bear, to professionals whose job is to know the difference.

Caveats on the venue models themselves

  • Coverage is uneven. Venues built on a few hundred documents carry wider uncertainty, which is noted in the venue selector.
  • The corpus is weighted to the last fifteen years. Older conventions may be under-represented.
  • Published awards over-represent larger disputes. Procedural handling of small claims may differ.
  • Decomposition is a model’s reading of a text, checked by hand on a sample. Residual error of a few points per field should be assumed.

A good venue model tells you how the room usually works. Nobody’s published record tells you how one person in it will decide your case.

See it on a contract you have already reviewed.

Send us a draft your team has already redlined and we will show you what ZAAN catches, and what it misses.