Instead of one duel at a time, the LLM-as-a-judge receives K dossiers from the same indication in a single call and returns them in order of opportunity. Plackett–Luce reads that ranking as a chain of choices — best of K, then best of the remaining K−1 — and updates every target’s belief distribution at once, so one call buys K(K−1)/2 pairwise constraints.
Everything the judge is allowed to see about one target: genetic evidence, tractability, safety, biomarker and competitive situation, plus its modality and mechanism chips and the number of papers published on it up to the selected year. It is a description, never a score — the judge reads two (or K) of these and says which is the better opportunity.
The target’s quality is never observed, so the model carries a belief about it: a normal distribution. μ is the current best estimate of that quality; σ is how unsure the model still is. Every verdict moves μ and shrinks σ; every elapsed year without a verdict pushes σ back up.
In the chart below, μ is where a curve sits and σ is how wide it is. Each curve is the true density, so all of them enclose area 1: a well-judged target is a tall narrow spike, a barely-judged one a low broad hill. The area is integrated numerically and printed per curve in the table under the chart.
Publication counts are not a quality score — citing volume as merit would be circular. They are a statement about how much we know, so they enter as precision. Precision adds up with evidence, which is the conjugate-normal update, and it enters in three separate places:
| Target | Papers to date | New this year | βi | Rounds |
|---|
With c the round’s shared scale and A_r the number of dossiers sharing rank r:
Note what the sums run over: a dossier only accrues evidence from the choice steps up to and including its own placement. Coming last is informative precisely because it survived none of the earlier steps. As in the pairwise model the μ step scales with σi², so a well-characterised target resists a single surprising ranking.
Only differences of μ are identified, so a rating has no absolute meaning — it is a position relative to the other dossiers in the comparison chain. Two targets that were never compared, directly or through a chain of shared opponents, have no defined ordering. That is the statistical reason the default unit here is one indication: within NSCLC the judge is applying one consistent rubric, across indications it would silently be inventing one.
Ranking on μ − z·σ rather than μ is what stops a barely-examined target from jumping the queue: it has to earn certainty before its score can rise.