This page is an interactive illustration of the rating mathematics, not the QIW system. Publication counts, dossiers, latent qualities and — unless you rule on a round yourself — the judgements are all simulated. QIW runs the same equations, but over real dossiers assembled from sourced evidence and over verdicts from real LLM judges.
Target rating Model · Plackett–Luce Evidence · K-way LLM-as-a-judge ranking 2010 – 2025

Rating target dossiers from K-way rankings

Instead of one duel at a time, the LLM-as-a-judge receives K dossiers from the same indication in a single call and returns them in order of opportunity. Plackett–Luce reads that ranking as a chain of choices — best of K, then best of the remaining K−1 — and updates every target’s belief distribution at once, so one call buys K(K−1)/2 pairwise constraints.

Reading the model

what a dossier is, and what μ and σ mean here

Everything the judge is allowed to see about one target: genetic evidence, tractability, safety, biomarker and competitive situation, plus its modality and mechanism chips and the number of papers published on it up to the selected year. It is a description, never a score — the judge reads two (or K) of these and says which is the better opportunity.

The target’s quality is never observed, so the model carries a belief about it: a normal distribution. μ is the current best estimate of that quality; σ is how unsure the model still is. Every verdict moves μ and shrinks σ; every elapsed year without a verdict pushes σ back up.

f(x)=1σ2πexp((xμ)22σ2),f(x)dx=1f(x)=\frac{1}{\sigma\sqrt{2\pi}}\; \exp\!\left(-\frac{(x-\mu)^2}{2\sigma^2}\right), \qquad \int_{-\infty}^{\infty} f(x)\,dx = 1

In the chart below, μ is where a curve sits and σ is how wide it is. Each curve is the true density, so all of them enclose area 1: a well-judged target is a tall narrow spike, a barely-judged one a low broad hill. The area is integrated numerically and printed per curve in the table under the chart.

Comparison pool

a rating only means something relative to what it was compared against

Timeline — evidence arriving year by year

2025
A judging round is a re-appraisal, not an arbitrary repetition: it happens because something was published. Scrub the year to replay the portfolio as it looked at the time — the whole run is a deterministic function of the year, so the history can never disagree with the current ratings.

Rating history

click the chart to jump to a year
The ribbon is ±1 σ. A band that is still wide late in the timeline is the model telling you the evidence does not yet support the ordering it is showing you.

Belief distributions at the selected year

Judging round

one ranking
Rule on this round yourself and it is inserted into the selected year, then the whole timeline is re-simulated from 2010 so every downstream year reflects your verdict. The automatic rounds are drawn from the model’s own generative process, so the ratings converge on the latent qualities instead of picking up a bias that would look like a flaw in the algorithm.

Evidence entering the model

publication counts enter as precision, never as score

Publication counts are not a quality score — citing volume as merit would be circular. They are a statement about how much we know, so they enter as precision. Precision adds up with evidence, which is the conjugate-normal update, and it enters in three separate places:

σ0,i2=σprior21+niprior/nref(precision grows linearly with evidence)\sigma_{0,i}^{2}=\frac{\sigma_{\text{prior}}^{2}} {1+n_i^{\text{prior}}/n_{\text{ref}}} \qquad\text{(precision grows linearly with evidence)}
A target with 2 000 papers behind it does not deserve the same starting uncertainty as one with twenty. At n = n_ref the entry variance is halved.
βi2=β021+ni/nref(a well-documented target yields less noisy verdicts)\beta_i^{2}=\frac{\beta_0^{2}}{1+n_i/n_{\text{ref}}} \qquad\text{(a well-documented target yields less noisy verdicts)}
A thick dossier gives the judge more to go on, so its verdicts are less noisy. This is a strict generalisation of Weng–Lin: with one shared β it reduces to the published equations exactly, which is what the parity tests pin down.
Pr[i enters a round in year t]ni,tnew\Pr[\,i \text{ enters a round in year } t\,]\;\propto\; n_{i,t}^{\text{new}}
Targets with a busy year get re-judged more often — so their σ falls faster, without any change to the update rule itself.
σiσi2+τ2Δt(applied once per elapsed year)\sigma_i \;\leftarrow\; \sqrt{\sigma_i^{2}+\tau^{2}\,\Delta t} \qquad\text{(applied once per elapsed year)}
And every elapsed year pushes uncertainty back up, so a rating nobody has revisited since 2016 is not presented as if it were fresh.
Target Papers to date New this year βi Rounds

With c the round’s shared scale and A_r the number of dossiers sharing rank r:

c=q(σq2+βq2)c=\sqrt{\textstyle\sum_{q}\big(\sigma_q^{2}+\beta_q^{2}\big)} qir=eμi/cj:rankj  reμj/cq_{ir}=\frac{e^{\mu_i/c}}{\sum_{j\,:\,\text{rank}_j\ \ge\ r} e^{\mu_j/c}} ωi=σi2crranki𝟙[i chosen at r]qirAr\omega_i=\frac{\sigma_i^{2}}{c}\sum_{r\,\le\,\text{rank}_i} \frac{\mathbb{1}[\,i \text{ chosen at } r\,]-q_{ir}}{A_r} δi=σi3c3rrankiqir(1qir)Ar\delta_i=\frac{\sigma_i^{3}}{c^{3}}\sum_{r\,\le\,\text{rank}_i} \frac{q_{ir}\,(1-q_{ir})}{A_r} μiμi+ωi,σiσi1δi\mu_i \leftarrow \mu_i+\omega_i, \qquad \sigma_i \leftarrow \sigma_i\sqrt{1-\delta_i}

Note what the sums run over: a dossier only accrues evidence from the choice steps up to and including its own placement. Coming last is informative precisely because it survived none of the earlier steps. As in the pairwise model the μ step scales with σi², so a well-characterised target resists a single surprising ranking.

Why the pool matters

piq and qir depend on μiμq onlyμ is identified only up to a common shiftp_{iq}\ \text{and}\ q_{ir}\ \text{depend on } \mu_i-\mu_q \text{ only}\;\Longrightarrow\; \mu \text{ is identified only up to a common shift}

Only differences of μ are identified, so a rating has no absolute meaning — it is a position relative to the other dossiers in the comparison chain. Two targets that were never compared, directly or through a chain of shared opponents, have no defined ordering. That is the statistical reason the default unit here is one indication: within NSCLC the judge is applying one consistent rubric, across indications it would silently be inventing one.

rank score=μzσ(z=399.87% confidence bound)\text{rank score}\;=\;\mu - z\,\sigma \qquad (z=3 \Rightarrow \text{99.87\% confidence bound})

Ranking on μ − z·σ rather than μ is what stops a barely-examined target from jumping the queue: it has to earn certainty before its score can rise.