This page is an interactive illustration of the rating mathematics, not the QIW system. Publication counts, dossiers, latent qualities and — unless you rule on a round yourself — the judgements are all simulated. QIW runs the same equations, but over real dossiers assembled from sourced evidence and over verdicts from real LLM judges.
Target rating Model · Bayesian Bradley–Terry Evidence · pairwise LLM-as-a-judge 2010 – 2025

Rating target dossiers from pairwise judgements

Each target is described by a dossier. An LLM-as-a-judge reads two dossiers from the same indication and names the stronger opportunity. We never observe a target’s true quality — only a stream of noisy verdicts, arriving whenever new evidence is published — so each target carries a belief distribution: a mean μ for how good we currently think it is, and a standard deviation σ for how sure we are.

Reading the model

what a dossier is, and what μ and σ mean here

Everything the judge is allowed to see about one target: genetic evidence, tractability, safety, biomarker and competitive situation, plus its modality and mechanism chips and the number of papers published on it up to the selected year. It is a description, never a score — the judge reads two (or K) of these and says which is the better opportunity.

The target’s quality is never observed, so the model carries a belief about it: a normal distribution. μ is the current best estimate of that quality; σ is how unsure the model still is. Every verdict moves μ and shrinks σ; every elapsed year without a verdict pushes σ back up.

f(x)=1σ2πexp((xμ)22σ2),f(x)dx=1f(x)=\frac{1}{\sigma\sqrt{2\pi}}\; \exp\!\left(-\frac{(x-\mu)^2}{2\sigma^2}\right), \qquad \int_{-\infty}^{\infty} f(x)\,dx = 1

In the chart below, μ is where a curve sits and σ is how wide it is. Each curve is the true density, so all of them enclose area 1: a well-judged target is a tall narrow spike, a barely-judged one a low broad hill. The area is integrated numerically and printed per curve in the table under the chart.

Comparison pool

a rating only means something relative to what it was compared against

Timeline — evidence arriving year by year

2025
A judging round is a re-appraisal, not an arbitrary repetition: it happens because something was published. Scrub the year to replay the portfolio as it looked at the time — the whole run is a deterministic function of the year, so the history can never disagree with the current ratings.

Rating history

click the chart to jump to a year
The ribbon is ±1 σ. A band that is still wide late in the timeline is the model telling you the evidence does not yet support the ordering it is showing you.

Belief distributions at the selected year

Judging round

one duel
Rule on this round yourself and it is inserted into the selected year, then the whole timeline is re-simulated from 2010 so every downstream year reflects your verdict. The automatic rounds are drawn from the model’s own generative process, so the ratings converge on the latent qualities instead of picking up a bias that would look like a flaw in the algorithm.

Evidence entering the model

publication counts enter as precision, never as score

Publication counts are not a quality score — citing volume as merit would be circular. They are a statement about how much we know, so they enter as precision. Precision adds up with evidence, which is the conjugate-normal update, and it enters in three separate places:

σ0,i2=σprior21+niprior/nref(precision grows linearly with evidence)\sigma_{0,i}^{2}=\frac{\sigma_{\text{prior}}^{2}} {1+n_i^{\text{prior}}/n_{\text{ref}}} \qquad\text{(precision grows linearly with evidence)}
A target with 2 000 papers behind it does not deserve the same starting uncertainty as one with twenty. At n = n_ref the entry variance is halved.
βi2=β021+ni/nref(a well-documented target yields less noisy verdicts)\beta_i^{2}=\frac{\beta_0^{2}}{1+n_i/n_{\text{ref}}} \qquad\text{(a well-documented target yields less noisy verdicts)}
A thick dossier gives the judge more to go on, so its verdicts are less noisy. This is a strict generalisation of Weng–Lin: with one shared β it reduces to the published equations exactly, which is what the parity tests pin down.
Pr[i enters a round in year t]ni,tnew\Pr[\,i \text{ enters a round in year } t\,]\;\propto\; n_{i,t}^{\text{new}}
Targets with a busy year get re-judged more often — so their σ falls faster, without any change to the update rule itself.
σiσi2+τ2Δt(applied once per elapsed year)\sigma_i \;\leftarrow\; \sqrt{\sigma_i^{2}+\tau^{2}\,\Delta t} \qquad\text{(applied once per elapsed year)}
And every elapsed year pushes uncertainty back up, so a rating nobody has revisited since 2016 is not presented as if it were fresh.
Target Papers to date New this year βi Rounds

For target i judged against target q:

ciq=σi2+σq2+βi2+βq2c_{iq}=\sqrt{\sigma_i^{2}+\sigma_q^{2}+\beta_i^{2}+\beta_q^{2}} piq=11+exp((μqμi)/ciq)p_{iq}=\frac{1}{1+\exp\!\big((\mu_q-\mu_i)/c_{iq}\big)} siq={1judge prefers i12tie0judge prefers qs_{iq}=\begin{cases}1 & \text{judge prefers } i\\[2pt] \tfrac12 & \text{tie}\\[2pt] 0 & \text{judge prefers } q\end{cases} μiμi+σi2ciq(siqpiq)\mu_i \;\leftarrow\; \mu_i + \frac{\sigma_i^{2}}{c_{iq}}\,(s_{iq}-p_{iq}) σiσi1σi3ciq3piq(1piq)\sigma_i \;\leftarrow\; \sigma_i\sqrt{\,1-\frac{\sigma_i^{3}}{c_{iq}^{3}}\,p_{iq}\,(1-p_{iq})\,}

Two consequences worth pointing at. First, the μ step is proportional to σi²: an uncertain dossier moves far, a well-characterised one barely budges, so a surprising verdict about a settled target is treated as noise rather than news. Second, the σ shrink peaks at p = 0.5 — judging two dossiers we already believe are equals is the most informative call you can buy, which is the scheduling rule for an LLM-judge budget.

Why the pool matters

piq and qir depend on μiμq onlyμ is identified only up to a common shiftp_{iq}\ \text{and}\ q_{ir}\ \text{depend on } \mu_i-\mu_q \text{ only}\;\Longrightarrow\; \mu \text{ is identified only up to a common shift}

Only differences of μ are identified, so a rating has no absolute meaning — it is a position relative to the other dossiers in the comparison chain. Two targets that were never compared, directly or through a chain of shared opponents, have no defined ordering. That is the statistical reason the default unit here is one indication: within NSCLC the judge is applying one consistent rubric, across indications it would silently be inventing one.

rank score=μzσ(z=399.87% confidence bound)\text{rank score}\;=\;\mu - z\,\sigma \qquad (z=3 \Rightarrow \text{99.87\% confidence bound})

Ranking on μ − z·σ rather than μ is what stops a barely-examined target from jumping the queue: it has to earn certainty before its score can rise.