compere vs Bradley-Terry: Selection Policy vs Rating Model

compere pairs UCB1 pair selection with Elo ratings. Why that differs from Bradley-Terry and TrueSkill, and when each half of the ranking problem matters.

The problem

You have n items and you want them in order. The only reliable signal is a judge — a person, a model, an A/B test — looking at two of them and picking the better one. Each judgement costs money, time or attention. Exhaustive comparison of every pair needs n(n-1)/2 judgements: 4,950 for 100 items, 435 for 30. Past a few dozen items that budget is rarely available.

Most write-ups of this problem then compare “bandits” against “Bradley-Terry” against “TrueSkill” as if they were rival answers to one question. They are not. There are two separate questions, and the alternatives answer different ones:

  1. Which pair should I ask about next? This is a selection policy. Exhaustive enumeration, random sampling, round-robin tournaments and UCB1 are all selection policies.
  2. Given the outcomes so far, what is each item’s score? This is a rating model. Elo, Bradley-Terry, Thurstone, Plackett-Luce and TrueSkill are all rating models.

Every ranking system makes one choice from each list. compere’s choices are UCB1 for selection and Elo for rating, and it is explicit that it does not implement Bradley-Terry, Thurstone, TrueSkill or Glicko. This post explains what that means and when each half matters.

What compere is

compere is a Python 3.11+ package and FastAPI service for pairwise comparison ranking, with SQLAlchemy persistence (SQLite by default, PostgreSQL via DATABASE_URL). It uses exactly two well-known algorithms:

  • UCB1 for pair selection. Each entity gets a score of win_rate + c * sqrt(ln(N) / n_i), where the second term is an exploration bonus that shrinks as an entity accumulates comparisons. The two highest-scoring entities are paired next. New entities get a large weight so they are compared early. The exploration constant defaults to 1.414.
  • Elo for ratings. After each verdict, both entities’ ratings move towards the observed result by K * (actual - expected), with K = 32 and an initial rating of 1500 by default.

There is also a similarity-based pairing strategy that pairs close entities instead of UCB-optimal ones, for when you specifically want to resolve near-ties. Swapping it in does not change the Elo layer.

It ships as a service rather than only a library because the usual deployment is an interface that shows a human two things and records which one they preferred. The quickstart, using the documented HTTP API:

pip install compere
compere --port 8090    # OpenAPI docs at http://127.0.0.1:8090/docs

# register the things you want to rank
curl -X POST localhost:8090/entities/ \
  -H "Content-Type: application/json" -d '{"name": "headline-a"}'
curl -X POST localhost:8090/entities/ \
  -H "Content-Type: application/json" -d '{"name": "headline-b"}'

# ask UCB1 for the next pair
curl localhost:8090/mab/next_comparison
# {"entity1_id": 1, "entity2_id": 2}

# record the verdict, then read the Elo leaderboard
curl -X POST localhost:8090/comparisons/ \
  -H "Content-Type: application/json" \
  -d '{"entity1_id":1,"entity2_id":2,"selected_entity_id":1}'
curl localhost:8090/ratings

The judge can be anything that returns a winner: a human annotator, an LLM-as-judge call, or the outcome of an A/B test. compere does not care which.

Rating models, briefly

Elo updates ratings one comparison at a time. It is simple, online, and the leaderboard is a pure function of the votes — easy to explain to whoever has to act on it. Its weaknesses are well known: the result depends on the order in which votes arrive, and a rating is a single number with no attached uncertainty.

Bradley-Terry models the probability that i beats j as a function of two latent strengths and fits all strengths at once, usually by maximum likelihood over the whole set of outcomes. Fitted as a batch, it does not depend on vote order, and it is the standard statistical model for paired comparisons. Plackett-Luce generalises it to rankings of more than two items.

TrueSkill is a Bayesian rating model that keeps a mean and a variance for every player. The variance is the useful part: it tells you how sure the system is, which Elo cannot.

None of these says anything about which pair to ask about next. That is the selection policy’s job.

Why the selection policy matters

With a fixed budget, the selection policy decides where the budget goes. Random sampling spends much of it on pairs whose order was never in doubt. UCB1 starts broad, because nothing is known, then concentrates on entities whose position is still uncertain and leaves settled regions alone.

We are deliberately not quoting a savings figure. compere publishes none, and the honest answer depends on how many items you have, how close they are, and when you decide to stop. compere has no built-in “confidence reached” switch; the compere site documents stopping heuristics — top-k stability, the UCB1 recommendations settling into a small cluster, and bootstrap re-runs of the vote history — that you apply yourself. If you need a number, measure it: run the selection policy against a synthetic set whose true order you know and plot ordering error against budget. Accuracy at an unlimited budget tells you almost nothing, because every reasonable method converges eventually.

When compere is the right answer

  • Preference data for RLHF or reward models. Annotators pick between two completions; you want the votes spent where the ordering is still open.
  • Evaluation leaderboards. Ranking models, prompts or systems from head-to-head LLM-as-judge or human verdicts, where an interpretable leaderboard matters more than a statistically optimal one.
  • A/B content ranking. Ordering headlines, designs or copy from pairwise picks.
  • You want a service, not a notebook. An HTTP API with persistence, optional JWT auth and rate limiting is ready to sit behind a voting interface.

When compere is NOT the right answer

  • You need uncertainty on each rating. Elo gives a point estimate. If “how sure are we that A is above B?” is the question, TrueSkill or a Bayesian Bradley-Terry fit answers it and compere does not.
  • You are analysing a fixed batch offline. If the votes are already collected, adaptive selection has nothing left to do, and a batch Bradley-Terry fit (for example with the choix library) can use the data more efficiently than sequential Elo.
  • Raters disagree systematically, or preferences are not transitive. Elo assumes a single latent strength per item and a particular noise model. Structured rater noise is better handled by a model that represents it.
  • Implicit signal already exists. If you have clicks, conversions or dwell time, asking people is the expensive way to learn something you are already being told.
  • The items have features that predict order. A learning-to-rank model trained on those features will generalise to new items; pairwise voting will not.
  • The item count is tiny. With a handful of items, exhaustively comparing every pair is cheap and simpler to reason about.

Composing the halves

Because selection and rating are separate, they can be mixed. Nothing stops you collecting votes through compere’s UCB1 endpoint and then exporting the comparison history (GET /comparisons/) to fit Bradley-Terry or TrueSkill offline for a second opinion. If the two leaderboards agree on the decision you need to make, you can stop worrying about the rating model. If they disagree, that disagreement is itself the finding: your comparisons are noisier, or less transitive, than Elo assumes.