LLM cost–quality benchmark

A controlled benchmark for CTOs and inference leads deciding whether to cut model spend or tail latency. We fix the quality bar first, test candidates on your traffic, and report cost per accepted result, even when the answer is to keep the current model.

The decision

Can we cut model cost or tail latency without an unacceptable loss of quality?

Usually commissioned by

CTO · Inference or platform lead · FinOps sponsor

You receive

  • Quality threshold and grading method agreed in writing before any measurement
  • Candidate set: smaller models, routed mixes, prompt changes and caching, run on your traffic
  • Cost per accepted result, not per call, for every candidate
  • Results sliced by task type, with the hard slice reported separately
  • Latency at p50 and p95, and a failure analysis of where candidates disagree
  • A recommendation, in which keeping the current model is a valid outcome

The question, and the version of it we can answer

“Can we use a cheaper model?” is usually asked as if it had one answer. It does not. A smaller model may match the current one on routine requests and fall apart on the minority of requests that carry most of the business risk. A routed mix may cut spend while adding a classification step that worsens tail latency. A shorter prompt may save tokens and quietly drop an instruction a downstream system relies on. Every one of those trade-offs is real, and every one is invisible in an average.

So we turn the question into one that can be answered with evidence: on a defined sample of your traffic, which candidates deliver results that pass a quality bar you set in advance, at what cost per accepted result, and at what latency? The answer can be “this candidate, for these task types, with this fallback”. It can also be “none of them”, and that is a useful answer.

AI spend has become routine work for finance and platform teams. The FinOps Foundation’s State of FinOps 2026 reports that 98% of respondents now manage AI spend and names AI cost management as the top skill set teams need to develop (data.finops.org). The difficulty is rarely seeing the bill; it is knowing what a cheaper configuration would do to quality before you commit to it.

Step one: fix the quality bar before measuring anything

The most common way a model-cost study goes wrong is choosing the threshold after seeing the results. We prevent that by agreeing, in writing, three things before any candidate runs:

  1. What counts as an accepted result for each task type: an exact match, a passing validator, a rubric score above a set level, or a human reviewer’s sign-off.
  2. The minimum acceptance rate each task type must reach, and whether any single failure type disqualifies a candidate outright.
  3. The latency budget, stated as a percentile, because a p50 inside budget says nothing about the users who wait at p95.

Where outputs are open-ended and no reference answer exists, absolute scores are unreliable. Pairwise judgements (“which of these two is better for this request?”) are easier to make consistently. Compere chooses which pairs to compare with UCB1 and turns outcomes into Elo ratings, which assumes reasonably transitive preferences; we use it when that assumption suits the task and a simpler rubric when it does not.

The candidates

Candidate typeWhat it testsTypical risk
Smaller or cheaper modelWhether quality holds on your task mixCollapse on the hard slice
Routed mixWhether easy requests can go cheap and hard ones stay expensiveThe router misclassifies; extra hop adds latency
Prompt changeWhether fewer tokens or a different structure keep qualitySilent loss of an instruction downstream code needs
CachingWhether repeated or near-repeated work can be served without a model callStale or wrong-context answers
Current configurationThe baseline every candidate must beatNone; it is the control

For embedding-heavy workloads, EmbedCache tests a specific cost lever: generating embeddings locally with ONNX models on CPU and caching them by content hash and model id, so the same text is never embedded twice. Whether that beats a hosted embedding API depends on your volume, your hardware and whether a local model’s retrieval quality is good enough, which is part of what the benchmark measures.

Why the hard slice is reported on its own

Averages hide the failures that matter. We split results by task type and by difficulty, and we always report the hardest slice separately, defined before the run, not chosen afterwards. If a candidate is within threshold overall but fails the hard slice, the recommendation says so and usually proposes routing that slice to a stronger model rather than rejecting the candidate entirely.

Where the current model and a candidate disagree, we read the disagreements. Some are cases where the candidate is wrong. Some are cases where the current model was wrong and nobody had noticed. Both go into the failure analysis, because the second kind changes what “keeping the current model” actually means.

Instruments, and their limits

Route-Switch is a self-hosted, OpenAI-compatible gateway that routes by configured strategy, logs per-prompt success, cost and latency to DuckDB, and can rerun MIPROv2 prompt optimisation on captured traces (claim record). It is useful when you have no per-request logging of cost and outcome, or when the candidate under test is a routed mix. It publishes no latency, cost-saving or quality benchmark, and we will not quote one. The whole point of this benchmark is that the trade-off has to be measured on your traffic; the routing-triangle essay sets out why.

If your existing gateway, tracing or observability stack already records what we need, we use that instead. The harness we leave behind should run on the infrastructure you already operate.

What we need from you

  • A sample of production requests large enough to cover each task type, with the hard slice represented. We help size it; redaction is fine if it preserves the task.
  • Current pricing, including any negotiated rates, so cost figures reflect your contracts and not list prices.
  • API access or credentials to the candidate models, held by you where possible.
  • One or two people who can say whether an output is acceptable, for grading and for checking any model-based grader.
  • A sponsor who owns the budget line and will act on the recommendation.

Limits

The benchmark tells you how candidates performed on the sample, at the prices and model versions in force during the run. It does not predict how providers will change prices or models, and it does not cover task types missing from the sample. If your traffic mix shifts materially, the result needs re-running. Our measured, inferred and unknown findings are kept separate in the report, following the methods we apply to all published work.

Typical shape

When the quality bar and traffic sample already exist, the work starts as an evaluation sprint (£8,000–£20,000): candidates run, cost per accepted result calculated, slices and disagreements analysed, recommendation written. When they do not, a technical diagnostic (£2,500–£5,000) comes first to agree the threshold, size the sample and choose candidates. Recurring re-evaluation reruns the harness when a provider ships a new model, changes pricing, or your prompts change. The smaller-model replacement guide covers the decision in more depth, and procurement covers contracting and data terms.

How engagements are priced
EngagementYou receivePrice (GBP)
Technical diagnostic Problem framing, baseline review, experiment plan, scope and recommendation £2,500–£5,000
Evaluation or benchmark sprint Controlled comparison, reproducible harness, failure analysis, decision report £8,000–£20,000
Applied R&D project A bounded prototype or new method, experimental results, limitations, handover £25,000–£75,000
Recurring re-evaluation Agreed testing after model, data, infrastructure or policy changes £2,000–£8,000 / month
Sponsored research package Defined research outputs with disclosed funding and publication terms £15,000–£50,000

Prices exclude VAT. Compute, external reviewers, travel and third-party licences or data are quoted separately. A reproducible negative finding satisfies the contract; nothing is priced on a favourable result.

Decision guides

Instruments we may use

Chosen only where they suit the question. Maturity and licence are checked per engagement.

Questions buyers ask

Why measure cost per accepted result instead of cost per call?
Because a cheap call that produces an answer you have to reject, retry or escalate is not cheap. Dividing total spend, including retries and fallbacks, by the number of outputs that pass the agreed quality bar gives the figure a budget owner can act on.
Will you tell us how much we will save?
Only after measuring on your traffic. We do not quote expected savings in advance, and Route-Switch, our routing gateway, publishes no cost or quality benchmark for exactly this reason: the trade-off depends on the workload.
What if no candidate meets the threshold?
Then the recommendation is to keep the current model, with the evidence showing where each candidate fell short. That result satisfies the engagement and stops a migration that would have cost more than it saved.
Do we need to route traffic through Route-Switch?
No. If your gateway or observability stack already logs per-request cost, latency and outcome, we use it. Route-Switch is an option when you have no such logging, and it is self-hosted, so traffic does not pass through Skelf.
How do you grade quality when there is no single right answer?
We agree a rubric with your domain owners and, where outputs are open-ended, use pairwise comparisons graded by people or by a model whose agreement with people we have checked on a sample.

Is this decision on your desk?

Tell us the decision, the deadline and what evidence would settle it. We reply with whether we can help, and if so the smallest investigation that would.