Services
Decisions we help resolve.
You bring a technical decision that is too consequential to make from a demo, a vendor claim or an internal hunch. We design the investigation, run it reproducibly, and tell you what works, what fails and what is worth building — including when the answer is no.
Agent & RAG evaluation
“Can we release this agent or retrieval change — and what will break?”
Release evidence for agents and retrieval: task suites, regression harness, failure taxonomy, recommendation.
Instruments: Promptel, Blogus, MPL, EmbedCache, Memista, Polymathy, Slorg, l0l1
Scope, method and deliverables → 02LLM cost–quality benchmark
“Can we cut model cost or latency without an unacceptable loss of quality?”
A cost–quality frontier on your own workload, with the quality threshold agreed before anything is measured.
Instruments: Route-Switch, EmbedCache, Compere
Scope, method and deliverables → 03Systems validation
“Are these kernels silently wrong — and is NUMA behind our p99?”
GPU kernel correctness campaigns and NUMA tail-latency investigations, narrowly scoped and reproducible.
Instruments: GPUEmu, NumaPerf
Scope, method and deliverables →How an engagement runs
- Brief. You describe the decision, the deadline, the current baseline and what evidence would change your mind, through the commissioning brief. No confidential detail is needed at this stage; we sign an NDA before it is.
- Diagnostic. A short, fixed-price piece of work that frames the question, reviews the baseline, fixes the acceptance threshold before anything is measured, and produces an experiment plan with a cost. Sometimes the diagnostic is the whole answer.
- Sprint or project. The controlled comparison or bounded prototype itself, with a reproducibility package you keep: harness, pinned versions, configurations, seeds and raw outputs.
- Decision report. A recommendation, the evidence for it, what was measured, what is inferred, what remains unknown, and the conditions under which the recommendation expires. See the sample report.
- Re-evaluation, if it earns its keep. Agreed testing when something material changes — a model upgrade, a new dataset, an infrastructure migration — rather than a retainer for being available.
What you are paying for
The work, not a conclusion. A reproducible finding that the proposed approach does not meet the agreed threshold satisfies the contract, and is often the most valuable result we deliver: it stops an expensive build before it starts. We recommend the right approach even when it is not one of our own tools, and we say so when a question is outside what we can answer well.
Our open-source systems are instruments, not the product. They let us build harnesses quickly and show how we work; whether one belongs in your stack is a separate question that the evidence answers. Every claim we make about them is listed, with its source and limits, in the evidence register.
| Engagement | You receive | Price (GBP) |
|---|---|---|
| Technical diagnostic | Problem framing, baseline review, experiment plan, scope and recommendation | £2,500–£5,000 |
| Evaluation or benchmark sprint | Controlled comparison, reproducible harness, failure analysis, decision report | £8,000–£20,000 |
| Applied R&D project | A bounded prototype or new method, experimental results, limitations, handover | £25,000–£75,000 |
| Recurring re-evaluation | Agreed testing after model, data, infrastructure or policy changes | £2,000–£8,000 / month |
| Sponsored research package | Defined research outputs with disclosed funding and publication terms | £15,000–£50,000 |
Prices exclude VAT. Compute, external reviewers, travel and third-party licences or data are quoted separately. A reproducible negative finding satisfies the contract; nothing is priced on a favourable result.
Also investigated, by arrangement
These questions sit close to our published research. We take them on when a qualified decision and the necessary data access exist, but we have not built dedicated pages for them yet.
- Private-data AI feasibility. Can this workflow run without sending data outside our boundary, at acceptable quality? Related research: EmbedCache, Memista, Polymathy, l0l1
- Untrusted code and client-side AI. How should generated code, browser features and credentials be isolated? Related research: ZViz, Perishable, Anouk
- Constraint-based decisions and ranking. Can these requirements be turned into decisions a solver can check? Related research: Savanty, Compere
- Warehouse-robotics policy evaluation. Which dispatching policy merits a real-world pilot? Related research: WareMax
- Generative-media pipeline feasibility. Can this pipeline meet our quality, cost and turnaround requirements? Related research: Direktor
Sponsored benchmarks, reference implementations and research work packages for funded programmes are covered on the partners page. Confidentiality, publication terms and our relationship with affiliated implementation businesses are set out under procurement.
Is this decision on your desk?
Tell us the decision, the deadline and what evidence would settle it. We reply with whether we can help, and if so the smallest investigation that would.