Service
Systems validation
Correctness campaigns for teams shipping optimised or generated GPU kernels, and NUMA tail-latency investigations for performance leads on multi-socket hosts. Both are narrow, reproducible, and end in evidence your engineers can rerun.
The decision
Are these kernels silently wrong, and is NUMA placement the reason our p99 is where it is?
Usually commissioned by
Kernel or compiler team lead · Inference platform lead · Performance engineering lead
You receive
- Kernel correctness campaign: scoped op list checked against fp64 references on adversarial shapes
- Calibrated per-op tolerances and reproducible failing seeds for every bug found
- A correctness regression harness that runs in CI without a GPU
- NUMA topology audit and controlled pinning and placement experiments on your hosts
- Cross-node traffic attribution and a recommendation, including 'NUMA is not your problem'
Two questions, one discipline
This offer covers two problems that look unrelated but share a method. In both, the system appears to work, the usual check passes, and the failure only shows up under conditions that ordinary testing does not reach: an awkward tensor shape, or a thread that happened to be scheduled on the far socket. In both, the job is to construct those conditions on purpose, measure carefully, and hand back something a team can rerun.
Both strands are narrow. We take them on because Skelf has built and published the instruments behind them. We do not offer general GPU performance consulting, compiler development, or whole-system performance audits, and if your question sits outside what is described here we will say so in the first conversation rather than stretch the scope.
Strand A: GPU kernel correctness campaigns
For: inference vendors, kernel and compiler teams, and anyone shipping hand-optimised, autotuned or LLM-generated CUDA or Triton kernels.
The question: are these kernels silently wrong?
A kernel that crashes gets fixed. A kernel that returns plausible numbers that are slightly wrong, on some shapes, ships. The usual check, comparing output against a reference with allclose on one or two convenient shapes, is blind to whole classes of bug. Tail-mask errors appear only when a dimension is not a multiple of the block size. Accumulator scaling and online-softmax rescaling errors may stay inside a loose tolerance on small inputs and leave it on large ones. A single fixed tolerance is either too loose for well-conditioned operations or too tight for reduced-precision ones.
What the published work shows, and where it stops
The method rests on three studies from GPUEmu, each recorded in the evidence register with its scope:
| Study | Measured result | Boundary |
|---|---|---|
| Correctness corpus (P1) | On a 26-op corpus across five GPU classes, a standard one-shape check accepted all ten deliberately buggy LLM-style kernels; the seeded fp64 oracle caught ten of ten with no false positives on sixteen correct controls. Nine of nine on the 24-op single-GPU corpus. | A small, deliberately constructed corpus. Not an estimate of recall on arbitrary kernels, and silent on bug classes it does not contain. |
| Tolerance calibration (P2) | Per-op calibrated tolerances raised bug recall from 65% to 82% against a single fixed atol/rtol, with no loss of precision. | Calibration is fitted to control kernels; a workload with different numerical regimes needs its own controls. |
| Input generation (P3) | Adversarial sampling reached 93% recall in a seven-strategy ablation; regular shapes only missed every tail-mask bug. | Measured on the research corpus with kernel and oracle held fixed. |
All three are about correctness. None says anything about kernel performance. The corpus write-up has the detail.
How a campaign runs
- Scope the op list. We agree which operations matter most, usually those on the hot path of your production models or those most recently rewritten or generated.
- Build fp64 references for each op, on the CPU, so the truth does not depend on the hardware under test.
- Generate adversarial inputs that respect each op’s schema: non-multiple dimensions, extreme magnitudes, degenerate shapes, the edges where masking and accumulation go wrong.
- Calibrate tolerances per op from known-correct controls, so a reduced-precision op is not failed for being reduced-precision, and a full-precision op is not passed for being nearly right.
- Reproduce every failure from a recorded seed, so your engineers can step through it.
- Hand over a CI harness that reruns the campaign without a GPU on every kernel change.
The kernel validation guide sets out the decision in more depth. If your existing test infrastructure can host the oracle and inputs, we build there rather than install new tooling.
Strand B: NUMA tail-latency investigations
For: performance and SRE leads running latency-critical services on multi-socket or multi-node hosts.
The question: is NUMA placement behind our p99?
On a NUMA machine, memory attached to another socket costs more to reach than local memory. A service whose threads and data drift apart can show a healthy median and an erratic tail. But many tail-latency problems that look like NUMA are something else: lock contention, allocator behaviour, garbage collection, noisy neighbours, the network. Changing placement on a hunch can make things no better and harder to reason about.
What we do
- Topology audit. What the hardware and firmware actually expose: node count, distances, which cores and devices (NICs, NVMe, GPUs) sit on which node, and firmware settings such as node interleaving that change the picture.
- Baseline under representative load, at p50, p99 and beyond, on the hosts you run in production or a faithful copy.
- Controlled placement experiments. Pin the hot path to one node, bind its memory locally, then vary one thing at a time. Each experiment is recorded so it can be rerun.
- Cross-node traffic attribution with numastat and hardware counters, so a change in tail latency can be tied to a change in remote memory access, or shown not to be.
- A recommendation: which placement change, if any, is justified, and the evidence for it.
NumaPerf is a Rust runtime offering topology discovery, RAII thread pinning, explicit placement policies, per-node scheduling and locality observability (claim record). It publishes no benchmarks of its own, and we will not quote any; whether placement moves a service’s p99 has to be measured on that service. Most investigations need only standard Linux tooling. NumaPerf becomes relevant only when the fix requires placement control inside a Rust codebase. The NumaPerf essay on when NUMA actually matters describes the filtering we start with.
A finding that NUMA is not your problem is a valid outcome. It comes with the experiments that ruled placement out, which tells your team where not to spend the next month.
What we need from you
| Kernel correctness | NUMA investigation | |
|---|---|---|
| Code | Kernels under test and the reference semantics for each op | The service, or a build of it we can run under load |
| Hardware | Access to the GPU classes you ship on, for confirmation | Production-class hosts, or a copy with the same topology |
| Load | Representative shapes from your models | A load generator or replayed traffic that reproduces the tail |
| People | A kernel engineer to triage findings | An engineer who can change pinning and deploy settings |
How it typically runs
A technical diagnostic (£2,500–£5,000) scopes the op list or audits the topology, reviews existing tests or latency data, and proposes the campaign or experiment plan. For NUMA, the diagnostic can settle the matter: a single-node box, or a throughput-bound workload, rarely needs more.
A sprint (£8,000–£20,000) runs the campaign or the placement experiments and delivers the harness, the findings and the recommendation. Recurring re-evaluation suits kernel teams whose op library changes often, rerunning the correctness campaign on new or regenerated kernels. Our methods and procurement pages cover reporting standards, confidentiality and contracting.
| Engagement | You receive | Price (GBP) |
|---|---|---|
| Technical diagnostic | Problem framing, baseline review, experiment plan, scope and recommendation | £2,500–£5,000 |
| Evaluation or benchmark sprint | Controlled comparison, reproducible harness, failure analysis, decision report | £8,000–£20,000 |
| Applied R&D project | A bounded prototype or new method, experimental results, limitations, handover | £25,000–£75,000 |
| Recurring re-evaluation | Agreed testing after model, data, infrastructure or policy changes | £2,000–£8,000 / month |
| Sponsored research package | Defined research outputs with disclosed funding and publication terms | £15,000–£50,000 |
Prices exclude VAT. Compute, external reviewers, travel and third-party licences or data are quoted separately. A reproducible negative finding satisfies the contract; nothing is priced on a favourable result.
Decision guides
Instruments we may use
Chosen only where they suit the question. Maturity and licence are checked per engagement.
Questions buyers ask
- Does a kernel correctness campaign tell us whether our kernels are fast?
- No. It checks numerical correctness only. Performance profiling is a different question, and we will say so if that is the one you actually need answered.
- How much of our kernel library can you cover?
- A scoped op list agreed in the diagnostic, prioritised by where silent errors would cost you most. Our published recall figures come from a small, deliberately constructed corpus and are not a promise of recall on your kernels.
- Do we need to give you GPU access?
- For the campaign itself, access to the GPU classes you ship on is useful so failures can be confirmed on real hardware. The regression harness we hand over runs against an fp64 CPU reference and does not need a GPU in CI.
- What if the NUMA investigation finds nothing?
- Then you receive a documented, reproducible finding that placement is not driving your tail, and the experiments that ruled it out. That narrows the search and is a valid outcome of the engagement.
- Do we have to adopt NumaPerf to act on the recommendation?
- No. Many placement fixes are configuration: numactl policies, cgroup cpusets, IRQ affinity or a BIOS setting. NumaPerf is a Rust library and only relevant if code-level placement control is the right fix for your service.
Is this decision on your desk?
Tell us the decision, the deadline and what evidence would settle it. We reply with whether we can help, and if so the smallest investigation that would.