GPU kernel correctness corpus: what a one-shape check misses

On a deliberately constructed 26-op corpus run across five GPU classes, the standard one-shape check accepted all ten buggy LLM-style kernels. A seeded fp64 oracle caught ten of ten with no false positives on sixteen correct controls. This page covers what that shows, and what it does not.

The question

The usual way to decide whether a GPU kernel is correct is to run it once and compare the output with a reference using allclose and a fixed absolute and relative tolerance. That uses one shape, one dtype and one seed. LLM-kernel benchmarks use the same check, and kernels that pass it ship.

The study asked a narrow question: when an LLM-style kernel contains a realistic numerical bug, does that one-shape check notice? And if not, does a stricter regime notice without flagging correct kernels as wrong?

The question matters to anyone deploying generated or hand-optimised CUDA or Triton kernels. A kernel that is silently wrong does not crash. It degrades a model’s output, often in long-context or edge-case inputs that nobody is looking at, and it can pass CI for months.

What was built

GPUEmu is Skelf Research’s open-source correctness oracle for CUDA and Triton kernels. It replaces the single comparison with a regime that has five parts:

  • an fp64 reference: kernel output is compared against a CPU reference computed in double precision, per dtype;
  • op-schema-aware inputs: each operator has a generator that knows its shape constraints and produces boundary, regular and adversarial values;
  • per-op calibrated tolerances: a tolerance class per operator and dtype, fitted to correct control kernels, in place of one global atol/rtol;
  • seed-deterministic replay: every failing case can be reproduced exactly, on any machine, from its seed and input snapshot;
  • static PTX/SASS lint: register pressure, spills and instruction counts compared against a baseline.

The validation step itself needs no GPU, because the reference runs on CPU. The kernel under test runs wherever it normally runs.

The corpus

The corpus was built deliberately around bug classes that show up in LLM-generated kernel code and that a one-shape check is structurally unlikely to catch:

Bug classExampleWhy a one-shape check misses it
Tail-mask leakA softmax that does not mask the last partial tileOnly fires when the hidden dimension is not a multiple of the block size; a shape such as 256 looks fine
Accumulator scaleA matmul that assigns to the accumulator instead of adding to itOn the chosen shape the result can land inside the relative tolerance
Missing normalisationAttention without the 1/√D scaleSoftmax saturates differently, but one shape can still look correct
Online-softmax rescaleFlash-attention that forgets to rescale the accumulator after a max updateOnly wrong when the sequence is longer than one tile

There are two versions. The 24-op single-GPU corpus has nine buggy kernels. The extended 26-op cross-GPU corpus has ten buggy kernels and sixteen correct control kernels, and was run on five GPU classes: RTX 3060, A10, L40S, A100 SXM4 and H100 NVL. The controls are there to measure false positives. An oracle that flags everything would catch every bug and be useless.

Method

Each kernel was checked two ways:

  1. Baseline: the standard one-shape allclose comparison with a fixed tolerance.
  2. Seeded oracle: GPUEmu’s fp64 reference, schema-aware generated inputs and calibrated tolerances, with every input derived from a recorded seed.

A buggy kernel counts as caught if the check rejects it. A correct control counts as a false positive if the check rejects it. Three further studies varied one part of the regime at a time: tolerance calibration (P2), input generation strategy (P3, a seven-strategy ablation with the kernel and oracle held fixed) and static PTX metrics (P4).

The primary write-up is study P1, “The correctness illusion in LLM-generated GPU kernels” (arXiv:2606.20128). All figures on this page come from the GPUEmu repository README at commit 696b510bc21f.

Results

StudyComparisonResult
P1, extended 26-op cross-GPU corpusOne-shape allclose checkAccepted 10 of 10 buggy kernels as correct
P1, extended 26-op cross-GPU corpusSeeded oracleCaught 10 of 10; 0 false positives on 16 correct controls
P1, 24-op single-GPU corpusSeeded oracleCaught 9 of 9
P2, tolerance calibrationFixed atol/rtol vs per-op p95-of-controls × 1.5 envelopeRecall rose from 65% to 82%, at no cost in precision
P3, input generationSeven strategies, kernel and oracle fixedAdversarial sampling reached 93% recall; regular shapes only missed 100% of tail-mask bugs
P4, static PTX metricsStructural vs semantic regressionsRegister and instruction deltas track structural performance regressions across 5 GPU classes; semantic bugs compiled to identical PTX

Each result corresponds to a register entry: P1, P2 and P3. P4 is qualitative and is reported here as the README states it.

Taken together, the studies say that three things contribute separately. Inputs have to reach the shapes where the bug lives (P3). Tolerances have to be tight enough per operator to see the error, without being so tight that ordinary floating-point noise trips them (P2). And static inspection of the compiled code cannot replace either, because a kernel with a wrong formula can compile to the same instructions as a correct one (P4).

What this does and does not show

What it shows. On this corpus, the standard check is blind to all four bug classes, and the seeded oracle is not. The result held on five GPU classes from consumer to datacentre parts, and the oracle produced no false positives on the sixteen controls.

What it does not show.

  • It is not a recall estimate for real-world kernels. The corpus is small and was built on purpose. Its bugs were chosen because they are the kind a one-shape check misses. “Ten of ten” describes this set, not the bugs in your codebase.
  • It says nothing about bug classes the corpus does not contain. Race conditions, memory-safety errors, non-determinism across runs and precision loss in unusual numeric regimes were not tested here. Memory errors in particular are the job of tools such as NVIDIA Compute Sanitizer.
  • It covers correctness only. Nothing here says whether a kernel is fast, or faster than another. Performance belongs on the target hardware.
  • The tolerance results depend on the controls. The P2 calibration is fitted to correct control kernels. A workload with different numerical behaviour needs its own controls before the envelope means anything.

A note on sources: the GPUEmu product site currently quotes different figures; the repository is authoritative and the site is being corrected.

How to reproduce

According to the README, all four studies ship as LaTeX manuscripts with replayable run records, and the 24- and 26-op kernel corpus lives in the GPUEmu research program linked from the README’s Research & evidence section. Ready-made fp64 references for the built-in operator schemas ship in the GPUEmu Python client, and every oracle run is seed-deterministic, so a failure recorded on one machine replays exactly on another.

We are not giving step-by-step commands here, because the research program’s materials are the authority on how to run them. If you cannot reach those materials from the README link, ask us and we will tell you what is available and in what form.

What a customer correctness campaign adds

This study tells you the method works on a constructed set of known bugs. It cannot tell you whether your kernels are correct. A systems-validation engagement uses the same method on the things that are specific to you:

  • Your operators. Fused kernels, custom attention variants and quantised matmuls each need a reference and an input schema. Writing the reference is often where the first disagreement turns up.
  • Your shapes. The sequence lengths, batch sizes, head dimensions and padding your serving stack actually produces, including the awkward ones that are not multiples of a block size.
  • Your dtypes. bf16, fp16, fp8 and mixed-precision accumulation each move the line between a real error and normal rounding.
  • Your tolerance budget. We calibrate against your correct kernels and agree with your team, before measuring, how much deviation is acceptable for each operator.
  • Your CI. The output is a gate that fails a pull request with a seed you can replay, not a one-off report.

The deliverable separates what was measured on your kernels, what is inferred from it and what remains unknown. If a kernel you planned to ship does not meet the agreed threshold, that is a valid result, and it comes with the reproducer that shows why. For the decision this supports, see validating optimised GPU kernels.

Related service: Systems validation · All evidence

Questions buyers ask

Does this study show that GPUEmu catches every kernel bug?
No. The corpus is small and was built on purpose around four bug classes. It shows that a one-shape check is blind to those classes and that a seeded oracle is not. It gives no recall figure for arbitrary kernels or for bug classes the corpus does not contain.
Why did the standard check pass kernels that were wrong?
It tests one shape, one dtype and one seed. Each bug in the corpus only shows up under particular conditions, such as a dimension that is not a multiple of the block size or a sequence longer than one tile. If the chosen shape avoids those conditions, the wrong kernel matches the reference.
Is this a performance benchmark?
No. The study measures numerical correctness only. A related study found that semantic bugs can compile to identical PTX, so static performance metrics cannot stand in for a correctness check.
Which numbers should I trust if the GPUEmu website says something different?
The GPUEmu repository README, pinned at commit 696b510bc21f, is the authoritative source for these figures. The product site currently quotes different figures and is being corrected.
How would this apply to our own kernels?
Only a run on your kernels can tell you. A systems-validation engagement applies the same method to your operators, shapes, dtypes and tolerance budget, and wires the result into your CI.

Is this decision on your desk?

Tell us the decision, the deadline and what evidence would settle it. We reply with whether we can help, and if so the smallest investigation that would.