How do we know an optimised or LLM-generated GPU kernel is correct?

The short answer: not from one allclose

You cannot know it from the check most teams use. A single-shape allclose against a reference is blind to whole classes of kernel bugs, and on Skelf’s deliberately constructed corpus it accepted every buggy kernel it was given. What does work, for the bug classes studied, is an fp64 reference compared across adversarial, schema-aware input shapes and values, with per-operation calibrated tolerances and seeds that let any failure be replayed. That raises confidence substantially; it does not prove a kernel correct, and it tells you nothing about whether the kernel is fast.

How much validation a kernel needs

How much validation a kernel needs depends on where it came from and what it is allowed to touch.

Kernel provenanceMinimum validation before mergeWhy
Generated or rewritten by an LLMfp64 reference, adversarial shapes including non-block-aligned sizes, calibrated tolerances, seeded replayThe generator optimises for passing the check it sees; bugs cluster in masking, accumulation and normalisation
Hand-optimised by a kernel engineerSame oracle, plus targeted shapes for the tiling and masking paths the engineer changedHuman bugs are fewer but fall in the same places: edges of tiles and reductions
Autotuner picked a new block configurationRe-run the oracle across shapes relative to the new block sizesA shape that was block-aligned under the old config may not be under the new one
Compiler, Triton or CUDA toolkit upgradeRe-run the full corpus with fixed seeds and compare against the previous runCode generation changes can alter numerics without any source change
Vendor library call (cuBLAS, cuDNN) used as-isSpot-check at your shapes and dtypesLower risk, but your shapes may not be the ones the vendor tested most

If the kernel sits on a training or serving path where a silent numerical error would run for weeks at scale, use the full regime regardless of provenance. If it is an experiment that will be thrown away, a careful reference comparison at a handful of awkward shapes may be enough.

What Skelf has measured, and its limits

This is the one decision guide where Skelf has measured evidence. It comes from the gpuemu research studies, and each result has a stated boundary. The full write-up is on the GPU kernel correctness corpus evidence page.

Measured.

  • One-shape checks pass buggy kernels. On an extended 26-operation corpus measured on five GPU classes (RTX 3060, A10, L40S, A100 SXM4, H100 NVL), a standard one-shape check accepted all ten deliberately buggy LLM-style kernels as correct. The seeded gpuemu oracle caught ten of ten, with no false positives on sixteen correct control kernels. On the 24-operation single-GPU corpus it caught nine of nine (claim).
  • Calibrated tolerances matter. Replacing a single fixed atol/rtol with per-operation tolerances, set as an envelope of 1.5 times the 95th percentile of error on correct controls, raised bug recall from 65% to 82% with no loss of precision (claim).
  • Input generation matters more than people expect. In a seven-strategy ablation holding the kernel and oracle fixed, adversarial sampling reached 93% bug recall, while testing on regular shapes only missed every tail-mask bug (claim). The gpuemu product site currently quotes a different figure for this study; the repository README’s 93% is treated as authoritative until the two are reconciled.

The bug classes behind those numbers. These are described in the gpuemu README and are the ones the corpus contains:

  • Tail-mask leak. A kernel processes data in tiles of BLOCK elements and forgets to mask the last partial tile. It only goes wrong when a dimension is not a multiple of the block size, so a test at 256 with a block of 64 looks perfect while 257 is wrong.
  • Accumulator scale. A matmul writes acc = where it should write acc +=. On some shapes the result still falls within a loose relative tolerance.
  • Missing normalisation. Attention without the 1/√D scaling saturates softmax differently; at one shape the output can look plausible.
  • Online-softmax rescale. A flash-attention kernel forgets to rescale its accumulator after the running maximum changes. It is only wrong when the sequence length N exceeds BLOCK_N.

Inferred. Because each of these bugs is gated on a shape or value condition, any test regime that does not deliberately construct those conditions will under-report them. That reasoning extends to other gated bugs, such as stride or alignment assumptions, but Skelf has not measured those.

Unknown. Recall on arbitrary production kernels, and on bug classes absent from the corpus — race conditions, out-of-bounds reads that happen to return zeros, precision loss in long reductions. The corpus is small and deliberately constructed. Settling the general question would need a larger corpus of independently sourced bugs, ideally real regressions from open-source inference engines, scored blind.

Where this regime stops protecting you

  • The bug is not numerical. Memory errors and data races need tools built for them, such as NVIDIA Compute Sanitizer. A numerical oracle and a sanitiser are complementary.
  • The reference is wrong. An fp64 reference only helps if it implements the intended operation. If the specification is ambiguous, for example about masking semantics or reduction order, both kernel and reference can agree on the wrong answer.
  • Your numerical regime differs from the controls. Calibrated tolerances are fitted to control kernels. New dtypes, very long reductions or unusual value ranges need their own controls.
  • The question is speed. None of this measures performance. A kernel can pass every correctness check and be slower than the one it replaces.
  • The input space is unbounded. No finite sample of shapes proves correctness. The oracle reduces the chance of a silent bug; it does not remove it.

How to build a correctness gate for your own kernels

A team can build a credible correctness gate in a week, with or without gpuemu. The steps below are what matters.

  1. Write an fp64 reference for each operation, in plain PyTorch or NumPy on CPU, as close to the mathematical definition as possible. Keep it boring; its job is to be obviously right.
  2. Generate shapes from the operation’s schema, not from habit. For every tiled dimension, include sizes equal to the block size, one more and one less, a prime, 1, and a size smaller than one block. For attention, include sequence lengths both below and above BLOCK_N. Include the shapes production actually uses.
  3. Generate values adversarially. Large magnitudes to stress softmax and exponentials, mixed signs, values near zero, and long rows for reductions. Uniform random values in [0, 1) hide many of these problems.
  4. Calibrate tolerances per operation and dtype. Run known-correct implementations, such as the vendor library or your reference cast to the target dtype, over many seeds. Set each operation’s tolerance from the observed error distribution, for example 1.5 times the 95th percentile, rather than one global atol/rtol.
  5. Seed everything and log the seed. Every failure should replay exactly on another machine. If your random number generation differs between languages or devices, fix that first.
  6. Prove the gate can fail. Plant the four bug classes above in copies of your own kernels and confirm the gate rejects each one, and that it accepts the correct versions. A gate that has never caught a planted bug has not been tested.
  7. Run it in CI. Generating inputs, evaluating the fp64 reference, comparing and replaying need no GPU; the kernel itself runs wherever it runs, on a GPU runner or, for small Triton cases, in Triton’s CPU interpreter mode. Keep correctness and performance as separate jobs with separate pass criteria, so a speed-up can never be used to excuse a numerical regression.

If the kernels are going into production

If you are about to ship generated or heavily optimised kernels into a training run or a serving fleet, Skelf can build and run this regime against your operators: write the references, design the shape and value generators for your schemas, calibrate tolerances on your controls, plant bugs to prove the gate works, and report what was caught, what was not tested and what remains unknown. GPUEmu is one instrument we may use; the report does not depend on it. See Systems validation.

Related service: Systems validation · How we run an investigation · Evidence register

Questions buyers ask

Why does torch.allclose on one shape miss kernel bugs?
Many kernel bugs only fire under specific shapes or values: a missing tail mask only matters when a dimension is not a multiple of the block size, and an online-softmax rescale bug only matters when the sequence is longer than one block. One well-behaved shape never exercises those paths, so the kernel looks correct.
What did Skelf's kernel correctness study measure?
On an extended 26-operation corpus across five GPU classes, a standard one-shape check accepted all ten deliberately buggy LLM-style kernels, while the seeded gpuemu oracle caught ten of ten with no false positives on sixteen correct controls. The corpus is small and constructed, so it is not an estimate of recall on arbitrary kernels.
Should we use one atol and rtol for every operation?
Preferably not. Different operations and dtypes have different legitimate error envelopes. In Skelf's study, per-operation calibrated tolerances raised bug recall from 65% to 82% compared with a single fixed atol/rtol, at no cost in precision.
Can kernel correctness be checked in CI without a GPU?
The validation step can. gpuemu compares kernel output with an fp64 CPU reference and replays failures from a seed, and its validation step runs without a GPU, so it fits ordinary CI runners.
Does passing a correctness oracle mean the kernel is fast?
No. Correctness and performance are separate questions. Performance-focused benchmarks such as KernelBench and TritonBench rank kernels on speed; a correctness oracle tells you whether the fast kernel computes the right thing.

Is this decision on your desk?

If this decision is consequential, we can measure it on your workload: systems validation, scoped to your deadline and the evidence that would settle it.