On an extended 26-operation kernel corpus measured across five GPU classes, a standard one-shape check accepted all ten deliberately-buggy LLM-style kernels as correct; the seeded gpuemu oracle caught ten of ten with no false positives on sixteen correct controls. On the 24-operation single-GPU corpus it caught nine of nine.
Evidence register
What we have measured, and what we have not.
13 claims, 5 of them measured. Each one names its source, the version it was read from, the method behind it, what it covers, and where it stops applying.
Why this page exists
Skelf Research sells one thing: evidence that is good enough to make a technical decision on. That only works if our own public claims meet the standard we would apply to a client's. So every number and every architectural statement this site makes about a Skelf system should trace to an entry on this page. If a page asserts something that is not here, treat it as an error and tell us.
The register is deliberately short. Several of our systems publish no benchmark at all — Route-Switch and NumaPerf say so on their own sites — because we have not run a measurement we would be willing to defend. For those projects the register records what the software does and the fact that the performance question is open, which is more useful to a reader than a number produced to fill the space. When a customer asks whether one of these tools would help on their workload, the honest answer is that we would have to measure it there.
How to read an entry
Kind separates what an experiment showed from what is true by construction and from what we are reasoning towards. Scope is the population the claim covers: a corpus, a scenario set, a kernel version. Method is how the result was produced. Source and version pin the claim to a commit, so it can be re-read even after the project moves on. Limitations state where the claim stops being true, which for a small, deliberately constructed test set is usually the most important line on the card.
A claim is reviewed whenever the code it describes changes, and the verification date shows when that last happened. Claims that a later version contradicts are marked stale rather than silently edited; claims we no longer stand behind are withdrawn and removed from this list. Where two of our own sources disagree — for example a product site quoting a different figure from the repository it describes — the repository wins and the disagreement is noted.
Nothing on this page comes from customer work. Commissioned investigations are confidential unless the client expressly agrees to publication, and they never feed this register or our public writing by default.
Measured
An experiment produced this result. The scope says what was measured; the limitations say what it does not show.
Per-operation calibrated tolerances (a p95-of-controls × 1.5 envelope) raised kernel-bug recall from 65% to 82% compared with a single fixed atol/rtol, at no cost in precision.
In a seven-strategy ablation of test-input generation, adversarial sampling reached 93% bug recall, while testing on regular shapes only missed every tail-mask bug.
On WareMax's built-in scenarios, learned dispatching policies match round-robin and nearest-robot heuristics but do not surpass them, because those scenarios are capacity- and destination-contention-bound.
In WareMax's deterministic warehouse simulation, a permutation-equivariant candidate-scoring dispatch policy trained with a reward targeting the delay its decision controls reached about 97% on-time attainment; a flattened MLP, or a naive dense or sparse reward, plateaued near the weakest heuristic at about 82–85%.
Design facts
How a system is built and what it deliberately does not do, read from its source and documentation.
Compere selects the next pairwise comparison with UCB1 and turns outcomes into ratings with Elo. It does not implement Bradley-Terry, Thurstone or TrueSkill.
Memista pairs SQLite metadata with a USearch HNSW index (inner product, F32) behind a three-endpoint Actix-web API. It is experimental at v0.1.x and has been tested below about 100,000 vectors.
MPL produces tamper-evident BLAKE3-hashed audit records, provenance and quality scores that map to common regulatory asks (SOX, GDPR, HIPAA, EU AI Act). These are evidence primitives, not certifications.
NumaPerf is a Rust runtime for latency-critical services that provides topology discovery, RAII thread pinning, explicit memory-placement policies, per-node scheduling and sharded structures, device locality, and locality observability.
Perishable's source is public, but its repository has no licence file, so Skelf does not describe it as open source until one is added.
Route-Switch is a self-hosted, OpenAI-compatible Go gateway that routes across providers by configured strategy, logs per-prompt success, cost and latency to DuckDB, and reruns MIPROv2 prompt optimisation against captured traces.
Savanty uses a DSPy-orchestrated LLM to translate an English description of a discrete constraint problem into Answer Set Programming over a canonical assign(Var, Value) contract, then solves it with Clingo, which finds a valid answer set or proves none exists.
ZViz is an OCI-compatible container runtime in Zig that layers namespaces, dropped capabilities, Landlock, seccomp-BPF and cgroups v2; 132 syscalls reach the host kernel natively, 24 are denied at seccomp, and socket is argument-filtered.
Reporting an error
If a claim here is wrong, out of date, or contradicted by the repository it cites, email contact@skelfresearch.com with the claim id. We would rather correct the register than defend it.