What we have measured, and what we have not.

13 claims, 5 of them measured. Each one names its source, the version it was read from, the method behind it, what it covers, and where it stops applying.

Why this page exists

Skelf Research sells one thing: evidence that is good enough to make a technical decision on. That only works if our own public claims meet the standard we would apply to a client's. So every number and every architectural statement this site makes about a Skelf system should trace to an entry on this page. If a page asserts something that is not here, treat it as an error and tell us.

The register is deliberately short. Several of our systems publish no benchmark at all — Route-Switch and NumaPerf say so on their own sites — because we have not run a measurement we would be willing to defend. For those projects the register records what the software does and the fact that the performance question is open, which is more useful to a reader than a number produced to fill the space. When a customer asks whether one of these tools would help on their workload, the honest answer is that we would have to measure it there.

How to read an entry

Kind separates what an experiment showed from what is true by construction and from what we are reasoning towards. Scope is the population the claim covers: a corpus, a scenario set, a kernel version. Method is how the result was produced. Source and version pin the claim to a commit, so it can be re-read even after the project moves on. Limitations state where the claim stops being true, which for a small, deliberately constructed test set is usually the most important line on the card.

A claim is reviewed whenever the code it describes changes, and the verification date shows when that last happened. Claims that a later version contradicts are marked stale rather than silently edited; claims we no longer stand behind are withdrawn and removed from this list. Where two of our own sources disagree — for example a product site quoting a different figure from the repository it describes — the repository wins and the disagreement is noted.

Nothing on this page comes from customer work. Commissioned investigations are confidential unless the client expressly agrees to publication, and they never feed this register or our public writing by default.

Measured

An experiment produced this result. The scope says what was measured; the limitations say what it does not show.

measured GPUEmu #gpuemu-p1-correctness-corpus

On an extended 26-operation kernel corpus measured across five GPU classes, a standard one-shape check accepted all ten deliberately-buggy LLM-style kernels as correct; the seeded gpuemu oracle caught ten of ten with no false positives on sixteen correct controls. On the 24-operation single-GPU corpus it caught nine of nine.

Scope
Ten (cross-GPU) and nine (single-GPU) deliberately constructed LLM-style bugs — tail-mask leaks, accumulator scale, missing normalisation, online-softmax rescale — on RTX 3060, A10, L40S, A100 SXM4 and H100 NVL.
Method
fp64 CPU reference oracle with seeded, schema-aware inputs compared against a single-shape allclose baseline; study P1, "The correctness illusion in LLM-generated GPU kernels" (arXiv:2606.20128).
Limitations
  • A small, deliberately constructed corpus; not an estimate of recall on arbitrary kernels.
  • Says nothing about bug classes the corpus does not contain.
  • Checks numerical correctness only, not performance.
  • The corpus repository the README links to (gpuemu-paper) is not publicly reachable as of the verification date, so independent reproduction currently needs a request.
Source
gpuemu@696b510bc21f · source dated 2026-07-02 · verified 2026-10-07 by Dipankar Sarkar
measured GPUEmu #gpuemu-p2-tolerance-calibration

Per-operation calibrated tolerances (a p95-of-controls × 1.5 envelope) raised kernel-bug recall from 65% to 82% compared with a single fixed atol/rtol, at no cost in precision.

Scope
The gpuemu research corpus; recall is over the study's seeded bug set.
Method
Study P2, operator-aware mixed-precision tolerance calibration.
Limitations
  • Calibration is fitted to control kernels; a workload with different numerical regimes needs its own controls.
Source
gpuemu@696b510bc21f · source dated 2026-07-02 · verified 2026-10-07 by Dipankar Sarkar
measured GPUEmu #gpuemu-p3-input-generation

In a seven-strategy ablation of test-input generation, adversarial sampling reached 93% bug recall, while testing on regular shapes only missed every tail-mask bug.

Scope
The gpuemu research corpus, holding the kernel and oracle fixed and varying only the input strategy.
Method
Study P3, test-input generation for tensor programs.
Limitations
  • The gpuemu product site currently quotes 99% for this study; the repository README (93%) is treated as authoritative until reconciled.
Source
gpuemu@696b510bc21f · source dated 2026-07-02 · verified 2026-10-07 by Dipankar Sarkar
measured WareMax #waremax-learned-vs-heuristic

On WareMax's built-in scenarios, learned dispatching policies match round-robin and nearest-robot heuristics but do not surpass them, because those scenarios are capacity- and destination-contention-bound.

Scope
WareMax built-in scenarios only; the simulator's tunable structure exists to study regimes where dispatch has more leverage.
Method
Comparison of trained policies against heuristic baselines in the accompanying paper.
Limitations
  • A negative result for these scenarios, not a general claim that learned dispatch never helps.
Source
waremax@abaa4fdaa5c6 · source dated 2026-07-02 · verified 2026-10-07 by Dipankar Sarkar
measured WareMax #waremax-representation-reward

In WareMax's deterministic warehouse simulation, a permutation-equivariant candidate-scoring dispatch policy trained with a reward targeting the delay its decision controls reached about 97% on-time attainment; a flattened MLP, or a naive dense or sparse reward, plateaued near the weakest heuristic at about 82–85%.

Scope
WareMax built-in robotic mobile-fulfilment scenarios, multi-seed.
Method
Reinforcement-learning ablation over policy representation and reward design, with multi-seed results committed under crates/waremax-gym/python/results in the repository.
Limitations
  • Simulation results; no physical-warehouse validation.
  • Figures are approximate as reported in the README and accompanying paper.
Source
waremax@abaa4fdaa5c6 · source dated 2026-07-02 · verified 2026-10-07 by Dipankar Sarkar

Design facts

How a system is built and what it deliberately does not do, read from its source and documentation.

design Compere #compere-ucb1-elo

Compere selects the next pairwise comparison with UCB1 and turns outcomes into ratings with Elo. It does not implement Bradley-Terry, Thurstone or TrueSkill.

Scope
The Python package and FastAPI service.
Method
Read from the repository and the project's documentation.
Limitations
  • No published measurement of comparisons saved; it depends on the items and the judges.
  • Elo assumes transitive preferences with a particular noise model.
Source
compere@9c27c6514bd0 · source dated 2026-07-02 · verified 2026-10-07 by Dipankar Sarkar
design Memista #memista-scope

Memista pairs SQLite metadata with a USearch HNSW index (inner product, F32) behind a three-endpoint Actix-web API. It is experimental at v0.1.x and has been tested below about 100,000 vectors.

Scope
The current crate; embedding dimensions are hardcoded and users fork to change them.
Method
Read from the repository README and source.
Limitations
  • No recall or latency benchmark has been published.
  • Above roughly a million vectors or at high query concurrency, a dedicated vector database is the better choice.
Source
memista@a66e8698a902 · source dated 2026-07-02 · verified 2026-10-07 by Dipankar Sarkar
design MPL #mpl-compliance-primitives

MPL produces tamper-evident BLAKE3-hashed audit records, provenance and quality scores that map to common regulatory asks (SOX, GDPR, HIPAA, EU AI Act). These are evidence primitives, not certifications.

Scope
The MPL sidecar proxy and SDKs.
Method
Read from the repository and the project's documentation.
Limitations
  • Using MPL does not make a system compliant; the operator's compliance programme owns the mapping.
Source
mpl@6694e1b60541 · source dated 2026-09-03 · verified 2026-10-07 by Dipankar Sarkar
design NumaPerf #numaperf-scope

NumaPerf is a Rust runtime for latency-critical services that provides topology discovery, RAII thread pinning, explicit memory-placement policies, per-node scheduling and sharded structures, device locality, and locality observability.

Scope
Linux x86_64 and aarch64, with graceful degradation on macOS.
Method
Read from the repository and the project's documentation.
Limitations
  • The project publishes no benchmark of its own; whether NUMA placement moves a service's p99 has to be measured on that service.
Source
numaperf@1a381808493c · source dated 2026-07-02 · verified 2026-10-07 by Dipankar Sarkar
design Perishable #perishable-licence

Perishable's source is public, but its repository has no licence file, so Skelf does not describe it as open source until one is added.

Scope
The GitHub repository as of the verification date.
Method
GitHub licence detection and a listing of the repository root.
Limitations
  • The project's own site currently says MIT; that statement is being corrected.
Source
perishable@e4e310b44379 · source dated 2026-07-02 · verified 2026-10-07 by Dipankar Sarkar
design Route-Switch #route-switch-mechanism

Route-Switch is a self-hosted, OpenAI-compatible Go gateway that routes across providers by configured strategy, logs per-prompt success, cost and latency to DuckDB, and reruns MIPROv2 prompt optimisation against captured traces.

Scope
The open-source gateway.
Method
Read from the repository and the project's documentation.
Limitations
  • The project publishes no latency, cost-saving or quality benchmark; the trade-off has to be measured on the user's own traffic.
Source
route-switch@33b54f35b1ef · source dated 2026-07-02 · verified 2026-10-07 by Dipankar Sarkar
design Savanty #savanty-asp-clingo

Savanty uses a DSPy-orchestrated LLM to translate an English description of a discrete constraint problem into Answer Set Programming over a canonical assign(Var, Value) contract, then solves it with Clingo, which finds a valid answer set or proves none exists.

Scope
Discrete constraint problems over finite domains — scheduling, assignment, timetabling, graph colouring, puzzles.
Method
Read from the source tree (savanty/asp_runtime.py, dspy_modules.py; clorm dependency) and the project's own documentation.
Limitations
  • The solver's guarantee covers the generated program, not whether the program faithfully encodes the English.
  • Not for continuous optimisation, machine learning, statistics or simulation.
  • No published translation-accuracy benchmark.
Source
savanty@e67b50ffd42e · source dated 2026-07-02 · verified 2026-10-07 by Dipankar Sarkar
design ZViz stale #zviz-selective-denial

ZViz is an OCI-compatible container runtime in Zig that layers namespaces, dropped capabilities, Landlock, seccomp-BPF and cgroups v2; 132 syscalls reach the host kernel natively, 24 are denied at seccomp, and socket is argument-filtered.

Scope
Linux 5.13+ with cgroups v2; weaker on kernels without Landlock.
Method
Read from the repository and the project's documentation of its syscall policy.
Limitations
  • The host kernel remains in the trusted computing base; it does not defend against a kernel exploit the way gVisor or a microVM can.
  • No overhead or cold-start benchmark has been published with raw data; the README quotes cold-start and syscall-speed figures that Skelf has not reproduced and does not repeat.
  • The README is internally inconsistent — its summary gives 132 allowed / 24 denied, while its architecture diagram shows 90 allowed, 22 denied and 5 mediated by a userspace broker. Under review until the repository is reconciled.
  • The repository's LICENSE file is MIT; the README badge says Apache-2.0.
Source
zviz@470e9cfa03bb · source dated 2026-07-02 · verified 2026-10-07 by Dipankar Sarkar

Reporting an error

If a claim here is wrong, out of date, or contradicted by the repository it cites, email contact@skelfresearch.com with the claim id. We would rather correct the register than defend it.