From a blocked decision to evidence you can defend.

Every investigation follows the same discipline, whether the question is about an agent, a model migration or a GPU kernel. The point is not ceremony. It is that the person who has to sign off the decision can see what was measured, how, against what, and where the conclusion stops applying.

1. Start from the decision, not the technology

We write down the mandate before designing anything. If a field cannot be filled in, that is usually the first finding: a question with no decision attached, or a decision with no deadline, rarely justifies an investigation.

Trigger
What happened that makes the decision necessary now.
Blocked decision
The decision that cannot yet be made, in your words.
Baseline
What runs today, and what is already known about it.
Constraints
Data handling, deployment, licensing, security, regulation, time.
Evidence required
The result that would change the decision — and what would not convince you.
Acceptance
How both sides will know the work is done, including if the answer is no.
Deadline
When the decision will be made, with or without us.

2. Fix the threshold before measuring anything

The acceptance threshold — how much quality loss is tolerable, what tail latency is acceptable, which failure classes are disqualifying — is agreed in writing with the people who own the decision before any experiment runs. Choosing the bar after seeing the results is the most common way an evaluation ends up confirming whatever someone hoped for. If the bar later turns out to be wrong, we say so and report results against both the original and the revised bar.

3. Measure against a real baseline, under controls

A candidate is only better or worse than something. We measure the system you run today, under the same conditions as every candidate: same data, same harness, same grading, versions pinned. Where results vary between runs, we repeat them and report the spread rather than the best run. Where an average could hide the failures that matter, we fix a hard slice of the workload in advance and report it separately.

4. Separate what was measured from what is believed

Every report, and every public page we write, sorts its claims into four kinds: measured (an experiment produced it, with a stated scope), design facts (how a system is built), inferred (reasoning from the first two) and unknown. Our own public claims are held to the same rule in the evidence register, where each entry names its source, pinned version, method and limitations. When two of our own sources disagree, the repository wins and the disagreement is recorded rather than smoothed over.

5. Look at the failures, not just the score

A pass rate tells you how often something worked; a failure taxonomy tells you whether the failures are tolerable. We classify every failure we observe, keep the examples, and report which classes the candidate introduces that the baseline did not have. For correctness work, a failure is only useful if it replays: every failing case ships with the seed or input that reproduces it.

6. Hand over something reproducible

The deliverable includes a reproducibility package: the harness, configurations, pinned model and library versions, seeds, references to frozen data (which stays in your environment where required), raw outputs and the analysis that turned them into conclusions. Anyone on your team should be able to rerun it without us. Our own tools may be part of the harness; they are open source, so nothing about the method depends on a black box.

7. Review, then state the limits

Before delivery, the method and conclusions are reviewed by someone who did not run the experiments — internally, or by an external reviewer you have approved where the stakes justify it. The report ends with its own limitations: what the evidence does not show, and the conditions under which the recommendation should be revisited.

8. Re-evaluate when something material changes

Recommendations expire. A model upgrade, a new class of input, a data migration or a change of hardware can invalidate a result that was sound when it was produced. Where it earns its keep, we agree in advance which changes trigger a re-run of the harness, rather than offering a retainer for being available.

What we will not do

  • Choose the threshold after seeing the results, or report only the runs that went well.
  • Publish a number we have not measured, about your system or our own.
  • Tie a conclusion to follow-on work, for us or for an affiliated business.
  • Recommend one of our own tools where the evidence favours something else.
  • Take on a question we cannot answer well. We will say so, and point elsewhere if we can.

See the method applied: a sample decision report, a published correctness study, and the decision guides for agent architecture, model replacement and kernel validation.

Is this decision on your desk?

Tell us the decision, the deadline and what evidence would settle it. We reply with whether we can help, and if so the smallest investigation that would.