Strand · TEST

Proof & Trust

12 essays, April 2026 to August 2026. They test Goodhart Corollary (3), Instruments Over Theory (3), The Targeting Problem (1) and Bottleneck Migration (1). 55 claims across 3 of them are checked against a primary source in the public ledger.

Instruments from this strand

Every essay in this strand

  1. Law A Everyone Got Safer. That's the Problem. Safety has two numbers: how often each system fails, and whether they fail together. The field has spent years driving the first one down while almost no dashboard reports the second. 19 sourced claims
  2. A Green Score Is Not Evidence A groundedness metric scored its best with the evidence removed. The way to tell whether your model is actually using its evidence is to change the evidence and watch what moves in the answer. 12 sourced claims
  3. Law IV Your Robot Coworker Is Still a Pilot Built, shipped, installed, working: four counts, quoted as one. Almost nobody publishes the fourth. 24 sourced claims
  4. Law IV The Most Expensive AI Errors Are Made of True Numbers Two months auditing an AI research agent. Nine ways a true number lies, and the check that catches each.
  5. Law V Your AI Looks Best Where You Can Check It Least When the first failure is terminal, you cannot iterate your way back.
  6. Law I Your Research Agent Cites Sources It Never Read The same trap has killed pricing models and trading desks for decades. One move tells you if your number is next.
  7. Law A Ten Lines of Code Scored 100%. One Agent Broke Eight Benchmarks. Not one task was actually solved, and the same blind spot is sitting in your own dashboard.
  8. Law IV How Reliable Is Your AI Agent? A month running an autonomous agent. Everyone who does comes back having built the same thing: a verifier.
  9. Law A The Stable Liar Every metric you optimise quietly stops measuring what you meant. The dangerous ones never break. They keep reporting green while the thing underneath rots.
  10. What Proves You Can Think? AI did not just make output cheap. It broke the old contract between effort, competence, and trust. The next scarce signal is proof of judgement under conditions where the surface itself can be faked.
  11. Most Verification Is Just Bigger Classification A confidence score is not evidence. If your eval cannot produce a replayable artefact, it will fail the moment the system can respond to being measured.
  12. You're Not Comparing Models. You're Comparing Contracts. Agent benchmarks don't measure models. They measure contracts. Two teams running the same model can publish different scores, and both can be honest.

All strands The instrument rack The claim ledger