The rack
Instruments
The publication’s working rule is instruments over theory. These are the runnable ones — free, no login, each built alongside an essay. Every instrument runs in your browser and reads only what you give it.
- The Shape Test Is your growth curve compounding or just accumulating? Drag through your own series and watch the verdict arrive, later than you expect. You bring Your own growth series, dragged point by point onto the chart. It returns The verdict, compounding or straight, and how late in the series it becomes callable. Built for Its essay is at press.
- The Structure Spotter Name a belief you hold about your own numbers. Three tests route it to the artefact that most likely produced it. You bring The load-bearing beliefs your work rests on, one per line. It returns A verdict per belief, the structure underneath it, and the fix a field that met it first already wrote. Built for Your AI Stack Has Three Bugs Other Fields Already Fixed
- The Potemkin Map Score what your system only appears to do. Five checks place each claim on the live map, with the move that closes the gap. You bring Your automated loops, scored on two questions each. It returns Each loop placed live on the fakeable × terminal map, and the move for any loop in the dangerous corner. Built for Your AI Looks Best Where You Can Check It Least
- The Metric Validity Audit Pick your metric. Get a bespoke audit: how it lies, the blind spot you missed, and what to do. You bring The metric you steer by, typed as you actually use it. It returns A bespoke audit: the specific way that number lies, the blind spot it buys, and the check that catches it. Built for The Stable Liar
- The Marathon Calculator Per-step reliability compounds over a long agent run. See the finish-rate gap and the cost per finished task. You bring Two per-step reliabilities, the length of the job, and the price gap between the models. It returns The finish-rate gap once steps compound, and the cost of a finished task with retries priced in. Built for Your Benchmark Measures a Sprint. Your Agent Runs a Marathon.
- The Multi-Agent Decision Most multi-agent systems are an org chart drawn in software. Four questions per task; a ranked call on whether a flat loop beats the fleet. You bring The tasks you are splitting across agents, one line each. It returns A per-task verdict, ranked: where a flat loop wins, where a fleet earns its overhead, and the four answers that decided it. Built for Your Multi-Agent System Is an Org Chart
- The Two-Rate Diagnostic Name the AI layer your advantage runs through. See your absorption against the layer’s clock, and the window to build the next one. You bring The layer your advantage runs through, and the pace you actually absorb its gains at. It returns A dated window: when the layer’s clock catches your absorption, and the layer to build next. Built for How Long Until Your AI Edge Stops Paying?
Built to be checked Every instrument is a single readable file. What you type stays on the page; open the source and check.
Download & run
On your own machine
These are not browser instruments. You download the file and run it locally, and it reads only what you point it at. Nothing leaves your machine.
-
The Grounding Pass
Scans a folder of notes and flags every one that links out but reaches no source: a source-orphan, a claim with nothing under it.
Needs Python 3.8+ · zero dependencies
-
Tests Worth Passing
Mutates your code one change at a time and re-runs your suite, so you see which tests actually catch the damage and which just pass.
Needs Python 3 · standard library