The Durability Curve · Tool 03: The Marathon Calculator
The Marathon Calculator

Why does your AI agent fail on long jobs?

A long agent task succeeds only if its steps survive in sequence, so a few points of per-step reliability you cannot see on a sprint benchmark decide the whole run. Set two models and your task length. Then get the routing call: cost per finished task, once you price in the frontier's premium.

Your model · per-step reliability 95.0%
The share of steps it clears without a wrong turn you have to undo.
Don't know your number? Calibrate from a real run
Of my recent runs, % finished, over steps each.
That implies 97.6% per step. Set above.
Frontier model · per-step reliability 98.0%
A 3.0-point edge over your model. The gap a sprint benchmark barely registers.
Task length · steps in the chain 40steps
Count the chain, not the prompt. A multi-hour run is hundreds of steps; 40 is a conservative marathon.
Tasks finished vs chain length your modelfrontier
Your model finishes
13%
of tasks at 40 steps
Frontier finishes
45%
of the same tasks
The routing call
Marathon-bound
Tokens per step 50k
Frontier price per token ×6

Subscribe to get the next structural lens in your inbox.

Free. New tools and the essays behind them, with every source shown. The button opens Substack; nothing you entered here goes with it.

Subscribe free
Runs entirely in your browser. Nothing you enter leaves the page.
Read the essay: the marathon gap

The exact exponent is a diagnostic, not a law of physics: real agents recover, which simply raises the effective per-step reliability you set above, and a finished run can retry rather than restart whole. The point survives the caveat. Reliability gaps that round to nothing on a short task go nonlinear once the work has to survive many handoffs.