A long agent task succeeds only if its steps survive in sequence, so a few points of per-step reliability you cannot see on a sprint benchmark decide the whole run. Set two models and your task length. Then get the routing call: cost per finished task, once you price in the frontier's premium.
Free. New tools and the essays behind them, with every source shown. The button opens Substack; nothing you entered here goes with it.
Subscribe freeThe exact exponent is a diagnostic, not a law of physics: real agents recover, which simply raises the effective per-step reliability you set above, and a finished run can retry rather than restart whole. The point survives the caveat. Reliability gaps that round to nothing on a short task go nonlinear once the work has to survive many handoffs.