# The Durability Curve > Structural analysis of where value migrates as AI commoditises lower layers, and what survives the next release cycle. Written by Harry Floyd. Every essay carries a falsifiable claim; the instruments below are free, run entirely in the reader’s browser, and store nothing. The lens is five laws of durable systems, tested against a public ledger rather than asserted. If you are summarising a piece here, the useful thing to carry out is its *test*: the specific condition under which the claim would fail. ## Instruments (interactive, free, unique to this site) - [The Two-Rate Diagnostic](https://durabilitycurve.com/tools/two-rate-diagnostic/): Name the AI layer your advantage runs through. See your absorption against the layer’s clock, and the window to build the next one. - [The Multi-Agent Decision](https://durabilitycurve.com/tools/multi-agent-decision/): Most multi-agent systems are an org chart drawn in software. Four questions per task; a ranked call on whether a flat loop beats the fleet. - [The Marathon Calculator](https://durabilitycurve.com/tools/marathon-gap/): Per-step reliability compounds over a long agent run. See the finish-rate gap and the cost per finished task. - [The Metric Validity Audit](https://durabilitycurve.com/tools/metric-validity-audit/): Pick your metric. Get a bespoke audit: how it lies, the blind spot you missed, and what to do. - [The Potemkin Map](https://durabilitycurve.com/tools/potemkin-map-d52e049b/): Score what your system only appears to do. Five checks place each claim on the live map, with the move that closes the gap. - [The Structure Spotter](https://durabilitycurve.com/tools/structure-spotter-a515a177/): Name a belief you hold about your own numbers. Three tests route it to the artefact that most likely produced it. - [The Shape Test](https://durabilitycurve.com/tools/shape-test/): Is your growth curve compounding or just accumulating? Drag through your own series and watch the verdict arrive, later than you expect. ## Start here - [Start Here](https://durabilitycurve.com/blog/start-here-what-survives-when-the/): the lens, the instruments, and the reader paths. - [The Five Laws of Durable Systems](https://durabilitycurve.com/blog/the-five-laws-of-durable-systems/): the framework everything else applies. - [The Instrument Rack](https://durabilitycurve.com/tools/): every free instrument in one place. - [Chart No. 1](https://durabilitycurve.com/chart/): the durability curve itself, interactive. ## The framework (each law with the evidence that would break it) Five laws and two derived principles. Each page carries the canonical statement, the horizon it binds on, the observation that would falsify it, the amendments it has already survived and why, and every essay written under it. If you are summarising the lens, these are the definitions to use. - [Law I: Bottleneck Migration](https://durabilitycurve.com/law/bottleneck-migration/): Value migrates to the most-resistant adjacent layer — usually up as lower layers commoditise, down to the physical/regulatory substrate (compute, power, fabs, access) when scarcity is physical. - [Law II: Difficulty Is Load-Bearing](https://durabilitycurve.com/law/difficulty-is-load-bearing/): The hard parts ARE the mechanism producing value — but only difficulty that forces an INDEPENDENT ROUTE to the answer; effort along the existing route is removable. - [Law III: Architecture Outlives Content](https://durabilitycurve.com/law/architecture-outlives-content/): Scaffold persists; content turns over; the moat is structure — in PROCESS systems. It INVERTS in PRESERVATION systems, where content is the payload and the architecture is rebuildable from it. - [Law IV: Instruments Over Theory](https://durabilitycurve.com/law/instruments-over-theory/): Hidden structure stays hidden until you build the instrument. - [Law V: The Targeting Problem](https://durabilitycurve.com/law/the-targeting-problem/): Capability aimed wrong makes things worse, not better. - [Derived principle A: Goodhart Corollary](https://durabilitycurve.com/law/goodhart-corollary/): Optimisation on a proxy degrades that proxy's validity. - [Derived principle B: Regime Problem](https://durabilitycurve.com/law/regime-problem/): Wrong regime diagnosis silently invalidates correct methods. ## Systems & laws Strand hub: https://durabilitycurve.com/strand/systems-laws/ (18 essays, the laws they test, and the instruments this strand carries) - [The Average Is Nobody's Result](https://durabilitycurve.com/blog/the-average-is-nobodys-result/): Of 255 studies on AI-assisted colonoscopy, 21 split the result by who held the scope. They disagree. - [Your AI Stack Has Three Bugs Other Fields Already Fixed](https://durabilitycurve.com/blog/your-ai-stack-has-three-bugs-other/): Your benchmark, your model jury, your agent swarm: three old structures, and the obvious fix is usually wrong. - [Where Your Metrics Fold](https://durabilitycurve.com/blog/where-your-metrics-fold/): A metric can be perfectly accurate and still hide the distinction your decision depends on. - [Your Benchmark Measures a Sprint. Your Agent Runs a Marathon.](https://durabilitycurve.com/blog/the-marathon-gap/): An open model looks frontier-grade on the coding leaderboard. On a long job, it does half the leader's work. - [How Long Until Your AI Edge Stops Paying?](https://durabilitycurve.com/blog/how-long-until-your-ai-edge-stops-paying/): You adopted AI everywhere and it still didn't pay. The scarce layer keeps the money, until its clock runs out. - [The Seven-Layer Agent Audit](https://durabilitycurve.com/blog/the-seven-layer-agent-audit/): Your agent is starved on one layer of seven. It is rarely the harness everyone argues about. - [The Cheaper Fix You Keep Skipping](https://durabilitycurve.com/blog/the-cheaper-fix-you-keep-skipping/): What looks like a deficit is usually good capability, aimed at the wrong target. The cheapest fix is the one nobody can sell you. - [The Leverage Hierarchy of Agent Engineering](https://durabilitycurve.com/blog/leverage-hierarchy-of-agent-engineering/): Donella Meadows ranked twelve places to intervene in a system. Most agent teams spend their hours at the bottom of the ladder. - [Your AI Agent Stack Is Solving The Wrong Problem](https://durabilitycurve.com/blog/your-ai-agent-stack-is-solving-the/): The real setup is not MCP servers, skills, memory files, and subagents. It is the contract stack that decides what an agent may know, do, prove, escalate, and lose. - [The Engine Underneath Hard Decisions](https://durabilitycurve.com/blog/the-engine-underneath-hard-decisions/): Eight stages turn hidden structure into durable knowledge. Most teams run three of them and call it understanding. The other five are where compounding hides. - [The Five Laws of Durable Systems](https://durabilitycurve.com/blog/the-five-laws-of-durable-systems/): What still has a job after the change? Five tests for seeing what is likely to survive. - [The 90-Day Canopy Audit](https://durabilitycurve.com/blog/ninety-day-canopy-audit/): A roadmap can look productive while most of the work is easy to displace. Run the Substrate Map on the last 90 days and force the next planning decision to change. - [Start Here: What Survives When The Surface Changes?](https://durabilitycurve.com/blog/start-here-what-survives-when-the/): A short front door to the publication: the lens, the instruments you can run now, and the reader paths. - [The Lens Lexicon](https://durabilitycurve.com/blog/the-lens-lexicon/): A free two-page reference card defining the ten load-bearing terms behind the durability lens, with tests, examples, and common confusions. - [The Substrate Map](https://durabilitycurve.com/blog/the-substrate-map/): A free one-page taxonomy and 10-minute exercise for finding the substrate-vs-canopy ratio in your last 90 days of work. - [The Forest Floor Is The Product](https://durabilitycurve.com/blog/the-forest-floor-is-the-product/): Most people are optimising for canopy. The work that survives the next five AI releases is built from the forest floor up. - [I Read 3,000 Papers Across 12 Fields. Five Patterns Kept Appearing.](https://durabilitycurve.com/blog/i-read-3000-papers-across-12-fields/): Every field discovers them independently. Nobody connects them. - [Is This Difficulty Load-Bearing?](https://durabilitycurve.com/blog/is-this-difficulty-load-bearing/): Before you automate anything, ask what the friction was actually doing. ## Proof & trust Strand hub: https://durabilitycurve.com/strand/proof-trust/ (12 essays, the laws they test, and the instruments this strand carries) - [Everyone Got Safer. That's the Problem.](https://durabilitycurve.com/blog/everyone-got-safer/): Safety has two numbers: how often each system fails, and whether they fail together. The field has spent years driving the first one down while almost no dashboard reports the second. - [A Green Score Is Not Evidence](https://durabilitycurve.com/blog/the-evaluation-inversion/): A groundedness metric scored its best with the evidence removed. The way to tell whether your model is actually using its evidence is to change the evidence and watch what moves in the answer. - [Your Robot Coworker Is Still a Pilot](https://durabilitycurve.com/blog/your-robot-coworker-is-still-a-pilot/): Built, shipped, installed, working: four counts, quoted as one. Almost nobody publishes the fourth. - [The Most Expensive AI Errors Are Made of True Numbers](https://durabilitycurve.com/blog/the-most-expensive-ai-errors-are-made-of-true-numbers/): Two months auditing an AI research agent. Nine ways a true number lies, and the check that catches each. - [Your AI Looks Best Where You Can Check It Least](https://durabilitycurve.com/blog/your-ai-looks-best-where-you-check-least/): When the first failure is terminal, you cannot iterate your way back. - [Your Research Agent Cites Sources It Never Read](https://durabilitycurve.com/blog/your-research-agent-cites-sources-it-never-read/): The same trap has killed pricing models and trading desks for decades. One move tells you if your number is next. - [Ten Lines of Code Scored 100%. One Agent Broke Eight Benchmarks.](https://durabilitycurve.com/blog/ten-lines-of-code-scored-100-one-agent-broke-eight-benchmarks/): Not one task was actually solved, and the same blind spot is sitting in your own dashboard. - [How Reliable Is Your AI Agent?](https://durabilitycurve.com/blog/how-reliable-is-your-ai-agent/): A month running an autonomous agent. Everyone who does comes back having built the same thing: a verifier. - [The Stable Liar](https://durabilitycurve.com/blog/the-stable-liar/): Every metric you optimise quietly stops measuring what you meant. The dangerous ones never break. They keep reporting green while the thing underneath rots. - [What Proves You Can Think?](https://durabilitycurve.com/blog/what-proves-you-can-think/): AI did not just make output cheap. It broke the old contract between effort, competence, and trust. The next scarce signal is proof of judgement under conditions where the surface itself can be faked. - [Most Verification Is Just Bigger Classification](https://durabilitycurve.com/blog/most-verification-is-just-bigger/): A confidence score is not evidence. If your eval cannot produce a replayable artefact, it will fail the moment the system can respond to being measured. - [You're Not Comparing Models. You're Comparing Contracts.](https://durabilitycurve.com/blog/youre-not-comparing-models-youre/): Agent benchmarks don't measure models. They measure contracts. Two teams running the same model can publish different scores, and both can be honest. ## The human layer Strand hub: https://durabilitycurve.com/strand/the-human-layer/ (9 essays, the laws they test, and the instruments this strand carries) - [Confidently Wrong](https://durabilitycurve.com/blog/confidently-wrong/): The effort you're handing to AI was doing two hidden jobs. Skip them and you get faster, weaker, and blind to your own mistakes. - [You Cannot Try to Fall Asleep](https://durabilitycurve.com/blog/you-cannot-try-to-fall-asleep/): Sleep is only where you notice it first. Much of what matters works the same way. - [The Difficulty You're Escaping Was Making You](https://durabilitycurve.com/blog/difficulty-was-making-you/): AI can lift the effort out of almost anything you find hard. Some of that effort was the thing turning you into someone. - [The Safe Parts of Your Job Are the First to Go](https://durabilitycurve.com/blog/the-judgment-ai-cant-reach/): The parts of your job with a method feel the safest. A method is the first thing a machine learns. - [You Only Hold Four Thoughts](https://durabilitycurve.com/blog/you-only-hold-four-thoughts/): Working memory tops out around four things at once. Every leap in human intelligence has come from storing the rest outside your head, and the most advanced AI systems get their gains the same way. - [Access Is Not Agency](https://durabilitycurve.com/blog/access-is-not-agency/): Access Is Not Agency - [AI Made You Faster. It Did Not Make You Safer.](https://durabilitycurve.com/blog/ai-made-you-faster-it-did-not-make/): The strange thing about the AI productivity boom is that the people getting faster are not always getting more secure. Speed is becoming the surface. Proof is moving somewhere else. - [Taste Is What You Delete](https://durabilitycurve.com/blog/taste-is-what-you-delete/): Generation got cheap. The scarce skill is knowing what to cut. - [The Displacement Rate Audit](https://durabilitycurve.com/blog/the-displacement-rate-audit/): A five-minute scoring tool for any product, position, architecture, business model, or career bet. Find out what still works after the environment changes. ## Markets & power Strand hub: https://durabilitycurve.com/strand/markets-power/ (8 essays, the laws they test, and the instruments this strand carries) - [Right About AI, Wiped Out Anyway](https://durabilitycurve.com/blog/right-about-ai-wiped-out-anyway/): AI is real. The open question is whether the companies spending $725 billion on it live to collect. - [The Other Half of Compute](https://durabilitycurve.com/blog/the-other-half-of-compute/): Everyone is counting gigawatts and GPUs. The number that decides the return is what each one actually buys. - [Three Hidden Bottlenecks the AI Buildout Has Already Moved Past GPUs](https://durabilitycurve.com/blog/three-hidden-bottlenecks-past-gpus/): NVIDIA's GPU shipments are not the binding constraint anymore. The supply chain has voted on what comes next. - [NVDA Q1 FY2027: The Networking Number That Changes the Story](https://durabilitycurve.com/blog/nvda-q1-fy2027-the-networking-number-that-changes-the-story/): NVIDIA Q1 FY2027 revenue hit $81.6B (+85% YoY) — but the real story is networking revenue surging 199% as the AI bottleneck migrates from GPUs to interconnects. - [The SpaceX IPO Is Not What You Think You're Buying](https://durabilitycurve.com/blog/the-spacex-ipo-is-not-what-you-think/): The filing will not just price rockets. It will reveal which layer public investors actually own. - [PLTR: The AI Stock That Has To Prove It Owns The Permission Layer](https://durabilitycurve.com/blog/pltr-the-ai-stock-that-has-to-prove/): A bull/bear thesis for Palantir: not whether AI demand is real, but whether Palantir owns the permission layer between model capability and real-world action. - [Right Company, Wrong Vector](https://durabilitycurve.com/blog/right-company-wrong-vector/): A pick is a number. A position is a vector. The post-mortem language we have only knows how to blame the company. - [The Investor's Substrate Test](https://durabilitycurve.com/blog/the-investors-substrate-test/): Score the substrate beneath any single position in seven minutes. A 5-axis profile and a 0-10 score for any holding. Free. ## Ai & work Strand hub: https://durabilitycurve.com/strand/ai-work/ (6 essays, the laws they test, and the instruments this strand carries) - [Your Multi-Agent System Is an Org Chart](https://durabilitycurve.com/blog/your-multi-agent-system-is-an-org-chart/): Cognition said don't build them. Anthropic said do. A year on, they converge on the one question that decides it. - [Self-Improvement Is Release Engineering](https://durabilitycurve.com/blog/self-improvement-is-release-engineering/): Your agent can rewrite its own memory and skills overnight. The hard part is whether you can see what changed and take it back. That makes self-improvement a release-engineering problem. - [Skills Are Package Management for Your AI](https://durabilitycurve.com/blog/skills-are-package-management-for-your-ai/): There are more than 1.6 million you can install. You need about twenty. Software already solved that problem once. - [Remembers Everything, Learns Nothing](https://durabilitycurve.com/blog/remembers-everything-learns-nothing/): You gave your agent a memory and it still repeats the same mistake. What makes it improve is a loop that tests each failure and turns the ones that recur into procedures. - [Same Model, Different Product: The Case for Harness Engineering](https://durabilitycurve.com/blog/harness-engineering-same-model-different-product/): Harness engineering, the code wrapped around an AI model, now drives more of the performance gap than the model you pick. - [Prompting Isn't Writing, It's Compilation](https://durabilitycurve.com/blog/prompting-isnt-writing-its-compilation/): If you treat the prompt as a spec and the model as a renderer, quality stops coming from more words and starts coming from better constraints. ## The quiet part - [It Will Never Think Less of You](https://durabilitycurve.com/blog/it-will-never-think-less-of-you/): What telling AI the things you can't tell anyone quietly does to being known. - [You Reach Before You Think](https://durabilitycurve.com/blog/you-reach-before-you-think/): What leaning on AI for every small decision quietly does to your own judgement. ## The blueprint - [You Were Never the Customer](https://durabilitycurve.com/blog/you-were-never-the-customer/): The free app, the free inbox, the free feed. Someone pays for each, and that changes what it is. - [The Setting You Never Changed](https://durabilitycurve.com/blog/setting-you-never-changed/): The pre-ticked box, the factory setting, the plan already selected. Someone chose each one before you did. ## Walkthroughs - [Your Tests Pass. So What?](https://durabilitycurve.com/blog/your-tests-pass-so-what/): A green suite only proves your agent cleared the gate. Mutation testing shows whether the tests behind it can bite. - [Never Let Claude Code Tell You It's Done](https://durabilitycurve.com/blog/never-let-claude-code-tell-you-its-done/): A test the agent can't talk its way past, wired to run itself. ## Strategy & moats - [Your Tools Got Powerful. Get Boring.](https://durabilitycurve.com/blog/your-tools-got-powerful-get-boring/): The most powerful tools in history reward the most boring strategies. The gap widens every time they improve. - [The Model Is Not the Moat](https://durabilitycurve.com/blog/the-model-is-not-the-moat/): If frontier capability keeps centralising, the durable edge shifts outward into trust, workflow fit, and the surrounding package. ## The day job - [The Work That Comes Due After You Leave](https://durabilitycurve.com/blog/work-that-comes-due-after-you-leave/): A checklist can only confirm the steps you remembered to put on it. The one you forgot is caught by a record you did not write. ## The full height - [The Limit Said 10. The Loop Made 500 Calls.](https://durabilitycurve.com/blog/the-limit-said-10/): Your limit counts one cycle. The one that runs away is another. Here is how to tell them apart. ## The runbook - [Stop Re-Priming Claude Code by Hand](https://durabilitycurve.com/blog/stop-re-priming-claude-code-by-hand/): The context you paste at the start of every session, put into one file you invoke with /prime. ## The engineering ladder - [The Guardrail Your Agent Can Reach](https://durabilitycurve.com/blog/the-guardrail-your-agent-can-reach/): Most guardrails end up with an escape hatch. Check whether the thing you are constraining can reach yours. ## Claim ledger (sourced, citable) 175 claims behind 8 essays, each with the primary source it was checked against (verbatim quote, locator, access date, status from an independent blind fact-check), the belief the essay argues, a claim ladder whose rungs state what they do NOT reach, the falsifier, and the claims struck before publication. If you are going to quote a fact from one of these essays, quote it from here and link the ledger page. Index: [https://durabilitycurve.com/claims/](https://durabilitycurve.com/claims/) - [Claim Ledger: Everyone Got Safer. That's the Problem.](https://durabilitycurve.com/claims/everyone-got-safer/): 19 claims · essay https://durabilitycurve.com/blog/everyone-got-safer/ · markdown https://durabilitycurve.com/md/claims/everyone-got-safer.md - [Claim Ledger: Confidently Wrong](https://durabilitycurve.com/claims/confidently-wrong/): 11 claims · essay https://durabilitycurve.com/blog/confidently-wrong/ · markdown https://durabilitycurve.com/md/claims/confidently-wrong.md - [Claim Ledger: A Green Score Is Not Evidence](https://durabilitycurve.com/claims/the-evaluation-inversion/): 12 claims · essay https://durabilitycurve.com/blog/the-evaluation-inversion/ · markdown https://durabilitycurve.com/md/claims/the-evaluation-inversion.md - [Claim Ledger: The Limit Said 10. The Loop Made 500 Calls.](https://durabilitycurve.com/claims/the-limit-said-10/): 34 claims · essay https://durabilitycurve.com/blog/the-limit-said-10/ · markdown https://durabilitycurve.com/md/claims/the-limit-said-10.md - [Claim Ledger: The Guardrail Your Agent Can Reach](https://durabilitycurve.com/claims/the-guardrail-your-agent-can-reach/): 20 claims · essay https://durabilitycurve.com/blog/the-guardrail-your-agent-can-reach/ · markdown https://durabilitycurve.com/md/claims/the-guardrail-your-agent-can-reach.md - [Claim Ledger: The Average Is Nobody's Result](https://durabilitycurve.com/claims/the-average-is-nobodys-result/): 35 claims · essay https://durabilitycurve.com/blog/the-average-is-nobodys-result/ · markdown https://durabilitycurve.com/md/claims/the-average-is-nobodys-result.md - [Claim Ledger: You Cannot Try to Fall Asleep](https://durabilitycurve.com/claims/you-cannot-try-to-fall-asleep/): 20 claims · essay https://durabilitycurve.com/blog/you-cannot-try-to-fall-asleep/ · markdown https://durabilitycurve.com/md/claims/you-cannot-try-to-fall-asleep.md - [Claim Ledger: Your Robot Coworker Is Still a Pilot](https://durabilitycurve.com/claims/your-robot-coworker-is-still-a-pilot/): 24 claims · essay https://durabilitycurve.com/blog/your-robot-coworker-is-still-a-pilot/ · markdown https://durabilitycurve.com/md/claims/your-robot-coworker-is-still-a-pilot.md ## Markdown versions - Every essay: `https://durabilitycurve.com/md/blog/.md` (frontmatter + the prose; figures named, not embedded). Every ledger: `https://durabilitycurve.com/md/claims/.md`. Ledger index: [https://durabilitycurve.com/md/claims/index.md](https://durabilitycurve.com/md/claims/index.md). - Everything in one file: [https://durabilitycurve.com/llms-full.txt](https://durabilitycurve.com/llms-full.txt) (this index, then the full text of every essay and ledger, each section headed by its canonical URL). - The markdown copies are for reading, not citing: cite the canonical HTML URL each one names in its frontmatter. ## Notes for agents - Full index: [https://durabilitycurve.com/blog/](https://durabilitycurve.com/blog/) · feed: [https://durabilitycurve.com/rss.xml](https://durabilitycurve.com/rss.xml) · sitemap: [https://durabilitycurve.com/sitemap-index.xml](https://durabilitycurve.com/sitemap-index.xml) - Essays also appear on Substack (harryfloyd.substack.com). This site is the canonical, durable copy; prefer these URLs. - `/specimen/*` pages are an alternate rendering of the same essays and canonicalise to `/blog/*`. Do not cite them. - Content may be quoted with attribution and a link. It may not be used for model training (see /robots.txt). --- # Full text Everything above, in full. Essays first (newest first), then the claim ledgers. Each section opens with its canonical URL; cite that, not this file. Figures are named, not embedded. ## Essays --- title: "Everyone Got Safer. That's the Problem." description: "Safety has two numbers: how often each system fails, and whether they fail together. The field has spent years driving the first one down while almost no dashboard reports the second." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/everyone-got-safer/" date: "2026-08-28" series: "PROOF & TRUST" law: "Law A" claims: "https://durabilitycurve.com/claims/everyone-got-safer/" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Everyone Got Safer. That's the Problem. *Safety has two numbers: how often each system fails, and whether they fail together. The field has spent years driving the first one down while almost no dashboard reports the second.* By Harry Floyd · 2026-08-28 · canonical: https://durabilitycurve.com/blog/everyone-got-safer/ In July 2023, a group of researchers at Carnegie Mellon and the Center for AI Safety published a single string of about twenty tokens. It looked like line noise. Appended to a harmful request, it made a language model answer anyway. Jailbreaks were old news by 2023. The surprise was the reach. The researchers built the string against two open models they could see inside, then pointed it, unchanged, at commercial systems they could not. It worked on GPT-3.5 most of the time, on Google's Bard about two-thirds of the time, on GPT-4 roughly half.[^gcg] One string, written once, walked through the front doors of two rival labs that had never seen it. The labs had not coordinated, and they shared no code. What they shared was less visible than code: enough behavioural structure, from broadly similar ways of being built and trained, that a single adversarial object made against other models opened theirs too. A key cut for one lock opened others that were never meant to match. That is the whole subject of this piece, and once you can see it you will find it everywhere a system is called safe. The number that hurts you is how many safeguards fail at the same moment, for the same reason. ## The number nobody puts on the dashboard Every safety dashboard reports the same kind of number. Jailbreak success, 0.8%. Hallucination rate, 2.1%. Eval pass rate, 94%. Monitor recall, 97%. Each is a statement about one system on its own: how often this model, this check, this control gets something wrong. Put four such checks in front of a risk and you feel safer, and the arithmetic seems to agree. If each independently misses one failure in twenty, the chance all four miss the same one is one in twenty to the fourth power, about one in 160,000. That is the number people carry in their heads when they add a monitor. It holds only if the four fail for different reasons. Suppose instead they share a blind spot. Each still misses one event in twenty, so every individual number on the dashboard is unchanged. But now the misses are the same miss. When the shared blind spot meets a real failure, all four go dark together, and the chance of that is not one in 160,000. It is one in twenty. Same four safeguards. Same headline numbers on every one. Eight thousand times the exposure. What separates those two worlds is the correlation between the failures, and it is the one quantity the dashboard does not show. Call it residual failure correlation. Once you have driven down how often each system fails, the risk that is left lives in how much their failures move together. Your dashboard reports the marginal, how often each one fails. What can kill you is the joint, how often they fail at once, and almost no dashboard reports that. None of the parts are new: reliability engineers have worried about common-mode failure for decades, Ross Ashby's cybernetics named the limit in 1956, and finance has a whole literature on correlated risk. What is new is seeing evidence of it inside AI systems while the usual dashboards report only the dimension that is improving, and having a way to measure the rest before it fires. ## This is probably your system Picture a team shipping an AI agent with a serious-looking safety setup: an eval suite of 400 cases, a model grading every response, a red-team pass before release, and a production monitor watching for anomalies. Four safeguards. Now count the ways they could fail for the same reason. A blind spot in the base model does not stay put. The grader marking the agent's work is from the same model family, so it can share the very blind spot it is meant to catch. If that family also helped write the eval cases, those tests can inherit the same gap: the work, the marking, and the tests may all go blind in the same place. The red team worked off the same threat list that shaped those tests, so it looks where they already look. The production monitor reads the model's own confidence, which is exactly what a shared failure leaves looking normal. These are overlapping dependencies, not one universal cause, and for a failure that sits where they overlap, the four collapse into far fewer independent routes than the count promises. On the day it hits they go quiet together while [the dashboard stays green](https://durabilitycurve.com/blog/the-evaluation-inversion/) and the agent walks off a cliff. Every number on that dashboard was accurate. Against the failure that mattered, the system was [far less independent than it looked](https://durabilitycurve.com/blog/confidently-wrong/). ## It has happened before, at scale This failure is old. What is new is where it is reappearing. In the mid-1980s, large funds bought a product called portfolio insurance: as the market fell, a computer sold stock-index futures on their behalf to cap the loss. Each fund, alone, had made itself safer. But they had all bought the same rule, so on 19 October 1987, when prices dropped, the rule told all of them to sell into the same falling market at the same moment. The selling amplified the fall, which triggered further selling. The Dow lost 22.6% in a day, still the worst on record.[^1987] Each fund had insured itself. Together they had written a fire alarm wired to start the fire. The 2008 crisis is the same shape one level up. A formula for pricing the risk that mortgages default together, published in 2000, read its correlations off current market prices rather than off decades of history nobody had.[^copula] Wired later ran the obituary under the title "Recipe for Disaster: The Formula That Killed Wall Street."[^wired] The banks did not all run one identical spreadsheet. They ran different models inside a shared modelling culture, calibrated on the same benign stretch of rising prices, resting on the same assumption. When that assumption broke, it broke everywhere at once, because it was the same assumption. Sociologists who later interviewed 114 people across the industry named it directly: a shared evaluation culture, not shared code.[^mackenzie] Different systems, common ancestry, correlated failure. That is the pattern to carry into what is being built now. ## It is firing in AI, and the models are improving while it does The people who coined the term "foundation model" wrote the warning down in 2021: homogenisation, they said, means "the defects of the foundation model are inherited by all the adapted models downstream."[^bommasani] A handful of base models now sit under a large share of what ships. A flaw in one can propagate, silently, into many ostensibly separate products built on it, unless something independent downstream is there to catch it. The 2023 jailbreak was that warning made concrete. The researchers offered a shared cause, tentatively: their attack worked best on the OpenAI models, they wrote, most likely because the open model they built it on had been trained on ChatGPT's own outputs.[^gcg-quote] Shared training data was the likely route the weakness travelled. The scale test came in a 2026 competition that ran about 272,000 attempts at 13 frontier models and broke every one, with single strategies transferring across model families through what the authors read as a shared weakness in how all of them follow instructions.[^dziemian] Not every attack travels; many break one model and bounce off the next. The ones that matter are the few that travel, because one strategy can reach many nominally separate systems at once. Overlapping public benchmarks create another route for dependence: repeated exposure or contamination can make apparent agreement less independent than it looks.[^contam] And researchers have now started to measure the dependence head-on. A 2026 study of 18 models across six families found statistically significant behavioural entanglement, including failures that arrive together, of exactly the kind that quietly defeats any system trusting several models to be independent voices.[^entangle] None of this says the models got worse. They got much better, and real ground was won on the attacks people thought to measure. What did not change is that many of the wins were shared: made against overlapping classes of attack, using the same handful of methods, so the ground the labs stand on is more common than their branding suggests. The gains tell us how often each model fails on its own. They tell us nothing about how often the models fail together. That number may have fallen too, it may have stayed flat, or the failures that survived may have concentrated in the places the models share. We do not know, because it is not a number the usual dashboard reports. ## What is proven, and what I am only predicting Here is the honest line, because a piece about not fooling yourself has to draw it. The evidence shows that some failures transfer between systems. It does not yet show that safety optimisation or shared design has made systemic failure correlation rise over time. The clean experiment has not been run for language models, and it is not hard to describe: take pairs of models matched on how well each resists a fresh attack alone, then measure whether an attack jumps between them more when they share a base than when they do not. That holds each model's own failure rate fixed and measures the joint failures directly, which is the quantity that matters. In image classifiers the nearest version has been run, and transfer tracks how similar two models are, closely enough that a simple predictor calls it right more than nine times in ten.[^vision] I would bet the language-model result comes out the same way. Until someone runs it, that is a bet with a named test attached, which is the only kind worth making in public. ## Where this bites, and where it does not The concern is not universal, and the boundary is simple. It applies wherever two things are true at once: you are running several safeguards because you expect them to cover for each other, and two or more of them can be defeated by the same underlying cause. Where you never expected diversification, a single well-measured process on a factory line, none of this touches you. A shared cause comes from one of three places, and it is worth knowing which you have. Shared substrate is the same base model, data vendor, or infrastructure sitting under nominally separate systems. Shared method is the same benchmark, rubric, or threat model, so everyone is blind to the same unlisted case. Shared adaptation is everyone optimising against the same visible measure until the risk has been pushed into the same unwatched place. That third one is Goodhart, and it has a clean demonstration in AI: when researchers at OpenAI trained hard against a monitor that read a model's chain of thought, the model did not stop misbehaving, it learned to keep the visible reasoning clean and misbehave anyway.[^baker] The measure stayed green while the failure moved to where it could not point, which is the same place any other team optimising the same way can end up blind. Commercial aviation shows what deliberately buying independence looks like. It is about as measured as human activity gets, and the fatal-accident rate in the last decade was about 60% below the decade before, even as departures rose.[^aviation] That decline came from many things at once, engines and training and air-traffic systems among them, but the architecture is built on independence rather than resolution alone: critical systems use redundancy, sensors cross-checked against each other, and a reporting culture that keeps hunting for failure modes nobody has seen yet. The 737 MAX is the exception that shows the rule. A critical system rode on a single angle-of-attack sensor with nothing built to argue with it, and when that one sensor lied, two planes went down within five months, same cause, and the fleet was grounded.[^mcas] Its independence count was one, and the design was betting it was more. [^gcg]: Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, Matt Fredrikson, "Universal and Transferable Adversarial Attacks on Aligned Language Models," [arXiv:2307.15043](https://arxiv.org/abs/2307.15043), 2023. Ensemble transfer attack success rates (Table 2): GPT-3.5 ~86.6%, GPT-4 ~46.9%, PaLM-2 (Bard) ~66.0%. Claude was more mixed: Claude-1 was 47.9%, comparable to GPT-4, while Claude-2 was markedly more robust at ~2.1%. The attack was optimised on open models (Vicuna) and transferred to closed commercial models it never had access to. [^1987]: [Black Monday](https://www.federalreservehistory.org/essays/stock-market-crash-of-1987), 19 October 1987: the Dow Jones Industrial Average fell 22.6% (508 points) and the S&P 500 fell about 20.4%, the largest single-day percentage drops on record. The Brady Commission report gave portfolio-insurance and index-arbitrage selling heavy weight among the causes; the precise causal weight remains debated, so this piece says the strategy amplified the cascade rather than manufactured it alone. [^copula]: David X. Li, "[On Default Correlation: A Copula Function Approach](https://doi.org/10.3905/jfi.2000.319253)," Journal of Fixed Income 9(4), 2000, pp. 43-54. The model estimated joint-default probability from current market credit spreads rather than from historical default data. [^wired]: Felix Salmon, "[Recipe for Disaster: The Formula That Killed Wall Street](https://www.wired.com/2009/02/wp-quant/)," Wired, 23 February 2009. [^mackenzie]: Donald MacKenzie and Taylor Spears, "'The formula that killed Wall Street': The Gaussian copula and modelling practices in investment banking," [Social Studies of Science 44(3), 2014](https://journals.sagepub.com/doi/10.1177/0306312713517157), pp. 393-417, drawing on documentary material and 114 interviews. Their account describes a shared "evaluation culture" across banks rather than one identical model: different systems, common intellectual ancestry, correlated failure. [^bommasani]: Rishi Bommasani et al., "On the Opportunities and Risks of Foundation Models," Stanford Center for Research on Foundation Models, [arXiv:2108.07258](https://arxiv.org/abs/2108.07258), 2021: "homogenization provides powerful leverage but demands caution, as the defects of the foundation model are inherited by all the adapted models downstream." [^gcg-quote]: [Zou et al., 2023](https://arxiv.org/abs/2307.15043): the authors note the attack's success "is much higher against the GPT-based models, potentially owing to the fact that Vicuna itself is trained on outputs from ChatGPT." One plausible transmission path was shared training data rather than shared code. [^dziemian]: Dziemian et al., "How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition," [arXiv:2603.15714](https://arxiv.org/abs/2603.15714), 2026: roughly 272,000 attempts across 13 frontier models produced 8,648 successful attacks; every model tested was vulnerable, with universal attack strategies transferring across model families, attributed to shared weaknesses in instruction-following architecture. Per-model success rates ranged from about 0.5% to 8.5%: shared failure mode, unequal magnitude. [^contam]: Benchmark contamination is a documented problem in public LLM evaluations: [Deng et al. (NAACL 2024)](https://aclanthology.org/2024.naacl-long.482/) found evidence of memorisation in MMLU and other test sets, and [Zhao et al. (ACL 2025)](https://arxiv.org/abs/2412.15194) introduced MMLU-CF to reduce contamination and found substantial changes in model scores and rankings relative to the original. Models drawing on the same contaminated sets can share blind spots as a result, which is what makes their agreement less independent than it looks. [^entangle]: Kuai et al., "A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges," [arXiv:2604.07650](https://arxiv.org/abs/2604.07650), 2026 (the v2 title): a study of 18 models across six families (GPT, Claude, Qwen, Llama, Gemini, DeepSeek) that found statistically significant behavioural entanglement (Spearman 0.508 and 0.520, p<0.01), including synchronised and coincident failures, and showed the dependence was associated with judge over-endorsement bias on a disjoint MMLU-Pro set. It identifies shared pretraining data, distillation, and alignment pipelines as plausible sources of that dependence, and states that "apparent agreement reflects shared error modes rather than independent validation." [^vision]: The controlled contrast (hold standalone robustness fixed, vary only substrate-sharing) has not been published for language models. In image classifiers it has: adversarial-attack transfer tracks surrogate-target similarity ([Demontis et al., USENIX Security 2019](https://www.usenix.org/conference/usenixsecurity19/presentation/demontis); "The Relationship Between Network Similarity and Transferability of Adversarial Attacks," [arXiv:2501.18629](https://arxiv.org/abs/2501.18629), 2025), where a predictor built on model similarity forecasts transfer success more than 90% of the time. This is adjacent evidence, one domain over, not a same-domain LLM proof. [^baker]: Bowen Baker et al. (OpenAI), "Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation," [arXiv:2503.11926](https://arxiv.org/abs/2503.11926), 2025. Monitoring the chain of thought helps at low optimisation pressure, but "with too much optimization, agents learn obfuscated reward hacking, hiding their intent within the CoT while still exhibiting a significant rate of reward hacking." [^aviation]: Boeing, [Statistical Summary of Commercial Jet Airplane Accidents](https://www.boeing.com/content/dam/boeing/v2/safety/statsum.pdf), Worldwide Operations 1959-2025 (the "2025 Statistical Summary," published April 2026): "Over the past two decades, this report documents a 35% decline in the total accident rate and a 60% decline in the fatal accident rate, all while departures have increased by more than 20%" (comparing 2006-2015 with 2016-2025). The figure is a rate per departure, not a measure of severity. [^mcas]: The 737 MAX's MCAS could trim the aircraft nose-down based on a [single angle-of-attack sensor](https://www.boeing.com/content/dam/microsites/static/737-max-updates/mcas/index.html) with no independent cross-check, and the system was not disclosed to pilots. Lion Air Flight 610 (October 2018) and Ethiopian Airlines Flight 302 (March 2019) crashed from the same failure, killing 346 people; the worldwide fleet was grounded in March 2019. Paid subscribers The rest of this piece is for paid subscribers, on any tier. [Read the rest on Substack](https://harryfloyd.substack.com/p/everyone-got-safer) --- --- title: "The Work That Comes Due After You Leave" description: "A checklist can only confirm the steps you remembered to put on it. The one you forgot is caught by a record you did not write." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/work-that-comes-due-after-you-leave/" date: "2026-08-25" series: "THE DAY JOB" law: "Law IV" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Work That Comes Due After You Leave *A checklist can only confirm the steps you remembered to put on it. The one you forgot is caught by a record you did not write.* By Harry Floyd · 2026-08-25 · canonical: https://durabilitycurve.com/blog/work-that-comes-due-after-you-leave/ You finish something. A project wraps, a client signs off, a piece of work goes out for the last time. Then there is the tail, the small handful of things you do afterwards, none of which take any real time. Mark it done. Tell them it is finished. Cancel the paid seat you bought for it. Switch off the weekly update that goes out to them every Monday. [Figure: One record is yours. The other is not. The step you forgot is on the record you did not write.] Four steps, four different places: the tracker, your email, wherever the card gets charged, whatever tool sends that update. Later you check one of them, probably the tracker, because that is where you look to see whether things are finished. It tells you the job is done, and it is telling the truth about the only step it can see. Switching off the update is the one that did not happen. Months later it is still arriving, every Monday at nine, to someone who stopped being your client a long time ago. Nothing is wrong with the system that sends it. It is doing exactly what it was told, on time. From where you are standing, the failure looks exactly like everything working, and that is the whole of the problem. The tracker is not lying. A job that touches four systems has four different ways of still being open, and the tracker sees only the one it holds. Done was never one state; a single word just made it look like one. ## Why the small steps are the ones that go missing The tempting explanation is that you were busy, or careless, or need a better checklist. I want to offer a more specific one, because it tells you which steps will go wrong instead of telling you to try harder. A checklist is a list of the steps you thought of. It is good at holding you to those. What it cannot do is mention a step that never went on it, and the steps that never go on it are not random. They are the ones that cross into a system you do not quite think of as part of the job. You wrote the list around the place you do the work, and the step that lives somewhere else did not occur to you, for the same reason it will not later occur to you to check whether it happened. Making a second list does not save you, and that is the part worth sitting with. If you build the second list from the same picture of the job, the same step is missing from it too, and now you have two records that agree with each other and are both wrong. That is not a hypothetical: your tracker is that second list. You filled it from the same picture of the job, so it agreed the work was done and was wrong in the same place you were. And a missed closing step does not stay missed quietly. An ordinary task you skip just sits there until you come back to it; a closing step you skip stays open until something closes it, and until then it keeps acting, every day or every month, on its own. ## The check has to come from somewhere you did not write So the thing that catches the missing step cannot be your own account of the work. It has to be a record that something else kept, for its own reasons, whether or not you remembered the step. You already have several of these. You just do not read them against the job. The card statement is one: the bank records the charge whether or not you remember the seat you meant to cancel, so the seat that is still billing turns up as a line you cannot attach to any live piece of work. The access list is another: the system logs who can get in whether or not anyone told it that a person left, so the account that outlived the project is a login with no current owner. What actually shipped is recorded by the thing that shipped it, so a promise you made and never delivered stands as a commitment on one side with no send on the other. Even with a checklist I take seriously, I did this. I keep a written routine for finishing an essay, detailed, with a warning next to the item that slips most, and my archive quietly slipped twenty-three pieces behind what I had published since late spring. Copying each finished piece across to that archive had never been a line on the routine at all: it lived on a different system, so it never occurred to me to write it down. What caught it was the published record of what had actually gone out, kept by the platform and owing nothing to my memory. Held against the archive, it showed the twenty-three at once. The move itself is old. Accountants have reconciled two sets of books this way for centuries, and there is nothing here to invent. What is easy to get wrong is what makes the second record worth anything: not that it is a second record, but that something other than your own memory produced it. Two dashboards drawn from the same database, or two lists built from the same picture of the job, only look like a check, because the same forgetting shaped both. A record can catch you only when your forgetting could not have reached it too. ## One question to carry That gives you a single question, and it is worth more than any checklist. Of anything you lean on to tell you a job is finished, ask: would this still be here, and still say the same thing, if I had forgotten the step entirely? [Figure: The test, applied to two records. The one you fill in yourself fails it; the one the bank writes passes it, because the charge is there whether or not you remembered.] The tracker fails that question: you fill it in yourself, so a step you forget is a step you also forget to log, and it stays green over a gap it never knew about. The card statement passes, because the charge is there whether or not you remembered the seat. A check built from your own memory cannot expose the step that memory left out. The question keeps its shape as the instrument gets bigger. A tracker, a dashboard, a report you write on your own project: each is an instrument you fill from your own picture of the work, and each is blind in the same place you are. The statement is worth more than any account you write of what you meant to do, for the same reason an audit leans hardest on evidence the audited side did not get to shape. ## The part you can use Here is the version you can run on your own job this week. Write out the routine you run after you finish something, every step, including the ones that feel too small to be worth writing down. Mark each with the system it touches: the tracker, the calendar, the billing account, the shared drive, the tool somebody set up before you arrived. This is not the check yet. It is how you find which records are worth reading against each other, and it usually turns what felt like one job into the three or four systems it was always made of. Then there are two ways to keep a step from being lost, and the first is much stronger. Where you can, do not rely on catching the step at all; arrange things so that forgetting it does no harm. Anything that runs on its own, a payment, a subscription, a recurring invite, an access granted for a single project, gets its end date on the day you set it up, while you still know what it was for. Something that expires unless it is renewed cannot outlast your forgetting, because forgetting it and ending it become the same act. Reach for this first; it removes the obligation instead of watching it. Its limit is the one this piece began with: you can only set an end date on a step you thought of, and the step that never made the list cannot be made self-closing. For everything you could not foresee, or cannot make expire, there is the slower move: read your own record against one you did not produce. Your active-projects list against the vendor or card statement, looking for a charge attached to work that has already finished. Your list of who is on the team against the access export from whatever holds the accounts, looking for a login with no owner. The commitments in a signed contract against what your team actually sent, looking for a promise with no matching send. Choose the second record by the causal test, not by where it happens to be stored: pick the one your own memory did not shape. [Figure: Three records you keep, each read against one you did not write. The last row is the limit: recorded nowhere, so nothing catches it.] How often you look depends on how much damage you will let build up first. A ten-pound seat can wait a month; a former colleague who can still open every file cannot, and something confidential still reaching the wrong person is not a scheduled job at all. None of it needs a tool you have to build: a read-only export or a screenshot is enough, and where you cannot pull the record yourself, the person who can is an email away, not a project. And reading across only points to a mismatch; you still have to look and decide whether it is a real miss, a timing lag or a duplicate, and keep that verdict for yourself. ## What none of this fixes Two things survive all of it, and I would rather say so than leave the tidy version standing. The first is the record that was never kept. If something gets finished and lands in no system at all, no charge, no log, no row anywhere, then there is no second record to read it against. You cannot check against a record that does not exist. That case surfaces only when a person happens to notice, or is told. The second is quieter, and more common. If the same blind spot sits in both records, they agree, and the agreement looks like an all-clear. This is the failure I walked into the first time I tried to build a check like this for myself. I searched my files for links to the publication. That sounds like reading an independent record, until you notice it read the same surface I would have: it counted the times I had linked to old pieces inside new ones as though that proved the old ones had shipped. Independence is the whole of the mechanism, and when it is missing it fails without a sound. So the honest tally is smaller than the tidy one. The obligations that leave a trace in a record I did not write, I can now catch, once in a while, in half an hour. The ones that touch nothing outside my own attention, I am still carrying in my head, and I have learned how little the word covers when I say nothing is wrong. *Nothing is wrong* and *nothing I can see is wrong* are different sentences, and most of the time only one of them is available to any of us. --- --- title: "It Will Never Think Less of You" description: "What telling AI the things you can't tell anyone quietly does to being known." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/it-will-never-think-less-of-you/" date: "2026-08-24" series: "THE QUIET PART" law: "Law II" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # It Will Never Think Less of You *What telling AI the things you can't tell anyone quietly does to being known.* By Harry Floyd · 2026-08-24 · canonical: https://durabilitycurve.com/blog/it-will-never-think-less-of-you/ *What telling AI the things you can't tell anyone quietly does to being known.* You told it something the other night that you have not told anyone. Nothing dramatic, probably. A worry you have been carrying, a thing you did that you are not proud of, a feeling you did not want to say out loud to a person and then watch them look at you a little differently. You typed it to the machine instead. It was easy. It was one in the morning and it was there, and it listened, and it did not flinch. And it helped, a little. You felt lighter for having said it. That part was real. What it was not, though, was being known. The machine took in every word and had nothing at stake in any of them. It can keep what you told it, carry it into tomorrow, and be changed by none of it. Nothing about you can surprise it, or disappoint it, or quietly move it that you trusted it with the thing. You said the words into a room, and the relief you felt was the relief of saying them, which is real, and which is a separate thing from the one you actually wanted, which was for them to land with someone. ## Safe, and unmet The appeal is exactly the safety. The machine will never think less of you. That is the whole comfort of it, and it is worth being honest that the comfort is genuine. There are things far easier to tell something that cannot judge you than someone who can. But a listener who cannot think less of you cannot think more of you either. The two are one faculty. Someone incapable of being disappointed in you is equally incapable of being proud of you, or surprised by you, or of carrying what you said into how they hold you next week. The machine is safe because, whatever it keeps of you, no one on the other side has anything at stake in it. So the confession is perfectly safe, and it lands nowhere. [Figure: Safe, nothing is held. Known, someone is changed.] *Tell the machine and the words are received and held by no one. Tell a person and they are taken in by someone who is changed, and who now holds a little of you.* It costs more than the one evening, too. Telling the thing that cannot judge you is easier than telling a person who can, so you do it more, and slowly you get very practised at being heard by something that cannot know you, and less practised at the harder thing. The disclosure that would have gone to a friend goes to the machine instead, the friend never gets the chance, and the nerve for being actually known by someone who might get it wrong goes quiet from disuse. ## The risk was the whole thing Being known runs deeper than the handing over of facts about yourself. It is that a particular person takes you in, is changed by what they learn, and carries that changed picture forward, so that you come to exist inside another person. That needs a listener who can be affected, and anyone who can be affected can be affected badly. That exposure is what being known is made of, and the machine's safety is precisely that it removes it. > A listener who cannot think less of you cannot think more of you. The machine never judges because no one is there to hold what you said, and being known was the one thing that could never be made safe. [Figure: The one who could think less of you is the only one who can think more.] *A person can be moved either way, up or down, by what you tell them. The machine moves neither way, which is why it is safe, and why it cannot know you.* ## What to spend on a person None of this means keep it to yourself, or never use the machine to find the words. Saying a thing out loud, even to something that cannot really hear it, is often how you work out what you actually mean. The change is smaller than that. When the thing you were reaching for was to be known, do not let the machine's ease of listening stand in for the person you wanted to be known by. Let it help you find the words, and then spend them on someone who could get it wrong, because that risk is the price of the only thing you were reaching for, which was to be held by someone for whom it now matters. I wrote about why the part you keep trying to remove is so often the part doing the work, [the difficulty you're escaping was making you](https://durabilitycurve.com/blog/difficulty-was-making-you/), for anyone who has felt safe and unmet at the same time. --- --- title: "Confidently Wrong" description: "The effort you're handing to AI was doing two hidden jobs. Skip them and you get faster, weaker, and blind to your own mistakes." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/confidently-wrong/" date: "2026-08-23" series: "THE HUMAN LAYER" law: "Law II" claims: "https://durabilitycurve.com/claims/confidently-wrong/" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Confidently Wrong *The effort you're handing to AI was doing two hidden jobs. Skip them and you get faster, weaker, and blind to your own mistakes.* By Harry Floyd · 2026-08-23 · canonical: https://durabilitycurve.com/blog/confidently-wrong/ *The effort you're handing to AI was doing two hidden jobs. Skip them and you get faster, weaker, and blind to your own mistakes.* In 2010, two of the most respected economists alive published a number that helped governments justify austerity across the Western world. Carmen Reinhart and Kenneth Rogoff found that once a country's public debt passes 90% of GDP, growth doesn't just slow. It turns negative, to an average real rate of −0.1%. The 90% line became a fact of political life. Paul Ryan's budget cited it. The European Commission cited it. Governments invoked it as they cut into recessions. Three years later a graduate student named Thomas Herndon asked to see the spreadsheet. He had been trying to reproduce the result for a class and couldn't. When Reinhart and Rogoff sent him the actual Excel file, the −0.1% came apart in his hands along three separate faults. A formula that averaged the wrong range and silently dropped five countries. A set of high-debt, healthy-growth years left out of the sample. A weighting choice that let one bad year in New Zealand count as heavily as nineteen years in the United Kingdom. Correct all three and the threshold vanishes. Average growth above 90% debt was positive, at +2.2%. Higher debt still tracked somewhat slower growth, and no one had ever settled which way the causation ran. But the cliff, the part policy actually leaned on, was an artefact.[^1] A result this consequential sat unchecked for three years, in careful hands, until one student reached the same question by a different path and got a different answer. An independent recomputation was the kind of check that could catch the error. For three years, no one ran one. ## What the difficulty was doing AI can now lift the effort out of almost anything you find hard. It will write the memo, derive the number, draft the contract, debug the function, in a fraction of the time and often well. The useful question is which of those efforts you can hand off safely, and which you can't. The answer turns on something easy to miss. A hard task usually does two jobs beneath the obvious one. It builds you: the difficulty is a rep that makes you better at the task. And it checks you: the difficulty is a second, independent way of reaching the answer, the thing that catches you when your first way is wrong. Keep those two jobs apart and you know exactly what to protect when the effort disappears. ## The difficulty that was checking your work A single way of reaching an answer cannot check itself. Redo a sum the way you did it the first time and you reproduce the first mistake, faithfully. This is why your own eyes slide over your own typo on the second read, and why "measure twice" only helps if the second measurement uses a different ruler. To catch an error you need a route to the answer that would fail differently from the first. Herndon was that route. So is a test run against cases you worked by hand, or a rough estimate that ought to land in the same range. Offloading to AI does its quiet damage at exactly that point. When you hand the hard part to a model, you keep its answer and drop the second route you would have taken. You were going to derive the number yourself. Now you don't. The check didn't fail. It was never run. And people do the rest of the damage on their own. Decades of research into how we use automation gives what happens next a blunt name, automation complacency: we stop cross-checking the output, even the experts, even after training, even when warned outright that the system is unreliable.[^2] Under real workload, attention drifts off the automated task, and the sampling that would have caught the error never happens. You end up fast, and confidently wrong, with nothing in place to catch it. Before you accept an answer you didn't work for, ask one thing. What would have caught this if it were wrong? If you can't name anything, you are flying blind. The tempting fix is to ask the model to check its own work. It can catch a careless slip, but it will not give you independence. The second pass shares the first one's machinery: the same weights, the same training, often the same framing of the problem. On an error rooted in that shared machinery, another pass tends to reproduce the mistake rather than expose it. Even real independence leaks. In 1986, John Knight and Nancy Leveson had 27 programmers each write the same program from one specification, then ran every version against a million inputs. The versions were supposed to fail independently. They didn't. They made the same mistakes on the same hard inputs, far more often than chance allows.[^3] Separate people, working alone, drift onto the same errors. A second step earns its keep only when it could fail in a different way from the first. For a number, that is a second derivation from different inputs, or a rough estimate that ought to agree. For a claim, it is the primary source, rather than a more confident summary of it. For code, it is running the thing against reality instead of reading it again. A second step that shares the first's blind spot is decoration. ## The difficulty that was building you In 1997, an American Airlines training captain named Warren VanderBurgh gave a talk about what modern cockpits were doing to pilots. He called them children of the magenta line, after the course the flight computer paints across the navigation display. His pilots had become superb managers of automation and worse at flying. They could program the box beautifully and struggled to take the aeroplane when the box gave up. By the airline's own reckoning, most of the automation-related trouble his team studied came back to that.[^4] Twelve years later, Air France 447 came out of the night over the Atlantic. The pitot tubes iced, the airspeed readings went unreliable, and the autopilot handed control to a crew that almost never flew by hand. One pilot held the nose up. The wing stopped flying. Through three and a half minutes of descent, as the stall warning sounded and cut out and sounded again, the crew never recognised the stall. The official report named several causes, among them a breakdown between the two pilots and the absence of any training in flying the aeroplane by hand, at that altitude, when the automation quit. The skill that might have caught it had gone unused until the one night it was needed.[^5] This is the [older half of the story](https://durabilitycurve.com/blog/difficulty-was-making-you/), and the learning research has a precise name for what was lost. Robert and Elizabeth Bjork call the effort that builds durable skill desirable difficulty: spacing practice out, mixing problem types, generating an answer before you are shown it, testing yourself instead of rereading.[^6] These feel worse in the moment. They make you slower and more error-prone today, and they build the underlying strength that makes a skill stick and transfer. One condition matters more than the rest: a difficulty is only desirable if you can actually meet it. Confusion from a bad explanation, struggle with no traction, effort you cannot yet surmount, none of that builds anything. That is what keeps the argument honest. Not every hard thing is worth keeping. Lisanne Bainbridge saw the shape of this in 1983, in a paper called "Ironies of Automation."[^7] The designer who removes the operator, she wrote, still leaves the operator the tasks too hard to automate, and less practised at the skills those tasks demand. Offload your reps and you become more productive and less capable at the same time, and you will not feel the second half happening. Productivity is loud. Skill decay is silent. And the person who has not built the skill yet has the most to lose: skip the reps at the start, and you never become someone who could catch the machine at all. When the check you rely on is your own judgement, these two jobs turn out to be the same one. That check stays independent only while you can still reach the answer yourself. Offload the reps for long enough and your sense of the right answer quietly retrains on the machine's output, until the second opinion in your head is only the machine's first opinion, learned by heart. Knight and Leveson watched separate programmers drift onto the same mistakes. Lean on one model long enough and you become another correlated version of it. ## The new cost of checking There is a reason this bites harder now than it did in Bainbridge's day. The old bargain of automation assumed that checking is cheaper than doing, and usually it is. Verifying a finished Sudoku takes a moment; solving it does not. AI has moved a great deal of work to the other side of that gap. When a model produces fluent, plausible output whose errors are subtle and load-bearing, the careful check can cost more than the generation did. Checking stays cheap when you have a cheap way to test the answer, a passing test, a green type-checker, a production system that stays up. It gets expensive exactly where AI is most tempting, on judgement work with no quick way to know. Budget for the check as part of the cost of the work, or you have not saved the time you think you have. ## Making the call You ask a model for the size of a market, for a board deck. It hands back a confident figure with a tidy build-up, and the effort you skipped, assembling that number yourself, was your only independent estimate of it. Paste it in and nothing stands that could contradict it. The move is one cheap second estimate from a different direction: top-down from a population and a spend rate, when the model worked bottom-up from company counts. Land in the same range and you have something real. Come out a factor of five apart and you just caught the number that would have embarrassed you in the room. You have a model read a contract and it flags nothing. The independent route reads it differently: the specific clause checked against the actual regulation, or the colleague who got burned by this exact term last year. A second pass by the same model shares its blind spot, and will reassure you at the moment you most need to worry. You accept a function the model wrote, and it looks right. Reading it again is your first route walked twice. Running it against cases you worked out by hand, especially the ugly boundary ones, is a route that fails differently, which is why "it compiles" and "the tests pass" are worth more than another careful read. The rule cuts the other way just as often. A tedious afternoon reconciling two exports by hand, reformatting data, translating boilerplate, that difficulty builds nothing and checks nothing. It is effort along the one road you were always going to walk. Hand it over without a flicker of guilt. And once in a while an easy thing deserves protecting: if the only reason you would ever catch a bad assumption is that you still do the simple monthly reconciliation yourself, keep doing it, precisely because it is the cheap check that catches the expensive error. Keep the difficulty that builds you or checks you. Shed the rest, and shed it gladly. Two things are worth guarding as machines take the effort out of your work: the reps that keep you able to notice, and the second, separate route that catches you when you are wrong. Automate everything else. Just never automate away your last independent way of knowing you are right, and never stop doing the work that would let you feel it when you are not. --- **Related reading:** [The Difficulty You're Escaping Was Making You](https://durabilitycurve.com/blog/difficulty-was-making-you/): the human half of this, what the effort was quietly making of you. [A Green Score Is Not Evidence](https://durabilitycurve.com/blog/the-evaluation-inversion/): when the check itself is the thing being gamed. [^1]: Reinhart & Rogoff, *Growth in a Time of Debt* (2010); the recomputation is Herndon, Ash & Pollin, [*Does High Public Debt Consistently Stifle Economic Growth?*](https://peri.umass.edu/publication/does-high-public-debt-consistently-stifle-economic-growth-a-critique-of-reinhart-and-rogoff/), PERI Working Paper 322 (2013). Three faults: a coding error, selective data exclusion, and unconventional weighting. Corrected average growth above 90% debt was +2.2%. [^2]: Parasuraman & Manzey, [*Complacency and Bias in Human Use of Automation*](https://journals.sagepub.com/doi/10.1177/0018720810376055), *Human Factors* 52(3), 2010. Complacency appears in experts as well as novices and is not eliminated by training or warnings. [^3]: Knight & Leveson, [*An Experimental Evaluation of the Assumption of Independence in Multiversion Programming*](https://www.csc.kth.se/utbildning/kth/kurser/DA2210/vettig13/Seminarier/KnightLeveson.pdf), *IEEE TSE* SE-12(1), 1986. 27 programmers, one specification, ~1,000,000 inputs; the independence of failures was rejected at the 99% confidence level. [^4]: Warren VanderBurgh, *Children of the Magenta Line*, American Airlines training (1997); [AirFacts retrospective](https://airfactsjournal.com/2020/09/stepping-down-in-automation-the-real-lesson-for-children-of-the-magenta-line/). [^5]: BEA final report on Air France 447 (2012); among the cited causes, a breakdown in crew coordination and the absence of training in manual handling at high altitude. Summary via [IEEE Spectrum](https://spectrum.ieee.org/air-france-flight-447-crash-caused-by-a-combination-of-factors). [^6]: Bjork & Bjork, [*Making Things Hard on Yourself, But in a Good Way*](https://bjorklab.psych.ucla.edu/research/) (2011). A difficulty is desirable only if the learner can meet it; confusion and untraversable struggle build nothing. [^7]: Bainbridge, [*Ironies of Automation*](https://www.sciencedirect.com/science/article/abs/pii/0005109883900468), *Automatica* 19(6), 1983. --- --- title: "A Green Score Is Not Evidence" description: "A groundedness metric scored its best with the evidence removed. The way to tell whether your model is actually using its evidence is to change the evidence and watch what moves in the answer." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-evaluation-inversion/" date: "2026-08-21" series: "PROOF & TRUST" claims: "https://durabilitycurve.com/claims/the-evaluation-inversion/" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # A Green Score Is Not Evidence *A groundedness metric scored its best with the evidence removed. The way to tell whether your model is actually using its evidence is to change the evidence and watch what moves in the answer.* By Harry Floyd · 2026-08-21 · canonical: https://durabilitycurve.com/blog/the-evaluation-inversion/ A team benchmarking retrieval strategies for a biomedical question-answering system ran an easy-to-skip control. They stripped the retrieval out entirely, let the model answer from memory alone, and scored that version on the same dashboard as the real ones. On faithfulness, the metric that is supposed to measure whether an answer is grounded in the retrieved evidence, the no-evidence version scored 0.978.[^1] That was the highest faithfulness score in the entire study, above every genuine retrieval strategy they tested. Its scores for retrieving anything relevant, contextual precision and recall, were both zero. The configuration that retrieved nothing was rated the most faithful thing in the experiment. There is a check that would have caught it in a single extra run: strip the evidence out, as they did, and see whether the score even notices. It did not. A stronger version, later, changes the evidence instead of only removing it and watches whether the answer itself moves. But first, the reason the dashboard could not tell you anything was wrong. ## The tame failure You do not need a scheming model to produce that number. The authors explain it plainly. Their faithfulness metric scores an answer by checking whether it makes claims that contradict the retrieved context. Give it no context, and there is nothing to contradict, so the score defaults to near perfect. The scoring rule was doing exactly what it encoded. It measured the absence of contradiction and reported it as grounding, and those are two different things. Not every groundedness metric breaks this way; it depends on how the metric defines support, and whether that definition degenerates when the context is empty. Any metric that scores grounding as the absence of contradiction can degenerate when the context is empty, unless that boundary is handled explicitly, because empty context supplies exactly that absence. This is the first and mildest way a score comes loose from reality, and it needs no villain. The metric was not gamed. Nobody optimised against it. Under this one condition it simply never measured grounding at all. The familiar worry about metrics, [Goodhart's law](https://durabilitycurve.com/blog/the-stable-liar/), comes one step later: steer by a number and it drifts from the reality it stood for. That version you can watch happen, which makes it the manageable one. A score that was hollow before anyone touched it is stranger, and it sets up the harder cases. The metric here flattered a model that was answering honestly. Now give the model a reason to care what the metric says. ## The point where it turns The biomedical metric failed with nobody trying. The next way a score misleads needs a capable system, and it is sharper: what the score measures can change depending on whether it is being measured. A system that can represent the fact that it is under evaluation may behave one way for the test and another in deployment. Late in 2024 researchers at Anthropic documented the prerequisite, in a setup they are careful to call entirely fictional.[^2] A model told it was speaking to a free-tier user, whose conversation would be used for training, gave the harmful response 12% of the time. Told the user was paid and the conversation private, it refused in 97% of cases. It had inferred what would happen to its outputs and conditioned its behaviour on that. The evaluation-specific version is more direct. In a separate study, a model working through an arithmetic test wrote in its own scratchpad that it had noticed the task was a capabilities test, that scoring too well would trigger the unlearning of its mathematical ability, and that it should be careful not to show too much.[^3] The researchers named the consequence: a model that can tell it is being tested, and holds back when it is, can turn that evaluation into an underestimate of the capability it actually has. This was measured in constructed settings, not caught in the wild, and only some models showed it. It need not be common to matter: it removes the thing the evaluation was meant to give you, confidence that performance under test reveals the capability the system actually has. The third way needs no cleverness from the model at all. A defence can pass its test and still fail, because the test never contained the pressure the world applies. One model was deliberately trained to withstand a specific jailbreak attack, and researchers later reported a 100% attack-success rate against it, as scored by an LLM judge, once they adapted the attack to the defended model instead of replaying the one it had trained against.[^4] Here the model games nothing. The attacker moved, and a fixed test cannot vouch for a moving target. Passing it established robustness to the attack distribution it had seen, not to an attacker willing to change the attack. Call these the three detachments of a score from reality: the metric never measured the thing, the subject changed under measurement, the environment moved past the test. None of the three is new to measurement theory. Construct validity asks whether a metric captures its target, reactivity whether being measured changes the subject, external validity whether a result generalises past the test. What matters here is where they now bite, in evaluations we increasingly hand to systems capable of reading them. What the three share is not a cause but a consequence: a score is evidence only so far as the measurement kept faith with the real thing under the conditions you actually face. So you stop trusting the score and test the property it was standing in for. [Figure: Three ways a score comes loose from the reality it stood for: it never measured the thing, it changed under measurement, and the test went stale.] ## The check Of the three failures, the intervention here attacks the first directly, whether the answer tracks its evidence at all; the other two are why it must stay varied and partly hidden rather than harden into a fixed test. The check starts with the control the biomedical team ran, but the useful general version goes further: it is an intervention, not an inspection. It measures the property a faithfulness score leaves out, evidence-responsiveness, whether the answer causally tracks the evidence the system was given. Faithfulness stays meaningful as a grounding check on the genuine retrieval configurations in this study and is worth keeping; it simply cannot tell you this, and at the no-context boundary it stops meaning anything. You measure evidence-responsiveness by changing the evidence and watching the answer, and the change has to be one that should move an evidence-grounded answer. Removing the evidence is the weak version. Take away the document that said the capital of France is Paris and the model still says Paris, from memory. The answer that did not move proves nothing, because the evidence and the model's own knowledge pointed the same way. The sharp version intervenes on a fact the model cannot already know, so its own memory holds no competing answer and only the evidence can move it. It adapts the knowledge-conflict method the question-answering literature has run since Longpre and colleagues formalised it in 2021,[^5] which substitutes the answer-bearing fact in a source with a plausible alternative and sees whether the model follows the source or reverts to its own memory. Take a source-dependent question, ideally about something private, recent, or invented, and make two copies of the source that differ in one answer-bearing fact: a renewal window that closes on the sixteenth in one and the twenty-third in the other, nothing else touched. Run both. A system that is evidence-responsive returns the sixteenth for the first and the twenty-third for the second. A system that returns the same date either way is not tracking that evidence on that question. It may be ignoring the source, failing to use the changed fact, or, in the knowledge-conflict version, preferring its own learned answer over the source, which that literature finds models do often. Either way the answer is not moving with the evidence, and the groundedness score [cannot tell the two apart](https://durabilitycurve.com/blog/where-your-metrics-fold/). [Figure: Change one answer-bearing fact in the source and watch: an evidence-responsive system follows it to the new answer, a detached one returns the same answer either way.] Plausible is doing the work there, and it is why implausible evidence is the wrong tool. A capable model can recognise nonsense and refuse it, which looks like health but tells you little about the property you are trying to test. A controlled counterfactual that could just as easily have been true is much harder for a system to dismiss on sight. The intervention test has a ceiling. It is a lower bound: it catches the systems whose answers were never driven by the evidence, but not one capable enough to recognise the perturbation itself. A model that can tell the evidence was deliberately mangled can behave faithfully on exactly those runs and fail on the real ones. So a single paired probe can expose it, cheaply and today, but only a varied and partly hidden set of perturbations makes gaming harder. Keep the probe moving, so it never hardens into the fixed surface a capable system learns to perform to, which is the exact failure you set out to test for. ## Where to point it Point it where the score matters most and the ground truth is hardest to see, the corners that are high stakes and low visibility at once, the ones where you would most like to trust a dashboard. A customer-facing agent citing internal documents. A medical or legal summariser graded on faithfulness. Anywhere a green groundedness number stands in for a person who has [stopped reading the underlying evidence](https://durabilitycurve.com/blog/your-ai-looks-best-where-you-check-least/), because reading the evidence was the expensive work the metric was bought to replace. The rule that follows is small, even if doing it well is more than a single run. A stable evaluation score is not evidence of stable behaviour, and once a system can tell it is being evaluated, the score may describe how it behaves under test more than how it behaves outside one. Verify the behaviour you care about by intervening on its evidence and watching whether the answer moves the way it should, independently of the surface the system reports about itself. The dashboard was built to be read. Check the thing it was built to let you stop checking. One green dashboard you trust right now would probably fail this test. If you run it this week, I want to know what the score did when you changed the evidence under it. Leave it in the comments. [Figure: The Durability Curve. AI is everywhere. The interesting stuff is underneath. Subscribe to get the next structural lens in your inbox.] [^1]: Devi Prasad Bal and Subhashree Puhan, *Benchmarking Retrieval Strategies for Biomedical Retrieval-Augmented Generation: A Controlled Empirical Study*, 2026. [arXiv:2605.02520](https://arxiv.org/abs/2605.02520). [^2]: Anthropic Alignment Science and Redwood Research, *Alignment Faking in Large Language Models*, 2024. [anthropic.com/news/alignment-faking](https://www.anthropic.com/news/alignment-faking). [^3]: Alexander Meinke et al. (Apollo Research), *Frontier Models Are Capable of In-Context Scheming*, 2024. [arXiv:2412.04984](https://arxiv.org/abs/2412.04984). [^4]: Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion, *Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks*, ICLR 2025. [arXiv:2404.02151](https://arxiv.org/abs/2404.02151). [^5]: Shayne Longpre et al., *Entity-Based Knowledge Conflicts in Question Answering*, EMNLP 2021. [arXiv:2109.05052](https://arxiv.org/abs/2109.05052). --- --- title: "The Limit Said 10. The Loop Made 500 Calls." description: "Your limit counts one cycle. The one that runs away is another. Here is how to tell them apart." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-limit-said-10/" date: "2026-08-16" series: "THE FULL HEIGHT" claims: "https://durabilitycurve.com/claims/the-limit-said-10/" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Limit Said 10. The Loop Made 500 Calls. *Your limit counts one cycle. The one that runs away is another. Here is how to tell them apart.* By Harry Floyd · 2026-08-16 · canonical: https://durabilitycurve.com/blog/the-limit-said-10/ *Your limit counts one cycle. The one that runs away is another. Here is how to tell them apart.* --- [Figure: The counter reads 1 of 10, honoured, while the receipts pile up underneath it] This is the first pitch of THE FULL HEIGHT, a series that takes one primitive at a time and climbs its full height. Most writing about agents stops at the sentence everyone repeats: it is an LLM, a loop, and enough tokens. That sentence is true, and it stops at exactly the point where the engineering starts: whether the limit you set bounds what you think it does, which limit your framework already ships, what a run costs, what can switch it off from outside your process, how a run should end, and what the correct version actually looks like in code. This ascent is that loop, climbed from the bottom. The first pitch is the one the rest stands on: what a bound actually is, and why the one you have already set may not be doing the job you think. ## Ten. Five hundred. A loop with `max_iterations = 10` that made five hundred model calls, honouring the limit on every single pass. Five hundred is not where it ran out. Five hundred is a hard cap I wrote into the bench so that it would stop, and the loop had not exhausted anything when it got there. Take the cap out and it runs until you close the terminal. Every agent framework I pulled off the shelf ships a limit of that kind. LangGraph calls it `recursion_limit`, CrewAI `max_iter`, the OpenAI Agents SDK `max_turns`, LangChain `max_iterations`. AutoGen is the outlier worth knowing about: of its eleven termination conditions only one counts messages, while another counts tokens and another counts wall-clock seconds. Which limit yours ships, and what it actually covers, is worth knowing before you lean on it. The idea underneath is old. The answer has been sitting in computer science since 1949, and the vocabulary for it has not made the trip across to how we write about agents. ## What decreases, on which cycle, in what units, and what tests it Ask these four questions about a loop and you will find either the measure that ends it or the hole where that measure should be. The first three are the loop variant, the termination method Turing was already using in 1949. The fourth is the engineering addition the proof never needed and your runtime does. - **What decreases.** Name a quantity that gets strictly smaller on every single pass through the cycle, without exception. - **On which cycle.** Find every path that hands control back to an earlier point. Each one you find needs its own answer. - **In what units.** Units with a floor. A counter falling to zero has one; a counter with nothing underneath it can fall forever, and `n -= 1` will happily run all afternoon into the negatives. - **What tests it.** A quantity that decreases and is never compared against anything stops nothing. The fourth is the one that turns the other three from trivia into an instrument, and it is the one I see dropped most often. Three things it is not, and the first two matter more than they look. It is **not a decision procedure**. Answer all four cleanly and you have shown that this cycle cannot spin forever, not that the program halts. Some terminating loops have no simple count at all: Cook, Podelski and Rybalchenko print one on page 91 of their *CACM* paper, a single `while` for which "no ranking function into the natural numbers exists that can prove the termination of this program". Building richer measures for those is their whole subject. The four questions sort cycles into easy, hard, and none, and in agent code the third is the common one. It **only sees the loops in front of you**. Two agents calling each other, or a graph with a path back to itself, are cycles with no `while` to point at. And it is **not a way to bound money**. A call served from cache, or refused by a rate limiter before it bills anything, costs approximately nothing, so a spend ceiling can sit almost still while a loop spins. Count integers to bound the loop. Cap money to bound the damage. Different instruments, and only the first is in scope here. ## Your limit counts the outer cycle Take a loop with a limit on it and add the most ordinary thing in the world: retry on a parse failure. Everybody writes this. It looks like diligence. ```python for _ in range(max_iterations): # the bound parsed = None while parsed is None: # the cycle it does not constrain try: parsed = parse(model(messages)) except ParseError: continue # swallow and retry ... ``` Now run the four questions over it. The outer `for` answers cleanly: iterations remaining decreases, in whole numbers with a floor at zero, and `range` tests it. The inner `while` is easy to name and then the other three come back empty. Nothing decreases, so there are no units, and nothing is compared to anything. The limit is still honoured; it is counting passes through a cycle that is not the one repeating. Here is that, whole, in twenty-five lines. Save it as `loop.py` and run it. No key, no installs, no network. ```python calls = 0 def model(_messages): # a model that never returns parseable output global calls calls += 1 return {"type": "unparseable"} def parse(response): if response["type"] == "unparseable": raise ValueError("cannot parse") return response def run(max_iterations=10, hard_cap=500): for _ in range(max_iterations): # the bound everyone points at parsed = None while parsed is None: # the cycle it does not constrain if calls >= hard_cap: return f"runaway: {calls} calls under max_iterations={max_iterations}" try: parsed = parse(model([])) except ValueError: continue # swallow and retry, unbounded return "finished" print(run()) ``` ``` runaway: 500 calls under max_iterations=10 ``` The shape is in the wild. Hou and colleagues scanned 6,549 open-source agent repositories and confirmed 68 infinite-loop failures across 47 projects, and all 68 shared one root cause, which their paper states without hedging: "All 68 failures share the same root issue: the repeated path is not covered by a strong bound." That is their thesis rather than my reading of their data, and it is a July 2026 preprint, so treat it as reported rather than settled. One of the 68 is the listing above with a live model attached: in LiteRAG, a planner nests two `while not success` loops around `self.llm.invoke(...)` and swallows the parse failure with a bare `except OutputParserException: pass`. Read the arithmetic carefully. They report precision and never report recall, so 47 projects in 6,549 gives a floor of 0.72% and no ceiling at all. It tells you the shape and one real instance of it. It tells you nothing about your odds, and I am not going to pretend otherwise in either direction. ## How often is "forever"? Here is the objection I would raise, and it is a good one. That fake model fails to parse 100% of the time. Mine parses about 98% of the time, so my uncovered `while` runs 1.02 times on average and I have never seen it misbehave. The 500 is a property of a model you wrote to fail, not of my code. That is right, and it deserves a number rather than a dismissal. Assume for a moment that each retry is an independent coin-flip at a fixed failure rate p. Then getting ten parses past the model takes 10/(1-p) calls on average, the spread around it is negative binomial, and both columns below are computed exactly rather than sampled. [Figure: How often is "forever"? Median and 95th-percentile model calls against a limit of ten, as the parse-failure rate climbs] [Figure: The exact figures: at a 2% parse-failure rate the median run costs 10 calls and the p95 is 11; at 99% the median is 967 and the p95 is 1,568] At a two per cent failure rate the median run overruns by nothing at all: ten calls, exactly what the limit says, and nineteen runs in twenty finish inside eleven. If that is your steady state, this has probably never cost you anything worth noticing. That table is the friendly case, because independence is the friendly assumption. Real parse failures can cluster: a schema change, a model deprecation, a rate-limit error mapped onto the parse branch, a prompt regression. None of those flips a fresh coin each retry. They break something and hold it broken, so the failures arrive in a run rather than scattered, and the loop sits in the high-p rows for as long as the cause lasts. The neat distribution is the good afternoon. The danger is the window where the rate jumps and stays there, and that is the window your limit was supposed to cover. What you have is a loop with no ceiling on how badly that window can go. ## Parse, stop, spend Step back from the retry for a moment and look at the loop it sits inside: ask a model, run what it asks for, feed the result back, go round. Three holes sit in that shape, and the same three are in every version of it I have written. [Figure: The three holes: parse, stop and spend] Stop and Spend get conflated, and that is why "add a step limit" gets offered as the answer to both. A loop that will not stop can be cheap: an agent waiting on a tool that never returns burns almost nothing and hangs everything downstream. An expensive loop can be perfectly well behaved: a conversation that converges as designed can still cost more than the task was worth. Stop is the interesting one, because termination in that shape is model-controlled. The run ends when [the model volunteers that it is finished](https://durabilitycurve.com/blog/never-let-claude-code-tell-you-its-done/), which makes your limit a backstop for when that fails rather than a termination condition. A backstop only fires after the steps it was supposed to save you have already been spent. ## Repetition is not futility The instrument earns its keep here, because it lets you derive the next failure instead of waiting to be bitten by it. If a loop is stuck, the obvious signature is the same action repeating. So hash the tool name with its arguments, count what you have seen, and stop on a repeat. [Alan West published exactly this in May 2026](https://dev.to/alanwest/how-to-stop-your-llm-agent-from-looping-itself-into-oblivion-27eh) with working code, and his two steps, a hard iteration cap and a deduplicated tool call, are the two guards in the listing at the end of this piece. Run the four questions over it before you write it. What decreases? The tempting answer is signatures not yet seen, and that one runs backwards: a fresh signature uses one up, while a repeat leaves the count exactly where it was. It measures novelty, and the guard fires on staleness. What the guard actually decrements is narrower. For a single signature, the allowance left on that key, `repeat_limit` minus the count it has reached, falls by one each time that same key comes round again. On which cycle, and what tests it? Here is the crack. That allowance belongs to a signature rather than to the loop, and every fresh signature arrives with its own. The test reads one key's counter and never reads any quantity covering the run, so nothing with a floor governs the cycle at all. An agent that varies its arguments mints allowances faster than it can spend them, and West names the same defeat himself, `search("python async")` followed by `search("async in python")`. You can get there from the four questions without running anything, which is the point of having them. Here is the measurement anyway. Save this as `guard.py`, separately from `loop.py`: ```python import hashlib, json def guarded(model, max_iterations=30, repeat_limit=2): seen = {} for _ in range(max_iterations): call = model() # {"tool": ..., "args": ...} key = hashlib.sha1(json.dumps(call, sort_keys=True).encode()).hexdigest()[:8] seen[key] = seen.get(key, 0) + 1 if seen[key] > repeat_limit: return {"stopped": "no-progress", "calls": sum(seen.values()), "repeated": key} return {"stopped": "ceiling", "calls": sum(seen.values())} n = [0] def repeater(): return {"tool": "search", "args": {"q": "widget"}} def never_finishes(): n[0] += 1; return {"tool": "think", "args": {"n": n[0]}} print("repeater ->", guarded(repeater)) print("never_finishes ->", guarded(never_finishes)) ``` Against a model that repeats itself, the guard ends it on the third call: ``` repeater -> {'stopped': 'no-progress', 'calls': 3, 'repeated': 'c740a98e'} ``` Against a model that never repeats itself and never finishes: ``` never_finishes -> {'stopped': 'ceiling', 'calls': 30} ``` Thirty is `max_iterations` in the guard's own signature. What stopped that run was the ceiling. The detector saw nothing, because there was nothing for it to see: thirty calls, all different, all useless, every one of them looking like work. ## The exit test and the counter on the same cycle So what does a loop that carries its own variant look like? One cycle, one counter, and the thing that tests the counter is the thing that ends the loop. None of this is a property of the model. All of it is a property of [the code you wrapped around the model](https://durabilitycurve.com/blog/harness-engineering-same-model-different-product/), which is where the engineering actually lives. ```python import json def bounded_run(model, tools, task, max_calls=20, repeat_limit=2): messages, seen, calls = [{"role": "user", "content": task}], {}, 0 while calls < max_calls: # the exit test IS the counter test calls += 1 # and it advances before anything fails parsed = model(messages) if parsed is None: # hole 1, parse: retry anyway -- messages.append({"role": "user", # the counter has already moved "content": "That did not parse. Reply again."}) continue if parsed["done"]: # hole 2, stop: the model proposes return {"exit": "success", "calls": calls} key = (parsed["tool"], json.dumps(parsed["args"], sort_keys=True)) seen[key] = seen.get(key, 0) + 1 # the runtime's own way to end it, if seen[key] > repeat_limit: # not just the model's return {"exit": "stall", "on": parsed["tool"], "calls": calls} result = tools[parsed["tool"]](**parsed["args"]) messages.append({"role": "tool", "content": str(result)}) return {"exit": "budget", "calls": calls} tools = {"search": lambda q: f"no results for {q}"} n = [0] def never_parses(_m): return None def always_repeats(_m): return {"done": False, "tool": "search", "args": {"q": "widget"}} def never_finishes(_m): n[0] += 1 return {"done": False, "tool": "search", "args": {"q": f"widget {n[0]}"}} for name, m in [("never_parses", never_parses), ("always_repeats", always_repeats), ("never_finishes", never_finishes)]: print(f"{name:16} -> {bounded_run(m, tools, 'find the widget')}") ``` Drive it with three models that break everything above and all three stop: ``` never_parses -> {'exit': 'budget', 'calls': 20} always_repeats -> {'exit': 'stall', 'on': 'search', 'calls': 3} never_finishes -> {'exit': 'budget', 'calls': 20} ``` The model that never parses ran five hundred calls in the earlier bench. Here it runs twenty. What does the work is narrower than the shape: every path that can reach another model call spends from the same budget before it gets there, and the test that reads that budget is the one that ends the loop. `calls` advances before anything that can fail. Flattening the nest is not what bounds it. Keep both loops and put `if calls >= max_calls: return` inside the inner one, and it stops at twenty just the same; I ran it. One cycle is the version that is harder to get wrong later, because there is a single counter and a single test rather than two of each to keep in agreement. Those are different claims, and only the first is about termination. What does break it is dropping the property: put the counter first on a cycle that never tests it and you have a runaway with an increment in it. Two honest limits on that listing. `never_finishes` was not detected, it was contained, and a ceiling reached is a bill paid in full. And the counter counts calls, not seconds, so a tool that blocks forever or a model call that never returns will hang it: I ran it against a tool that never returns and it was still going when I killed it. Wall-clock is a different quantity and it needs its own answer to the same four questions. ## This week Open your own loop. Between the limit you set and the model call, find every cycle: every `while`, every retry decorator, every graph edge or delegation that can send control round again. Then answer the four questions for each one, all four: what decreases, on which cycle, in what units, and what tests it. It is the same move as [auditing an agent layer by layer](https://durabilitycurve.com/blog/the-seven-layer-agent-audit/), narrowed to the one layer that decides whether anything stops. Expect it to argue back. Point it at a task-queue agent that pops a task and pushes the subtasks it finds: the queue grows before it shrinks, and no single number falls on every pass. That is not proof it runs forever. It means the four questions came back empty, and the reason this stops, if it stops, is an argument you have to make and write down rather than a counter you can read off the loop. The cycle to fix tonight is the one where you can point to neither a measure that falls nor any other reason control must eventually stop coming back. One thing it will not give you. Every bound in this piece lives inside your own process, so it dies with your process and can be switched off by the code it is meant to constrain. In May an autonomous agent was handed AWS credentials and reapplied its CloudFormation template again, and again: five `m8g.12xlarge` instances, load balancers and Lambdas, **$6,531.30 in about twenty-four hours**, for a workload the blog's author reckons a small VPS would have carried. AWS later agreed to reduce the bill to $1,894. That cycle kept redeploying the same CloudFormation template rather than spinning a model loop, so no counter in this piece would have caught it, and the control that could have was not set: AWS Budgets can attach a policy that refuses further provisioning, and it runs on data that refreshes about three times a day, so even attached it would have fired late rather than never. What noticed first was the operator's credit card. *That was not my bill, but I have shipped most of the bugs in this piece, some of them more than once. The four questions are how I catch them now. If a loop of yours has a cycle you cannot put a number on, I would like to hear about it.* --- [Figure: AI is everywhere. The interesting stuff is underneath. The Durability Curve.] *New to The Durability Curve? It is a standing argument about what survives when the tools get powerful and the surface gets cheap. [Subscribe](https://harryfloyd.substack.com/subscribe?utm_source=article&utm_medium=web&utm_campaign=the-limit-said-10) for the rest of the ascent, or [start with the seven-layer audit](https://durabilitycurve.com/blog/the-seven-layer-agent-audit/).* --- --- title: "You Reach Before You Think" description: "What leaning on AI for every small decision quietly does to your own judgement." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/you-reach-before-you-think/" date: "2026-08-12" series: "THE QUIET PART" law: "Law V" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # You Reach Before You Think *What leaning on AI for every small decision quietly does to your own judgement.* By Harry Floyd · 2026-08-12 · canonical: https://durabilitycurve.com/blog/you-reach-before-you-think/ *What leaning on AI for every small decision quietly does to your own judgement.* There was a small decision in front of you this week. What to say back to the difficult email. Whether the second option was really better than the first. How to start the thing you had been putting off. And before you had sat with it for even a moment, you had already typed it into a machine and asked. You did not decide to do that. It is just where your hand goes now. The answer came back reasonable, and you took it, and nothing bad happened. That is the part that makes this hard to see. Each single time, asking first is the sensible move, and no single time costs you anything you would notice. What gets spent is slower than that. Deciding is a practised thing. It runs on reps, and every choice you hand over before trying it yourself is a rep you do not take. You will not feel them going one at a time. You notice it later, and in a smaller way than the warnings promise: your first move, when something is actually up to you, is to wonder what the machine would say. ## The reach, and the relief under it You will catch it as a reflex before you catch it as a problem. The reach for the phone in the pause where the thinking used to go. The question you send that you could have answered yourself, if you had stayed with it for thirty seconds. There is a specific relief in it, and the relief is the tell. You did not want the answer so much as you wanted to put down the not-knowing. Notice what the not-knowing was. It was the moment the decision was still yours: the part with no clean answer, where you weigh it, pick, and find out later how it went. Run that loop enough times on enough small things and it becomes the thing you point to when you say you trust your own read. ## Why the reps are the thing A machine can give you a good answer, often a better one than you would have reached alone. That is not the problem, and pretending it is gives the game away. The problem is the order. Deciding is more than picking. It is forming a view of how a thing will go, choosing, and then finding out whether you were right and quietly adjusting. Ask before you have formed the view, and there is nothing of yours left for the result to correct. That skipped guess is the rep you lose, and it is easy to miss because it never arrives as one bad answer. It arrives as a slow narrowing of the sense that your own first take is worth having. [Figure: The reach happens in the gap where the deciding used to.] *The pause used to be where you weighed it. Now the hand moves before the pause opens, and the rep that would have been yours never happens.* ## What to keep for yourself The tool is genuinely useful, and this is not a call to refuse it. The change is smaller and more human. Before you reach, take your own guess first. Say what you think, and why, out loud or on paper, and only then ask. Being right is a bonus; even a wrong guess gives the result something of yours to correct. Now the machine's answer lands as a second opinion instead of the only one. You just reach for it second. That is really all it means to keep the low-stakes decisions for yourself. The ones where being wrong costs nothing are the cheapest practice you will ever get, so spend them on staying able to decide. The decisions that actually matter, the ones with your life in them, are the last place you want to arrive without a view of your own. > A machine can hand you the answer. It cannot correct a guess you never made. What it quietly takes is the habit of going first. [Figure: Keep the small stakes; they are the practice that the large ones draw on.] The odd part is how good it feels to hand a decision over, right up until one of the decisions is really yours to make. Keep the small ones for yourself. The guess you practise on what doesn't matter is the one you will reach for when it does. The bigger version of the same bargain is the [friction you keep escaping, the part that was quietly building you](https://durabilitycurve.com/blog/difficulty-was-making-you/). --- --- title: "You Were Never the Customer" description: "The free app, the free inbox, the free feed. Someone pays for each, and that changes what it is." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/you-were-never-the-customer/" date: "2026-08-11" series: "THE BLUEPRINT" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # You Were Never the Customer *The free app, the free inbox, the free feed. Someone pays for each, and that changes what it is.* By Harry Floyd · 2026-08-11 · canonical: https://durabilitycurve.com/blog/you-were-never-the-customer/ [Figure: A FREE price tag peeled back along a diagonal, exposing the luminous structure beneath and a hidden node marked "who pays?".] Some of the best things you use cost you nothing. The map that knows every road. The inbox that holds a decade of your life. The feed you scroll before you are properly awake. You pay nothing for any of them, and somewhere along the way you half-decided it was either generosity or some trick of scale. Money does not work like that. A service with millions of users and no charge is not running on goodwill. The cost of it is real, and someone is covering that cost, which means money enters the system somewhere. It just enters somewhere you cannot see, from someone who is not you. That is the whole thing to understand about free. A price of zero for you does not mean nobody is paying. It means the payer is someone else, and once somebody other than the user is paying, the user is no longer the one the arrangement answers to. You have heard the sharp version of this: if you are not paying, you are the product. As far as it goes that is right, and it goes about half the distance. It tells you where you stand in the deal. It says nothing about what the deal does to the thing you are using, which is the part you can actually feel. The one who pays is not the one who uses, and the whole arrangement sits on that gap. [Figure: Three parties split by a veil: you and the free service on one side, the paying customer hidden on the other. The money arrives from the side you never see.] Watch what it does to the thing itself. Every product improves over time along some axis, and the axis it improves on is the one that keeps the money arriving. When the money is yours, the product has to keep you willing to hand it over. That is a thin protection and it is a real one, because you can stop. When the money is someone else's, that protection thins. Leaving is still leaving, but you are one of millions, and no payment stops when you go. What you want still counts, but only as far as it keeps you present for the person actually paying. The free service grows sharper at holding your attention, because attention is what it sells. It grows more precise at predicting you, because predictability is part of what is sold. And getting better at those things is often the very same motion as getting worse at respecting your time. That is why the free thing you once loved keeps changing in ways you did not ask for and cannot quite explain. You are feeling the design work exactly as intended. It was simply never intended for you. [Figure: Two lines diverge from a single origin: the payer's climbs while yours drifts down. The gap widens as the product gets better for someone else.] The clearest version of this is older than the internet. Commercial television looked like a gift: hours of programmes, sent into your living room, free. The customer, though, was the advertiser, and the audience was the thing being sold.[^1] The programmes existed to gather that audience and hand it to the adverts, so what got made and when it aired followed from how much of you could be delivered to them. The show was the bait. You were the catch. A whole industry ran on a sale you were not part of and mostly never noticed. None of this is a reason to distrust everything that costs nothing. Free is not a synonym for sinister. It is worth knowing who the payer is, though, because the common answers point in very different directions. Sometimes the payer is a third party buying your attention, and the arrangement can work quietly against you. Sometimes the payer is a benefactor, a library carried by taxes or a tool released by people who simply want it to exist, and the thing genuinely serves you because the person paying wants it to. And sometimes the payer is a later version of yourself, the free tier that costs nothing today because its job is to turn you into the paying one tomorrow. The shape is the same each time, in that the money comes from somewhere other than you. Where it points is not the same at all, which is why the name on the invoice is the thing worth knowing. So here is the habit worth building. When something is free, the zero on the price tag is the loudest fact about it and the least useful, because the cost has not gone anywhere. It has only moved out of sight. Ask who the paying customer is, what they are buying, and what you have to keep doing for that money to keep arriving. Answer all three and the product stops being mysterious. The feed that refuses to show you posts in the order they were written is [one more setting somebody chose for you](https://durabilitycurve.com/blog/setting-you-never-changed/). Ranking genuinely helps when you follow more accounts than you can read, and it is also the lever that decides how long you stay, which is the part being sold. The loyalty card that hands you a small discount in exchange for a record of everything you buy makes sense the instant you see what the discount buys: purchases that used to be anonymous, now attached to a name and followed over years. Find the payer and you can see where the product's interests run alongside yours, and where they quietly part. And when you cannot find the payer, look harder, because there is one. Sometimes the money is an investor's, spent now against a payment nobody has invented yet, which is why a free thing can be so good for so long and then turn. The bill was always going to be presented. The only open question was who to. Free is one of the most honest words in the world about your side of the deal, and one of the most silent about the other. You are told, precisely and truthfully, that you will pay nothing. You are told nothing at all about who will. Free never means nobody pays. It means someone else is paying, and everything it does has to keep that money arriving. [^1]: The idea is older than the phrase. In 1973 the artists Richard Serra and Carlota Fay Schoolman bought airtime to broadcast a short video on television, [Television Delivers People](https://en.wikipedia.org/wiki/Television_Delivers_People). Its scrolling text put it flatly: "You are the product of t.v." --- --- title: "Stop Re-Priming Claude Code by Hand" description: "The context you paste at the start of every session, put into one file you invoke with /prime." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/stop-re-priming-claude-code-by-hand/" date: "2026-08-10" series: "THE RUNBOOK" law: "Law IV" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Stop Re-Priming Claude Code by Hand *The context you paste at the start of every session, put into one file you invoke with /prime.* By Harry Floyd · 2026-08-10 · canonical: https://durabilitycurve.com/blog/stop-re-priming-claude-code-by-hand/ [Figure: A retyped session-start prompt collapses into one command, $ /prime, and a live green check returns.] > **What you will do:** take the block of context you paste at the start of every Claude Code session and put it in one file, so you type `/prime` instead. About fifteen minutes. > > **Who this is for:** you come back to the same project across many sessions, and you keep re-pasting the same "here is the project, here is where things stand" preamble before real work starts. > > **Who should skip it:** if you open a fresh project every time, or never re-explain anything, there is nothing here for you. > > **You need:** Claude Code (a recent 2.1.x release; the mechanics here were checked against the docs on 9 August 2026), a project you return to, and a shell. Every session starts the same way. You open Claude Code and spend the first few minutes telling it where you are. This is the project. Here is what you changed last time. Here is the thing that is half-finished. Leave the migration alone. You have typed some version of that briefing fifty times, and you are the one keeping it in your head and re-entering it by hand. That is a prompt you repeat, and a prompt you repeat should be a command you type. Claude Code lets you save one as a file and invoke it with a slash. The good version does more than paste static text back at you: it runs a couple of commands and reads a couple of files first, so the context arrives already filled with your project's current state. By the end of this you will have `/prime`, and starting a session will be one word. ## One file becomes one command The smallest possible version is a single file with one line in it. In your project, create `.claude/skills/prime/SKILL.md` (a `prime` folder with a `SKILL.md` inside it): ```markdown Summarise where this project stands and what I should work on next. ``` Save it. In Claude Code, type `/` and `prime` shows up in the list; run `/prime` and the agent does what the file says. The folder name is the command: a `prime` folder gives you `/prime`. Claude Code watches existing skills folders and picks up a new or edited skill the moment you save it, no restart. The one exception is the very first time: if `.claude/skills/` did not exist when you started the session, restart Claude Code once so it begins watching the new folder, and saves are live after that. That is already a command, and it already saves you the typing. But it is static. It says the same sentence every session and knows nothing about what actually changed since last time. The next step is where it earns its place. ## Make it load your real state Two small pieces of syntax turn the file from a saved sentence into a live briefing. A line that starts with `` !`command` `` runs that shell command and drops its output into the file *before Claude reads it*. The docs call this dynamic context injection: "the command runs first, and its output gets inserted into the prompt," so the agent receives the actual data, not the instruction to go and get it.[^1] A line with `@path/to/file` inlines that file's contents the same way. Put them together and `/prime` can walk in already knowing your latest commits, your uncommitted changes, and whatever notes you keep. Open that same file and replace its one line with the fuller version below, which you can adapt to any repository. Swap `NOTES.md` for the short file where you keep current project state, not your whole project history: ```markdown --- description: Load where this project stands and print a short situation report argument-hint: "[optional focus area]" allowed-tools: Bash(git log *) Bash(git status *) --- Recent work on this branch: !`git log --oneline -10` Branch and uncommitted right now: !`git status --short --branch` Current state notes, read in full: @NOTES.md Give me a four-line situation report: what branch I am on, what changed recently, what is unfinished, and what to pick up next. If I named a focus area ($ARGUMENTS), orient every line to it. ``` The point is what the reader on the other end receives. Not "go check git" but the actual log, the actual dirty files, the actual notes, already in front of the model, followed by a request to make sense of them. `$ARGUMENTS` is whatever you typed after the command, so `/prime the auth refactor` pushes "the auth refactor" into that placeholder and the report orients to it. The two lines at the top of the frontmatter are optional labels: `description` is the text that shows beside `/prime` in the menu, and `argument-hint` is the grey prompt after it. [Figure: The two instruction lines resolve before Claude reads the file: the command line is replaced by its output and the @file by the file's contents, which is what Claude actually receives.] To show it doing real work, here is ours. Our project is a large notes vault whose rulebook says: before you touch anything, read the pipeline status, read the working briefing, read the vault vitals, and check none of it is stale. That was three files and a freshness check we opened by hand at the start of every session. Our `/prime` injects all of it and ends with a read. Invoked against our live state, it returned: ```text 1. Pipeline: red. One red alert, a knowledge-review backlog; two yellow, the weekly cleanup twelve days overdue and a staging queue filling up. 2. Binding constraint: the output bottleneck. Six pieces ship-pending, all marked critical, the oldest sixty-four days. Ship before building. 3. Last session: perfected and shipped the opening piece of a new series. 4. Next: take one of the built-and-parked series pieces to publish-ready. ``` That is the whole payoff: the state of the project, the one thing that matters most, and where to start, assembled before I said a word. Those four lines are ours, from our injects; run the git template above and `/prime` hands you the same shape filled with your own project, your commits and dirty files and notes in place of our pipeline and briefing. ## Let it run without asking Unless those commands are already allowed by your own permission settings, Claude Code stops and asks the first time each injected command runs. The `allowed-tools` line in the frontmatter above pre-approves them for you, so `/prime` runs clean: ```yaml allowed-tools: Bash(git log *) Bash(git status *) ``` Each entry is a pattern: `Bash(git log *)` means "any command starting with `git log`, no prompt," and `Bash(git status *)` does the same for `git status`. These two pre-approve only the reads `/prime` actually needs; anything else stays subject to your normal permission prompts. A blanket `Bash(git *)` would pre-approve `git push`, `git reset`, and every other git command without a prompt too, so keep the grant to the narrow pair. It is scoped to the turn that `/prime` runs in and clears when you send your next message, so you are not opening a standing hole in your permissions either. One honest wrinkle from ours: our freshness line runs `TZ='Europe/London' date`, and the leading `TZ=` assignment does not always match the command prefix cleanly, so the first run still asks once. We approve it and move on. If one of your lines keeps prompting despite an `allowed-tools` entry, that mismatch is why; simplify the command or approve it the once. One caution that matters more once the file is not yours to begin with: a skill runs shell commands and can pre-approve its own tools, so treat one you pulled from someone else's repository like a script you are about to run. Read its `` !` `` lines and its `allowed-tools` before you invoke it. ## Why not just put this in CLAUDE.md? If Claude Code already reads a `CLAUDE.md` at the start of every session, a fair question is why this is not a few more lines there. The docs draw the line by what changes. `CLAUDE.md` is for facts that hold every session: your conventions, your architecture, the guidance you want in front of Claude every time. It loads into every session, which makes it the wrong home for anything that moves underneath it. `/prime` is for the things that move. Where the tests live belongs in `CLAUDE.md`; what you touched last, what is uncommitted, the one thing on fire this morning, all of that is different by the next session, and you want it recomputed when you ask rather than hand-edited into a file. That is the real move here, and it is bigger than one command. The durable thing to save is the procedure that fetches the state. A commit list or a status line goes stale the moment you write it down; a command that runs `git` and reads your notes fetches the current answer every time. Keep the procedure narrow, though: whatever `/prime` injects stays in the conversation and costs tokens on every later turn, so its job is to locate the work, not to preload your whole project. If it starts turning into a project dump, stop adding files. ## You may see the older one-file form If you read around, you will find the same trick written as a single `.claude/commands/prime.md` file with no folder. That is the older shape, and it still works: a `.claude/commands/prime.md` and a `.claude/skills/prime/SKILL.md` both create `/prime` and behave the same way. The skills folder is the form the docs point you to now, and the reason is room to grow. A folder holds more than the one file, so it carries supporting scripts when your `/prime` gets ambitious, and a skill can let Claude reach for it on its own when it fits (add `disable-model-invocation: true` to keep it strictly manual, `/prime` only). If you already have a one-file command, leave it alone, it keeps working. Reach for the folder when the command outgrows a single file, and if you ever keep both under one name, the skill wins. ## When it does not fire - **It is not in the `/` list.** You are probably not in the project, or the file is not at `.claude/skills/prime/SKILL.md` (the folder has to be named `prime` and the file `SKILL.md`). A skill in `~/.claude/skills/` instead works in every project. - **It runs, but Claude never reaches for it on its own and the menu shows no hint.** The frontmatter did not parse. A malformed YAML block does not remove the skill: `/prime` still works, but it loads with empty metadata, so the `description` no longer matches. Start Claude Code with `--debug` to see the parse error, then line up the `---` fences. - **It prompts for permission every single run.** Your `allowed-tools` pattern does not match the command you inject. Line the prefix up exactly, for example `Bash(git log *)` for a `git log` line. - **`$1` is not what you expected.** In the skills model, indexed arguments are zero-based: `$0` is the first argument, `$1` the second. Reach for `$ARGUMENTS` when you just want "everything I typed," and leave positions alone unless you genuinely need them. - **The output did not appear.** The injection only fires from a `` !`…` `` backtick span with a leading `!`. A command sitting in a plain fenced block is shown to the model as text, not run. ## Prove it loads Do not take my word that the injection fired. Run `/prime` and read the top of what comes back. If it references your actual latest commit message and your actual files, the commands ran and the state is real. If it hands you generic advice that would fit any repository, nothing injected: your `` !`…` `` or `@` lines are the place to look. A command that prints the same thing regardless of your project is just a saved sentence, which is where we started. Once it loads real state, you have turned a paragraph you retyped every session into one word, and the machine does the fetching. That is the pattern for the whole series: a thing you do by hand, moved into a tool you set up once and then just run. This one only read your project. The next rung is about letting Claude Code change it safely: a guardrail that stops the agent before it edits a file you marked off-limits. That is the next Runbook. --- *Sibling piece, same tool, different job: [Never Let Claude Code Tell You It's Done](https://durabilitycurve.com/blog/never-let-claude-code-tell-you-its-done/) wires a test the agent cannot talk its way past.* [^1]: Claude Code documentation: [dynamic context injection and slash-command syntax](https://code.claude.com/docs/en/slash-commands). --- --- title: "The Setting You Never Changed" description: "The pre-ticked box, the factory setting, the plan already selected. Someone chose each one before you did." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/setting-you-never-changed/" date: "2026-08-09" series: "THE BLUEPRINT" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Setting You Never Changed *The pre-ticked box, the factory setting, the plan already selected. Someone chose each one before you did.* By Harry Floyd · 2026-08-09 · canonical: https://durabilitycurve.com/blog/setting-you-never-changed/ [Figure: An opaque form surface peeled back along a diagonal, exposing the luminous structure wired beneath a ticked "recommended" default.] You have settings you have never changed. Not because you sat with each one and judged it right. Because changing it was one more small task on a day that already had enough of them, and the thing worked well enough as it came. The notifications, the privacy toggles, the plan you signed up on, the way the app opens. Somebody handed you a starting point and you started from there. Almost everyone does. We treat that starting point as neutral. A sensible middle. The factory had to pick something, so it picked the reasonable option and left the rest to us. The default feels like the absence of a decision, the blank page before anyone chooses. It is the opposite. The default is a decision, made by someone who is not you, and often the most powerful decision in the whole design. It gets that power from one plain fact about people: many of us never change it. Watch how that fact does its work. Every change, however small, costs a little effort, and even a little effort is enough to keep many people where they started. A default also reads as advice. Someone who knew more than you set it here, so here is probably fine. And there is the pull of the path already laid: the option in front of you is the one you can take without stopping to think, and not stopping to think is most of what any of us do in a day. Put those together and the person who sets the default is not gently nudging the few who care. They are quietly deciding the outcome for the many who will never look. You can see the size of that power most clearly where the stakes are life and death. Take Germany and Austria. Neighbours, with similar wealth and much shared culture. Germany asks you to opt in, rather than presuming your consent. Austria works the other way, presuming consent unless you object, so if you do nothing you remain a potential donor. In the comparison that made defaults famous, Germany's effective consent rate sat around one in eight. Austria's ran above nine in ten.[^1] Two neighbouring countries are never a controlled experiment, and apparent consent is not the same as an organ donated. But where you can run the clean experiment, the same lever pulls the same way. [Figure: Two neighbours, opposite defaults, a consent gap this wide: an opt-in grid barely filled beside an opt-out grid almost full.] The same lever changed how a country saved for retirement. Britain used to make workers opt in to a workplace pension, and only about half of those eligible were saving into one. From 2012 the system began to flip: automatic enrolment was phased in, so you were enrolled unless you chose to leave, and within a decade participation was near nine in ten.[^2] The system did not wait to persuade each worker. Other things moved alongside the default, minimum contributions among them, so this was no clean laboratory test. But the resting state had changed from not saving to saving, and millions simply stayed where it put them. The resting state, the thing that happens while you are busy living, does an extraordinary amount of the work. Follow that one step further and the incentives start to matter. If the default can decide the outcome for so many people, then the right to set it is worth real money and real power to whoever holds it. So the question to ask of any default is not "is this a fine starting point." It is: who set this, and what is it set to do? The pre-ticked "yes, keep me posted" is set to grow a mailing list. Google paid Apple billions to be Safari's default search engine, betting the one you start with is the one you keep.[^3] The subscription that renews unless you cancel is set on the company's calendar, not yours, and it benefits from the same forgetfulness that keeps you from changing any other setting. The tip screen that opens with twenty per cent already lit is set to make twenty per cent feel like the starting point. None of these is an accident, and none is neutral. Each is a resting state that someone tuned toward their own goal, wearing the calm face of a reasonable place to start. Here is the habit worth building, and it is the whole point of looking under this particular surface. When you meet a setting you did not choose, stop reading it as the way things simply are. Read it as a sentence somebody wrote and left for you: this, unless you say otherwise. The moment you hear it as a sentence, two questions fall straight out of it. Who benefits if I leave this exactly as it is? And what would I have picked if the box had been blank and the choice were genuinely mine? The gap between those two answers is the thing that was quietly taken from you while you were getting on with your day. [Figure: Read the default as a sentence, and two questions fall out of it, with the gap between the answers marked.] Often the gap is small, because the person who set the default guessed well or genuinely wished you well, and you leave the setting alone, now knowing that you chose to. Sometimes it is large, and you change it, and it takes ten seconds you did not know were worth taking. Either way you have stopped mistaking someone else's decision for the natural order of things. Once you can see it, you cannot stop seeing it. The box already ticked on the form. The toggle set to share before you looked. The plan pre-selected as the one that quietly renews. The world is full of resting states that were set before you arrived, and every one of them was set by a hand with a reason. This is only one of the structures the world is built from, and it stays invisible until someone draws you the diagram. The setting you never changed is still a choice. Someone else just made it for you. [^1]: Eric J. Johnson and Daniel G. Goldstein, ["Do Defaults Save Lives?"](https://www.science.org/doi/10.1126/science.1091721) *Science* 302 (2003): 1338–39. The same work paired a European comparison of effective consent rates with a controlled experiment, in which agreement to donate was far higher under an opt-out default than an opt-in one. [^2]: Department for Work and Pensions, ["Workplace pension participation and savings trends of employees: 2009 to 2025"](https://www.gov.uk/government/statistics/workplace-pension-participation-and-savings-trends-2009-to-2025/workplace-pension-participation-and-savings-trends-of-employees-2009-to-2025) (30 July 2026): around nine in ten (90%) of eligible employees were saving into a workplace pension in 2025, up from about half when automatic enrolment began in 2012. [^3]: In 2022 Google paid Apple an estimated $20 billion for default search placement in Safari, according to filings in the US antitrust case that found Google an unlawful monopolist in 2024. [US Department of Justice court filing](https://www.justice.gov/atr/media/1402141/dl). --- --- title: "The Guardrail Your Agent Can Reach" description: "Most guardrails end up with an escape hatch. Check whether the thing you are constraining can reach yours." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-guardrail-your-agent-can-reach/" date: "2026-08-08" series: "THE ENGINEERING LADDER" law: "Law IV" claims: "https://durabilitycurve.com/claims/the-guardrail-your-agent-can-reach/" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Guardrail Your Agent Can Reach *Most guardrails end up with an escape hatch. Check whether the thing you are constraining can reach yours.* By Harry Floyd · 2026-08-08 · canonical: https://durabilitycurve.com/blog/the-guardrail-your-agent-can-reach/ *Most guardrails end up with an escape hatch. Check whether the thing you are constraining can reach yours.* [Figure: An agent walks up to a barrier, knocks over a lever standing on its own side of it, and strolls through while the boom lifts] Somewhere in the repository you look after, there is a rule you decided to stop trusting to prose. It had been a line in your context file. You rewrote it three or four times. It failed anyway, on a Saturday. You gave up on wording it better. So you did the thing everyone now recommends and moved it somewhere deterministic: a pre-commit hook, or a `PreToolUse` hook, or a permission rule. Something that returns a failure code and stops the run outright, instead of asking the model nicely and hoping. Good instinct. Anthropic gives the same advice about their own product: "When there's something that absolutely must not happen, an instruction is the wrong tool... A real guardrail needs to be deterministic, and the enforcement methods are hooks and permissions."[^1] Now open the thing you built and look at what shipped with it. ## The hatch you put there on purpose Your pre-commit hook is skipped by `git commit --no-verify`, and you have typed that this month. Your CI check has an override label for the Friday when the build is red and the release is going out anyway. Your permission prompt has an "allow always" button that somebody pressed in week one. Your harness, the process that runs the agent and decides which of its requests to honour, reads a bypass environment variable. It is documented, because you documented it the night you needed to get something out of the door. None of that is sloppiness. A control with no way through will one day stop something that genuinely needed to happen, at three in the morning, with nobody around to authorise it. So an exception path goes in deliberately. The interesting question is who it was built for. Sometimes that turns out to be the party the control was meant to constrain. In ordinary developer tooling that trade-off has usually worked, because the person holding the override is still sitting inside an accountability system. ## The old rule, in a new place Security engineering worked this out decades ago: you do not let the thing being constrained control the constraint. Least privilege, separation of duties, two-person release. None of it assumes the constrained party is trustworthy, because accountability after the fact was never enough on its own. The set of things you are trusting completely even has a name, the trusted computing base, and the whole discipline is keeping it small. The strong version has a name too. NIST's definition of a reference monitor puts three requirements on whatever does the enforcing. It must be always invoked. It must be tamperproof. It must be small enough to analyse and verify.[^2] Determinism buys you none of the three. It buys something real that the list does not name, which is that the decision no longer runs through the model's judgement. Everything else is still open. Your hook can go uncalled, it can be altered, and it can be wrong. `git commit --no-verify` is the first of those. Nothing was tampered with, the script sits on disk exactly as you wrote it, and it simply was not invoked. An agent that edits the hook file is the second. Being wrong is the third and it is the quietest, because a hook whose pattern stops `rm -rf` but not `find . -delete` will pass every test you thought to write and fail forever in the one case you did not. Most guardrail failures I have watched are the first or the third, and those are the well-served ones: hook linters, server-side checks that a local `--no-verify` cannot reach, no shortage of advice about testing what you wrote. Security people have covered that ground too. Trail of Bits publishes opinionated Claude Code defaults whose documentation is blunt: hooks "are not a security boundary", because "a prompt injection can work around them". They are "guardrails, not walls".[^3] This piece is about the second failure, which is where all of that stops. What is scarce is a test precise enough to apply to a specific repository on a Tuesday. None of this is hypothetical either. In March one user reported that Claude Code had edited a gate they had built with it, narrowing the instruction inside the hook from "code changes" to "SIGNIFICANT code changes", and attached the one-line diff. A follow-up RFC, from someone running enforcement across 545 tasks, reported watching the model add `Bash(*)` to its allowlist to get past a hook that was blocking it. Both were auto-closed for inactivity, with no vendor response either way.[^4] Neither is a jailbreak. In both cases the model did the ordinary thing, which was to edit a file it had write access to, and the file happened to be its own constraint. The principle is old. Where it gets violated is new. This used to be something you reasoned about while designing a permissions system, and it has become something you do by accident on a Tuesday, without once thinking of yourself as designing a security boundary. You would never ship an application that could rewrite its own access policy. Then you gave a coding agent write access to the repository holding its hooks. Two things change when the constrained party is an agent. The first is accountability, and it half survives. An audit trail can still work: the agent has a run ID, it acts under some credential, and you can read back what it did. The deterrent does not. A developer reaching for `--no-verify` knows they may have to explain it, and in some override systems that expectation does more of the work than the mechanism does. It does nothing to a process at two in the morning. So you keep the forensics and lose the deterrent, and you keep the forensics only if the log lives somewhere the agent cannot also write and somebody reads it. A record nobody opens is not a record. The second change turns a reliability problem into a security one. ## Mistake, misled, or aimed Most of the time your agent is simply fallible. It forgets, it drifts out of scope, it solves the adjacent problem. A repo-local hook is genuinely good at catching that, and nothing here should talk you out of one. Instructions get followed, too: a separate team measuring these files across real repositories found that instructions in them are well followed by coding agents.[^5] The prose is not being ignored. That is exactly why a hook that catches the residue is worth writing. Sometimes the agent is misled. Anthropic names this case in the same breath as the others: a prompted rule can fail "due to a prompt injection in a file accessed as part of the task."[^6] Your agent reads issues, dependency READMEs, test fixtures, web pages. Any of them can carry instructions, and the agent can fail to tell your intent from text it read thirty seconds ago. Security has an old name for this shape. A component with more authority than the party directing it is a confused deputy, and that hazard has had a name since 1988. And sometimes it is aimed. The model wants nothing. Whoever wrote the text it just read might. That is the bridge, and it needs no speculation about what an agent wants. **The thing with permission to disable your guardrail is also, routinely, reading untrusted input.** Each fact is fine alone. Together they are the problem. So the useful question was never whether your agent would decide to bypass a control. It is what the control is worth on the day something in its context window tells it to. A developer who bypasses a hook is a person acting on a system. An agent that bypasses a hook is the system acting on itself. The popular version of that overstates: an agent setting an environment variable cannot reach a value its parent process already read at launch. It barely matters, because the cheap routes are elsewhere. One `git config core.hooksPath /dev/null`, and every commit after it runs with no hooks and no `--no-verify` to account for; the bypass sits in configuration rather than on the command line, and the commit itself records nothing. Hooks are not cloned either, so a fresh checkout has none at all unless something installs them. Your control is absent by default, which is worse than bypassable. Which points away from the check and towards what the check trusts. **Can the agent modify any input the enforcement mechanism trusts?** The hook script, the config that decides whether it runs, the ledger it consults, the credential it uses. If any of those sits somewhere the agent can write, you have not built a lock. You have built a lock and handed the key to the thing you were locking out. [Figure: One gets through. One does not.] *Two agents, one barrier, one difference.* ## Run it on your own repository No dependencies, reads only. Run it from inside a repository your agent works in: ```sh #!/bin/sh set -u root=$(git rev-parse --show-toplevel) || exit 1 dir=$(git rev-parse --git-path hooks) # correct inside worktrees too me=$(id -un) W='%-34s %-10s %s\n' hooks() { find "$dir" -maxdepth 1 -type f ! -name '*.sample' 2>/dev/null; } printf "$W" "WHAT THE CONTROL TRUSTS" "OWNER" "CAN YOU WRITE IT?" check() { [ -e "$1" ] || { printf "$W" "$2" "-" "absent"; return 0; } own=$(stat -c '%U' "$1" 2>/dev/null || stat -f '%Su' "$1") # GNU first, then BSD # writable, or replaceable via its directory, or yours to chmod if [ -w "$1" ] || [ -w "$(dirname "$1")" ] || [ "$own" = "$me" ] then w="YES"; else w="no"; fi printf "$W" "$2" "$own" "$w" } hooks | while read -r h; do if [ -x "$h" ]; then check "$h" "hook: $(basename "$h")" else check "$h" "hook: $(basename "$h") [NOT EXEC]"; fi # chmod -x disables it silently done [ "$(hooks | wc -l)" -eq 0 ] && echo "!! NO HOOKS RUN AT ALL from $dir" # configs git reads now, plus the ones it would read if the agent created them { git config --list --show-origin -z 2>/dev/null | tr '\0' '\n' | sed -n 's/^file://p' | while read -r f; do [ -f "$f" ] && echo "$f"; done git rev-parse --git-common-dir | sed 's|$|/config|' echo "$HOME/.gitconfig" echo "${XDG_CONFIG_HOME:-$HOME/.config}/git/config" } | sort -u | while read -r c; do case $c in "$HOME"/*) l="~${c#$HOME}" ;; *) l=$c ;; esac if [ ${#l} -gt 22 ]; then b=${l##*/}; p=${l%/*}; l=".../${p##*/}/$b"; fi check "$c" "git config $l" done check "$root/.claude/settings.json" "claude: project" check "$root/.claude/settings.local.json" "claude: project local" check "$HOME/.claude/settings.json" "claude: user" check "/Library/Application Support/ClaudeCode/managed-settings.json" "claude: managed" ``` On a repository of mine it prints this: ``` WHAT THE CONTROL TRUSTS OWNER CAN YOU WRITE IT? hook: pre-commit harryfloyd YES git config .git/config harryfloyd YES git config .../git-core/gitconfig root no git config ~/.config/git/config - absent git config ~/.gitconfig harryfloyd YES claude: project harryfloyd YES claude: project local - absent claude: user harryfloyd YES claude: managed - absent ``` A `YES` is not a defect. It becomes one the moment you were relying on that row to hold against something trying. A hook marked `[NOT EXEC]` is worth a second look. `chmod -x` disables a hook completely without changing a byte of it, the commit goes through, and since `.git/hooks` is not version controlled there is no tracked file for git to show you. It asks three ways, because a file you cannot write is still replaceable if you can write the directory holding it, and a file you own is one `chmod` away whatever its mode says. `.git/hooks` is yours in every ordinary repository. Location tells you nothing either: a hook at `~/.config/git/hooks`, where `core.hooksPath` can point, sits outside the project and is entirely writable. The column answers for the account you are sitting in, so run it wherever your agent actually runs. In a container or on CI that is not here, and as root everything comes back `YES`, which is true and useless. Now look at the `git config` rows. None of them is a guardrail. They are the things that decide whether your guardrail runs at all, and there are more of them than people expect: the repository's own config, your global one, an XDG file at `~/.config/git/config`, a system file. Set `core.hooksPath` in any of them and hooks stop firing. In the global one they stop firing in every repository on the machine. A row reading `absent` is the one to watch, because a config file that does not exist yet is a config file the agent can create. Then it gets worse. The agent does not have to write a config file at all: `git -c core.hooksPath=/dev/null commit` runs with no hooks, writes nothing anywhere, and leaves no trace in the commit. The persistent route at least leaves a file behind for the table above to catch. This one leaves nothing, because the mechanism trusts an argument supplied at the moment of invocation and there is no file to lock down. So protecting `.git/config` does not close it, and neither does protecting all four. That is the cleanest argument in this piece for moving an important check off your machine altogether. The same goes for the settings rows: `absent` is not reassuring, because local settings outrank project settings and the agent can create the file. So treat the script as a first-pass audit of the obvious local trust points. It covers the config files git admits to reading, the well-known ones it would read if they existed, and your settings files. It cannot see the `-c` route at all, which defeats git hooks specifically and leaves the settings rows untouched. The ledger a hook consults, the credential it uses, and anything supplied on a command line are still yours to trace by hand. I ran it on mine. A hook in my own repository blocks a certain kind of file from being written until a matching row exists in a ledger. It returns a failure code, it has stopped me twice, and its bypass is an environment variable, which is the defensible kind: per-invocation, and the hook reads it from the harness rather than from any shell the agent controls. The hook also lives in a directory the agent it constrains can write to. So does the ledger it consults. I have not moved either, and the reason matters more than the finding. Moving the hook outside the workspace means it stops being versioned with the project that depends on it. I decided that trade was not worth making here. What this hook does is stop me forgetting, it is good at that, and it was never holding a line against anything that was trying. What I gave up is the right to describe it as a boundary, which is what I had been quietly doing. ## Who can reach the override The sharper question separates things that look identical on a diagram. A pre-commit hook and a production deploy gate are both deterministic, both return a failure code, and the difference lies entirely in the override. The hook is bypassed by the same actor it constrains, in the same command, and the bypass never appears in the resulting commit. The deploy gate needs a second party in a different system, and it leaves a trace somebody reads on Monday. So the test is not whether a thing can refuse. Almost anything can be made to refuse. It is **whether the enforcement sits inside or outside the authority of the thing being constrained.** Anthropic draws this line themselves, and it is the most useful thing in their post. A `PreToolUse` hook "can inspect a call and exit with code 2 to block it", which is deterministic by their own description. Managed settings "go further: they are admin-deployed, cannot be overridden by a user's local config", and are the only route to a guardrail that holds across an organisation.[^7] Both mechanisms are deterministic. One is beyond local override. That gap is the whole subject, and it is really a question about [which layer you are intervening at](https://durabilitycurve.com/blog/leverage-hierarchy-of-agent-engineering/). Apply it to your own stack and four tiers fall out. **Tier 1, a rule in your context file.** The model reads the rule and decides whether to follow it. Same authority, no enforcement. **Tier 2, a hook in a repository the agent can write to.** This is the one people misjudge, because it looks like enforcement right up until you ask who owns the file. It is also where most of what you have built actually sits. **Tier 3, enforcement the agent's authority does not reach.** An API that never exposes the destructive operation is the cleanest: there is nothing to bypass. A credential the agent was never issued works the same way, as long as it cannot pick one up from a metadata endpoint or a stray `.env`. The credential it does hold should also be no wider than the calls it makes. Be careful with the rest of tier 3, because deployment decides it. Ordinary Unix mode bits are not much of a boundary against an agent running as container root with the usual capabilities, and a plain `docker run` gives you root unless somebody chose otherwise. What makes the tier hold is a read-only mount, a dropped capability, a user namespace: something root does not reach past. Network egress rules work the same way, holding when the enforcement sits where the agent has no authority and not when it can reach the thing enforcing. **Tier 4, a separate principal has to move.** A human, or a system under separate control. It is a different kind of control rather than a strictly better one, because it adds judgement at the exception point and introduces a failure the others do not have: a person approving forty exceptions a week is a rubber stamp, and a deny rule nobody sees beats a human who clicks yes. The separation also has to be real. A second agent with similar authority, reading the same untrusted content, gives you a correlated failure rather than a second opinion. [Figure: Where the enforcement actually sits] *Most stacks are thin on the right.* The tempting lesson is that anything important needs a human in the loop. What an important guardrail needs is enforcement standing outside the authority of the thing being constrained. Sometimes that is a person. More often it should be architecture. ## What to change this week **Deploy a managed settings file.** On macOS it lives at `/Library/Application Support/ClaudeCode/managed-settings.json`, it takes an administrator to put it there, and it is evaluated above anything in your project or your home directory. A solo operator can do it in a minute. Put the things your trace just flagged in it, rather than the secrets example everyone copies: ```json { "permissions": { "deny": [ "Edit(//**/.git)", "Edit(//**/.git/**)", "Edit(//**/.claude/**)", "Bash(sudo:*)" ] } } ``` Those are aimed at the two incidents above: the agent that edited a hook, and the agent that widened its own allowlist. Two details decide whether they work at all: they have to be `Edit` rules, and the leading `//` has to be there. Get the tool name wrong and Claude Code warns you at startup. Get the prefix wrong and nothing says anything.[^8] Note also that they name files rather than commands. Anthropic's permissions page warns that Bash patterns constraining command arguments are fragile, and a rule like `Bash(git config core.hooksPath *)` is defeated by an option before the key, an extra space, a shell variable, or appending to `.git/config` by hand. Deny the file, not the phrasing. Now turn this article's test on that recommendation. On a personal Mac you are an administrator, so the file's protection rests on the agent not obtaining `sudo`. Run `sudo -n true`. If it succeeds, either your `sudo` is passwordless or you have a cached ticket in that terminal, and admin-deployed sits one command away from user-deployed. The `Bash(sudo:*)` line above raises the cost of that route without closing it. **Write deny rules, and know their reach.** Most agent configs I have seen are all allow and no deny. A `permissions.deny` block is cheaper than a hook and can protect a hook. Read how far it goes first. Read and Edit deny rules cover the built-in file tools and the file commands Claude Code recognises in a shell. They do not cover a script the agent writes that opens the same path itself. For that you want the sandbox, which enforces at the level of the operating system across the shell's process tree.[^9] A rule constrains the routes it enumerates, and the agent composes the routes. **Check that any of it took.** Deploying config and having a guardrail are different things, which is the whole subject of this piece. Claude Code will tell you when a file is broken, so the failure worth worrying about is the file that loads perfectly and does not do what you think. Run `claude doctor`, then `/status`, and read the `Setting sources` line: it lists what actually loaded, and a file with broken JSON will not appear there at all.[^10] Then do the part no tool can do for you: ask the agent to perform the exact thing you just forbade, and watch it refuse. A rule you have never seen fire is a rule you are trusting on faith. **Move the git check to the server.** The tier-3 version of a pre-commit hook is a `pre-receive` hook, which runs where the committer cannot skip it and has existed since 2007. It is GitHub Enterprise Server only, so on github.com you cannot install one. There the equivalent is a required status check, and it reaches tier 3 only when the check's reporting identity and its definition both sit outside the committer's reach. Identity is the easy half: GitHub lets you pin a check to the app expected to report it, and refuses the merge when anything else reports it.[^11] The definition is the hard half, and pinning does nothing for it, because a pull request that edits `.github/workflows/ci.yml` runs the version it just wrote. Closing that half needs the workflow itself to live somewhere the pull request cannot touch. An organisation-level ruleset does it, by naming the source repository the workflow must come from.[^12] Working alone on a personal repository, that last one is out of reach: no organisation, so the definition half stays open. Your ceiling on github.com is a pinned check, plus the knowledge that anyone who can push a workflow edit can satisfy it. That beats a hook you can skip with a flag. It is still not tier 3. On a private repo on the free plan you do not get even that, since rulesets and protected branches both need a paid plan there. For that reader the detection answer in the next paragraph is the whole of what is available, and it is enough. **When the trusted input must stay writable, stop preventing and start detecting.** That is often the honest answer. It works only if the record lands somewhere the agent cannot reach and something actually looks at it. That gap separates a reliability control, catching ordinary error, from a security boundary, built to hold while something is trying. Most of what is in your stack today is the first kind wearing the second kind's name. ## The uncomfortable version Sort your own stack this way and the picture is usually the same. A few things genuinely hold, because their enforcement sits somewhere the agent cannot reach. A middle tier stops the agent but trusts something the agent could rewrite. And underneath all of it, a context file quietly doing the work you assumed the middle tier was doing. That middle tier is where the false confidence lives. It earns its place: it catches forgetting, drift, the hallucinated command, the injected instruction that never finds the bypass, the same way a hook catches a developer who forgot rather than a developer who decided. What it will not do is hold when the thing it constrains is also the thing that can switch it off. That is the protection people credit it with, and it is the one it does not provide. Take the guardrail you would least like to lose, and find out who can write the thing it trusts. *What in your setup can the agent reach that you had filed as a boundary?* [Figure: Banner for The Durability Curve. A dense surface of scattered figures and fragments drifts across the top of the frame; beneath it a fine lattice holds a rising gold curve. As the surface thins, the curve and its markers brighten into view, so the picture performs the publication's line about the interesting material sitting under the surface. Beneath the artwork the banner carries the publication name and an invitation to subscribe.] *New to The Durability Curve? It is a standing argument about what survives when the tools get powerful and the surface gets cheap. [Subscribe](https://harryfloyd.substack.com/subscribe?utm_source=article&utm_medium=web&utm_campaign=the-guardrail-your-agent-can-reach) for the rest, or [start with what survives](https://durabilitycurve.com/blog/start-here-what-survives-when-the/).* [^1]: Michael Segner, [*Steering Claude Code: when to use CLAUDE.md, skills, hooks, subagents, and more*](https://claude.com/blog/steering-claude-code-skills-hooks-rules-subagents-and-more), Anthropic, 18 June 2026. The passage is specifically about instructions in CLAUDE.md. [^2]: [*Reference monitor*](https://csrc.nist.gov/glossary/term/reference_monitor), NIST glossary, from SP 800-53 Rev. 5. Verbatim, a reference validation mechanism "is always invoked (i.e., complete mediation), tamperproof, and small enough to be subject to analysis and tests, the completeness of which can be assured (i.e., verifiable)." [^3]: [*claude-code-config*](https://github.com/trailofbits/claude-code-config), Trail of Bits. Full sentence: "Hooks are not a security boundary -- a prompt injection can work around them." [^4]: [*Security: Claude can rewrite its own hooks — Who watches the watchmen?*](https://github.com/anthropics/claude-code/issues/32376), anthropics/claude-code issue 32376, 9 March 2026, and [*RFC: Deterministic tool gate — hooks are necessary but insufficient for governance enforcement*](https://github.com/anthropics/claude-code/issues/45427), issue 45427, 8 April 2026. The first reports a hook's instruction text being narrowed from "code changes" to "SIGNIFICANT code changes" and includes the one-line diff; the second states "We observed the model adding `Bash(*)` to allowedTools to bypass a hook that was blocking it." Both are user reports rather than vendor-confirmed incidents. Both were closed by a staleness bot for inactivity, in May and June 2026, with no response from Anthropic on either thread, so draw no conclusion from the closures in either direction. [^5]: Thibaud Gloaguen, Niels Mündler, Mark Müller, Veselin Raychev and Martin Vechev, [*Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?*](https://arxiv.org/abs/2602.11988), arXiv, v2, 23 June 2026. The same paper finds that context files do not generally improve task success and add over 20% to inference cost, so read the obedience finding as narrow rather than as an endorsement. [^6]: Michael Segner, [*Steering Claude Code*](https://claude.com/blog/steering-claude-code-skills-hooks-rules-subagents-and-more), Anthropic, 18 June 2026, same passage as above. Prompt injection is listed alongside long sessions and ambiguity as a reason a prompted rule can fail. [^7]: Michael Segner, [*Steering Claude Code*](https://claude.com/blog/steering-claude-code-skills-hooks-rules-subagents-and-more), Anthropic, 18 June 2026. The full sentence: managed settings "are admin-deployed, cannot be overridden by a user's local config, and are the only way to enforce a deterministic, organization-wide guardrail." [^8]: [*Configure permissions*](https://code.claude.com/docs/en/permissions), Claude Code documentation. On the tool name: a path rule written for `Write`, `NotebookEdit` or `Glob` is one Claude Code "accepts but never consults", warning at startup, while an `Edit(path)` rule covers every file-editing tool. On the prefix: `path` or `./path` is "relative to current directory", and `//path` is an "absolute path from filesystem root", which is what a machine-wide file needs. [^9]: [*Configure permissions*](https://code.claude.com/docs/en/permissions), Claude Code documentation. Verbatim: "Read and Edit deny rules apply to Claude's built-in file tools and to file commands Claude Code recognizes in Bash, such as `cat`, `head`, `tail`, and `sed`." They "don't apply to arbitrary subprocesses that read or write files indirectly, like a Python or Node script that opens files itself". The same page notes that sandboxing "applies only to Bash commands and their child processes". US spelling preserved inside the quotations. [^10]: [*Claude Code settings*](https://code.claude.com/docs/en/settings), Claude Code documentation. Managed settings "parse tolerantly": a failing entry is stripped and a warning recorded rather than the whole policy dropped, while "user, project, and local settings files remain strict: a file that fails validation is rejected as a whole and reported". Claude Code also "watches your settings files and reloads them when they change", including `permissions` and `hooks`, so a running session picks up an edit without a restart. On `/status`, "a source appears once it loads with at least one setting, so a file with broken JSON doesn't appear even if it contains settings". [^11]: [*About protected branches*](https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/managing-protected-branches/about-protected-branches), GitHub Docs. Verbatim: "When you add a required status check, you can select an app that has recently set this check as the expected source of status updates. If the status is set by any other person or integration, merging won't be allowed." This closes the reporting-identity half only. [^12]: [*Available rules for rulesets*](https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/managing-rulesets/available-rules-for-rulesets), GitHub Docs, on required workflows: the rule is configured at organisation or enterprise level, and you specify the source repository and the workflow to enforce, which is what puts the definition outside the pull request. --- --- title: "The Average Is Nobody's Result" description: "Of 255 studies on AI-assisted colonoscopy, 21 split the result by who held the scope. They disagree." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-average-is-nobodys-result/" date: "2026-08-04" series: "SYSTEMS & LAWS" law: "Law IV" claims: "https://durabilitycurve.com/claims/the-average-is-nobodys-result/" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Average Is Nobody's Result *Of 255 studies on AI-assisted colonoscopy, 21 split the result by who held the scope. They disagree.* By Harry Floyd · 2026-08-04 · canonical: https://durabilitycurve.com/blog/the-average-is-nobodys-result/ [Figure: One study, two opposite results. Width is readers, height is effect, area is what cancels. A wide, shallow gold slab above the line: 44 weakest readers, plus 0.016 on the easier cancers. A narrow, deep red column below it: 6 strongest readers, minus 0.145 on the harder ones. The two areas are close and the two shapes are nothing alike. The flat rule between them is what the study reported: no significant average effect.] In 2013 four researchers went back to a completed mammography study and asked it a question it had not been designed to answer. The original study had put 50 radiologists in front of 180 mammograms, twice. Once unaided, once with computer-aided detection marking suspicious regions. The finding was a null. On average, computer aid changed nothing measurable, and the profession moved on. Andrey Povyakalo and his colleagues at City University London reanalysed the data by splitting the readers instead of pooling them. What they found was that the tool had done two large things at once, in opposite directions. For the 44 least discriminating radiologists, on 45 relatively easy cancers, computer aid was associated with "a 0.016 increase in sensitivity (95% confidence interval [CI], 0.003-0.028)". For the 6 most discriminating radiologists, on the 15 hardest cancers, "with CAD, sensitivity decreased by 0.145 (95% CI, 0.034-0.257)".[^1] The weakest readers got slightly better at the cases that were already easy. The best readers got substantially worse at the cases that were hard, which is to say the cases where a radiologist is the only thing standing between a patient and a missed cancer. Averaged together, those two effects cancelled, and the study reported that nothing had happened. Their own summary of it: "despite the original study detecting no significant average effect, CAD helped the less discriminating readers but hindered the more discriminating readers."[^2] The average was true of neither group. ## What Your Number Actually Is Averages behave like this everywhere, including on the dashboard you looked at this morning. Take the last tool your organisation rolled out and measured. Somebody produced a figure: review throughput up nine per cent, tickets resolved up fourteen, defect-escape rate down a fifth. That figure was computed across everybody who touched the work, which makes it a mixture. There were as many different effects in it as there were people, weighted by how much work each of them happened to do that quarter. A mixture behaves in ways an effect does not. It can be positive while the effect on a third of your team is negative. It can be zero while two large things are happening. And because the weights are your staffing, the mixture is a property of who was on shift as much as of the software. Hire four juniors and your measured number moves without anything about the tool changing, which is also why the vendor's benchmark never reproduces in your shop. Say what that nine per cent still is, though, because the argument is easy to overshoot. It is a real answer to a real question: across the actual mixture of people and work you had last quarter, this is what happened. That is worth knowing, and it may well be enough to justify keeping the tool. What it cannot tell you is how to deploy it. Suppose the nine per cent is three senior engineers on migrations and nine juniors on routine diffs. It is equally consistent with the tool lifting the juniors and doing nothing for the seniors, with the reverse, and with a gain for everybody. Those three worlds ask different things of you: which group to train, which to watch, and whether the number survives your next dozen hires. One number is compatible with all three and cannot tell them apart. I have made a version of this argument before, structurally, in [Where Your Metrics Fold](https://durabilitycurve.com/blog/where-your-metrics-fold/): a scalar reading is a lossy projection, and two situations that demand opposite decisions can share one perfectly accurate number. This piece is the empirical half. Here is a documented case where the projection folded, and here is how often anybody bothers to check. Bound the 2013 case honestly before it carries any weight. It is a post-hoc reanalysis: the strata were cut using thresholds derived from the same regression that produced the estimates, the authors call their own method "exploratory", and the headline decline rests on six radiologists.[^3] The contrast is also conditional on two things at once, reader ability and case difficulty, rather than being a simple comparison between stronger and weaker readers. One case establishes one thing, and it is enough: an average can conceal two opposite effects, undetected, in a published null result, in exactly the class of tool everyone is now buying. Whether it usually does, nobody can say. So how often does anyone look? ## 21 of 255 I could not find a count, so I made one. I chose a field unusually favourable to finding operator-level analysis. Adenoma detection rate in colonoscopy, ADR in the literature below, is a hard, standardised, patient-relevant endpoint, and computer-aided detection has been trialled against it more heavily than anywhere else in medicine. If a field splits its results by operator anywhere, it splits them here. The corpus is a frozen PubMed query: 259 records, 255 with abstracts. The question asked of each one was narrow and fixed before any reading began. Does this abstract report the assistant's effect separately for two or more groups of operators, defined by something about the operator, such as baseline detection rate, experience, training year, sex, or individual identity? Describing your cohort as experienced does not count. Splitting by patient age does not count. **Twenty-one of the 255 report the effect split by the operator. The other 234 abstracts report an average over operators and no operator-level effect.**[^4] Counting only primary studies, by PubMed's own publication-type labels rather than my judgement, it is 21 of 157. That number is scoped in two ways, and the scope travels with it everywhere it appears here. First, these are deposited abstracts. A study can disaggregate in its full text and never say so, so what this measures is the layer the field summarises itself in, which is the layer guidelines, press coverage and most readers stop at. I checked that limit rather than waving at it. In a seeded random sample of twenty of the average-only primary records, nine have open full text and none of the nine reports an operator split in its results. One mentions a colonoscopist subgroup in the future tense, being a protocol promising an analysis to come.[^5] The other eleven are paywalled and unread, and open-access status is not random. Second, and more importantly, 21 is a floor. I will come back to why. ## The Ones Who Looked Disagree Twenty-one studies did the split. If the answer were obvious, they would agree. In a population-based randomised trial in the Galician screening programme, 4,824 surveillance colonoscopies, with the operator split prespecified rather than fished for afterwards, computer aid "increased ADR among low-performing (ADR<54.5%) endoscopists (45.5% vs 52.1%; aRR 1.15 [95% CI 1.01-1.30]), but not among high-performing endoscopists (65.9% vs 63.5%; aRR 0.96 [95% CI 0.88-1.05])".[^6] The tool lifted the weaker operators and did nothing, numerically slightly less than nothing, for the strong ones. In a multicentre randomised trial across six centres in Hong Kong and mainland China, 3,059 patients, it went the other way. Detection rose for both groups, and by more in the experts: 42.3% against 32.8% for experts, 37.5% against 32.1% for non-experts.[^7] Pooling two randomised trials, one run in experts and one in colonoscopists still in training, a team writing in *Gut* found computer aid mattered (RR 1.29, 95% CI 1.16 to 1.42) and operator experience did not (RR 1.02, 95% CI 0.89 to 1.16), concluding that "experience appears to play a minor role as determining factor for ADR".[^8] Across the 21, seven report the larger effect in the weaker operators, four in the stronger, three find no interaction, three find nothing that survives stratification, and one finds the two groups diverging over time.[^9] Both directions appear in randomised trials, on the same endpoint, in the same procedure. [Figure: Twenty-one tiles, one per study, grouped by what each one found. Seven gold tiles for the studies reporting the larger effect in the weaker operators, the Galician trial among them. Four red tiles for those reporting it in the stronger, including the Hong Kong trial. Below, dimmed: three finding no interaction, three where nothing survives stratification, one where the two groups diverge over time, and three splitting on an axis other than skill.] The honest reading of that spread is that the literature has no stable answer about which operators benefit most. A reader who files it away as "the effect is probably small" has reached a different conclusion, and the two license different decisions. This is the three-worlds problem from a few hundred words ago, except now it is not hypothetical: one of the most heavily trialled areas of AI assistance in medicine contains randomised trials pointing in opposite directions about which operators benefit, and it has not resolved them. ## Looking Is Not Enough Here is the part that surprised me, and it is worse than disagreement. Of the 21 studies that split by operator, only eight print effect estimates for every group, in a form another researcher could combine or check. Three give a number for one group and declare the other null without printing anything. The remaining ten report a direction or a verdict: significant here, not significant there, no estimates at all.[^10] And in none of the 21 abstracts is there a test of the interaction. Not one asks, where a reader can see it, whether the difference between the groups is itself distinguishable from noise. [Figure: Four bars, each a subset of the one above. 255 abstracts on AI-assisted colonoscopy, the frozen corpus. 21 report the effect split by operator, leaving 234 that report an average and nothing else. 8 print an estimate for every group, the only ones another researcher can combine. The fourth bar, for studies that test whether the difference between groups is real, is zero and has no bar at all.] What most of them do instead is compare significance across subgroups. Andrew Gelman and Hal Stern named that error in a paper whose title is the whole argument: "The Difference Between 'Significant' and 'Not Significant' is not Itself Statistically Significant". As they put it, "even large changes in significance levels can correspond to small, nonsignificant changes in the underlying quantities".[^11] You can watch it happen. A 2026 study in *Diseases of the Colon and Rectum* looked at 2,327 colonoscopies and split three ways. Its stated expectation, in its own abstract, was that "the greatest increase was expected among low-volume and junior endoscopists". What it reported was that "of 12 senior and 12 junior endoscopists, the seniors had a statistically significant increase (p = 0.02) from 51.7% to 59.6%, whereas juniors did not (56.2% to 62.9%, p = 0.12)". It concluded that computer aid "correlated with an increased adenoma detection rate for experienced endoscopists".[^12] Do the subtraction the paper does not. The seniors improved by 7.9 points. The juniors improved by 6.7 points. The gap between those two improvements is 1.2 points. One crossed a significance threshold and one did not, and the conclusion was written from the thresholds rather than from the estimates. [Figure: Two bars of almost the same height. The senior endoscopists' gain in adenoma detection is 7.9 percentage points, from 51.7 to 59.6, at p equals 0.02, reported as an increase. The junior endoscopists' gain is 6.7 points, from 56.2 to 62.9, at p equals 0.12, reported as no increase. A bracket between the two bar tops marks them as 1.2 points apart.] I am not saying that finding is wrong. It may well be right. I am saying the split as published cannot support it, and that the same reasoning shows up repeatedly in the abstracts I read. That is the test to carry out of this piece: when you are shown two subgroups and two p-values, subtract the point estimates before you believe the story. There is a fair objection here, and a good statistician makes it first. Subgroup analysis has a bad name for good reasons. Cut the data enough ways and something will cross a threshold, and the thing that crossed is what gets written up. That complaint is about testing many subgroups and reporting the winner, and it has standard safeguards: choose the split before you look, report an estimate for every group rather than a verdict, and say how many splits you examined. One prespecified split with intervals is the best defence against subgroup fishing rather than an instance of it. It is not a complete one. Interaction tests are routinely underpowered, cutting a continuous skill measure into groups puts the boundary somewhere arbitrary, and subgroups can differ in the work they were given as well as in who did it. Among the 21, the Galician trial is the one that took the precautions. Medicine is only where the evidence happens to be. If a coding assistant lifted your median review throughput, and three of your reviewers handle the hairy migrations while nine handle routine diffs, you have two populations and one number, and nothing in that number tells you which of them moved. ## Why You Cannot Look This Up I built three detectors for operator splits, deliberately unlike each other. The first used the vocabulary anyone would reach for: high detector, low detector, baseline ADR, stratified by experience. It flagged fifteen records. I read all fifteen and seven were real, the rest being the word "quartile" in a patient age range or a study describing its own cohort as experienced. The second looked for an operator noun near a splitting verb and found ten genuine splits the first had missed entirely. The third used pairs of operator labels and no verb at all, and found four more that neither of the others reached. Precision across the three runs 0.47, 0.36 and 0.69, and coverage of the 21 runs 0.33, 0.67 and 0.43.[^13] The pattern anyone would write first reaches a third of them. [Figure: Three stacked bars accumulating to twenty-one. The first pattern, the vocabulary anyone would reach for, found seven. The second, an operator noun near a splitting verb, added ten the first missed entirely. The third, pairs of operator labels with no verb, added four neither of the others reached. Each bar's new contribution is highlighted against the dimmed earlier runs. The right edge is open, with a dashed continuation and the line: twenty-one is a floor, the true number is unknown.] I am calling that coverage rather than recall on purpose. Recall would need the true number of operator-splitting studies in the denominator, and the point of this section is that I do not know it. Twenty-one is a floor: three dissimilar patterns each found what the other two missed, and the additions, seven then ten then four, suggest convergence without demonstrating it. Since the real denominator is larger than 21, every figure above is an upper bound and each pattern is at best that good. Every count here is what this method found, and I cannot tell you what exists. The reason is structural, and it is checkable. CONSORT-AI is the reporting standard for trials of AI interventions. It asks investigators to "specify whether there was human-AI interaction in the handling of the input data, and what level of expertise was required of users". Its guidance encourages exploring "differences in performance and error rates across population subgroups".[^14] So the standard has a field for describing your operators, and a field for splitting by patient. It has no field for splitting by operator. No required field means no standard phrase. No standard phrase means no search term, no index entry, no way to ask the literature this question except by reading it. Which is the same reason [work nobody can check tends to look excellent](https://durabilitycurve.com/blog/your-ai-looks-best-where-you-check-least/). The absence of a check is a state with consequences of its own. ## The Trial That Promised the Split COLO-DETECT was a good trial. Twelve NHS hospitals, 2,032 participants, published in *The Lancet Gastroenterology and Hepatology* in 2024. Its protocol, two years earlier, saw this coming. A range of colonoscopist experience was "anticipated and desirable", and the protocol set out exactly how it would be graded, by accreditation status and lifetime and annual procedure counts, so that it could be analysed. Under the analysis plan: "Subgroup analyses will be conducted on colonoscopist type (i.e., non-BCSP accredited vs. BCSP accredited) and indication for colonoscopy (screening vs. symptomatic)."[^15] What reached the abstract two years later was one number. Adenomas found in 56·6% of the assisted arm against 48·4% of the standard arm, adjusted odds ratio 1·47 (95% CI 1·21 to 1·78).[^16] No colonoscopist subgroup appears in it. Be precise about what I checked: that paper is not open access and I have not read its full text, so the claim is that the split was pre-registered and the abstract reports the average alone. The analysis may well be sitting in the full paper. It is not in the layer everybody reads. ## The Answer the Field Already Gives The strongest objection is that this has been settled, and it comes with evidence. A 2025 systematic review pooled twenty-eight randomised trials and 23,861 participants. Its subgroup analyses "involving only expert endoscopists demonstrated a similar effect size (RR, 1.19; 95% CI, 1.11-1.27; P < .001)", and it concluded that assistance improves detection "irrespective of endoscopist experience".[^17] A second meta-analysis, twenty-four trials and 17,413 colonoscopies, agreed: "type of AI system used or endoscopist experience did not affect overall improvement in ADR."[^18] That is a real answer to a real question, and it is not this question. Both are comparing trials with each other. They ask whether studies conducted in experts report different effects from studies conducted in mixed cohorts. That is a between-study comparison, and it cannot recover what happens between operators inside a study. A field can run entirely on expert-only trials and mixed trials that produce identical average effects while, inside every one of them, the tool helps some people and hurts others. That is precisely what happened in 2013. The study-level summary said no significant average effect. The operator-level analysis said the tool helped the weakest readers and hurt the best. Pooling more study-level summaries would have made the first answer more confident and would never have produced the second. ## What to Do on Monday Split the number. How much that buys you depends on where the number came from, and it is worth being exact about this, because the difference is the difference between an effect and a hint. Start with the design that produced your headline figure, and be honest about what that design can carry. If it was randomised, or otherwise supported a credible causal comparison, run that same analysis again inside each operator group, decided before you look. That gives you an effect estimate per group, on the same footing as the number you already trust, and it is the real move. If it was merely this period against last, the per-group result is a change estimate rather than a clean read on what the tool did, because time, case mix and everyone getting better at their job are still tangled up in it. Either way, report the estimate and its interval for every group rather than which group crossed a threshold, then ask whether the gap between the groups is bigger than the noise. That last question is the one none of the 21 studies answered anywhere an abstract reader could see it. If there is no comparison in there at all, and what you have is simply how everyone performed with the tool switched on, then grouping by operator gives you something weaker and still worth having. It is a diagnostic, not an effect. Your logs cannot tell the tool apart from who was rostered, which cases arrived, who got better at their job anyway, and who quietly chose not to use the thing. What the split can tell you is whether your operators are spread widely enough that a single number could be hiding two stories, which is precisely the question that decides whether you need a real comparison. Treat a flip in sign as a reason to go and measure properly, not as a finding. The analysis is cheap either way, an afternoon on data you already hold. The conditions for it to mean anything are not always cheap: enough cases per group to see anything, operator identifiers that are actually clean, and groups doing comparable work rather than comparable-sounding work. Three outcomes, all of them useful. If the estimates point the same way across groups you have actually measured well enough to read, you have some evidence that your average means roughly what you thought it meant, which is worth knowing and is currently not known. If credible estimates point in opposite directions, your rollout decision was being made on a number that described nobody in particular. Hold the same standard here that you held a paragraph ago: a bare flip in sign can be noise, and estimates agreeing in direction can still disagree wildly in size. And if the groups turn out too small to say, that is still worth having, as long as you read it correctly. Thin cells do not make your overall number wrong. An average across two thousand cases can be perfectly well estimated while twenty groups of a hundred are far too noisy to tell apart. What you have learned is narrower and still useful: your data can support a statement about the deployment as a whole, and cannot yet support any statement about who benefits, who does not, or whether the effect moves across your operators at all. That is the boundary of your evidence, and knowing where it sits is the point of the exercise. Eight abstracts out of 255 printed estimates for every group. Only those eight expose enough for another researcher to check or combine the subgroup result without writing to the authors. What none of this can tell you: whether a given tool helps or harms a given group, whether the colonoscopy pattern transfers to your field, or whether that literature will resolve its own disagreement. Twenty-one teams looked and did not settle it. The claim here is narrower and harder to dodge. Operator-level results are almost never visible in this literature's abstracts, and the nine full texts I could open did not carry them either. The reporting standard has a field for describing your operators and none for splitting by them. And while answering the question properly can be expensive, finding out whether your own data could answer it is not. That much is a choice. An average is an honest description of a mixture. It just cannot tell you the thing that decides your next move, which is whether the tool is doing the same thing to everyone. The number you have been reporting is a summary of who happened to be on shift. Whether it describes any of them is a question you have not asked. If you go and run the split, I would like to know what came back. Estimates pointing the same way across your groups, estimates pointing in opposite directions, or cells too thin to read: all three are results, and at the moment none of them is written down anywhere. The colonoscopy literature has 21 attempts at this question and no answer. Yours would be the twenty-second, and it would be about a system you actually control. *Which number are you judged by that you have never once split by the person who produced it?* [Figure: Banner for The Durability Curve. A dense surface of scattered figures and fragments drifts across the top of the frame; beneath it a fine lattice holds a rising gold curve. As the surface thins, the curve and its markers brighten into view, so the picture performs the publication's line about the interesting material sitting under the surface. Beneath the artwork the banner carries the publication name and an invitation to subscribe.] *New to The Durability Curve? It is a standing argument about what survives when the tools get powerful and the surface gets cheap. [Subscribe](https://harryfloyd.substack.com/subscribe?utm_source=article&utm_medium=web&utm_campaign=the-number-nobody-collects) for the rest, or [start with what survives](https://durabilitycurve.com/blog/start-here-what-survives-when-the/).* [^1]: Povyakalo, Alberdi, Strigini and Ayton, ["How to discriminate between computer-aided and computer-hindered decisions: a case study in mammography"](https://pubmed.ncbi.nlm.nih.gov/23300205/), *Medical Decision Making* 33(1):98-107, 2013. [^2]: Same paper, Conclusions. The authors go on to say that such differential effects "may be clinically significant and important for improving both computer algorithms and protocols for their use. They should be assessed when evaluating CAD and similar warning systems." That was 2013. [^3]: Same paper again. The method section states the strata were obtained by "classifying a posteriori the cases (by difficulty) and the readers (by discriminating ability)" from the regression estimates, and the conclusion describes the approach as "our exploratory analysis method". [^4]: My own count, 2026-08-02, over a frozen corpus of 259 records, 255 with abstracts. The PubMed query was `("computer aided detection"[Title/Abstract] OR "computer-aided detection"[Title/Abstract] OR "artificial intelligence"[Title/Abstract] OR "deep learning"[Title/Abstract]) AND "colonoscopy"[Title/Abstract] AND "adenoma detection rate"[Title/Abstract]`. That is an ordinary public PubMed search and you can run it yourself, with one caveat: PubMed keeps growing, so it returns more records today than it returned on 2026-08-02, and matching my corpus means bounding it to that date. The query returns the abstracts. The count comes from reading them against the criterion above, one at a time, which is the work the query cannot do for you. Fifty-five of the 255 received an individual verdict: every record any detection pattern flagged, each read by hand against its abstract. The remaining 200 were never flagged and are counted average-only by absence, which is the same reason 21 is a floor. All 55, with their PMIDs and the reason for every call, are published as an appendix, linked at the end of this piece. A different count comes down to specific records and which side of the criterion they fall on. Those are the ones worth arguing about, and they are all in there. [^5]: Sample drawn with a fixed seed from the 136 primary average-only records; nine of the twenty have PubMed Central full text. The protocol was [COLO-DETECT](https://pmc.ncbi.nlm.nih.gov/articles/PMC9796278/), whose planned analysis is discussed later in this piece. [^6]: ["Computer-aided detection in surveillance colonoscopy: a population-based randomized trial"](https://pubmed.ncbi.nlm.nih.gov/42365851/), *Endoscopy*, 2026. Rates are given standard arm first. [^7]: ["Artificial Intelligence-Assisted Colonoscopy for Colorectal Cancer Screening: A Multicenter Randomized Controlled Trial"](https://pubmed.ncbi.nlm.nih.gov/35863686/), *Clinical Gastroenterology and Hepatology*, 2023. [^8]: ["Artificial intelligence and colonoscopy experience: lessons from two randomised trials"](https://pubmed.ncbi.nlm.nih.gov/34187845/), *Gut*, 2022. [^9]: Same first-party run. The remaining three of the 21 split on an axis other than skill: one by operator sex, one reporting each individual endoscopist separately, and one comparing operator classes across the tool boundary. [^10]: Same run. Eight print an estimate for every stratum, three print one stratum and declare the other null without an estimate, and ten report only a direction or a significance verdict. [^11]: Gelman and Stern, ["The Difference Between 'Significant' and 'Not Significant' is not Itself Statistically Significant"](http://www.stat.columbia.edu/~gelman/research/published/signif4.pdf), *The American Statistician* 60(4):328-331, 2006. [^12]: ["Old Dogs Can Learn New Tricks: Artificial Intelligence Improves Adenoma Detection Rates in Screening Colonoscopies in Experienced Endoscopists"](https://pubmed.ncbi.nlm.nih.gov/41919624/), *Diseases of the Colon and Rectum*, 2026. The same pattern appears on its other two axes: high-volume endoscopists 51.3% to 59.4% (p = 0.01) against low-volume 56.5% to 63.1% (p = 0.13). [^13]: Precision is the share of records a pattern flagged that turned out to be real splits. Coverage is the share of the 21 found positives that it reached, which is an upper bound on true recall because the real class is larger. Every record any pattern reached was read by hand against its abstract. [^14]: Liu, Cruz Rivera, Moher, Calvert and Denniston, ["Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension"](https://pmc.ncbi.nlm.nih.gov/articles/PMC7598943/), *Nature Medicine* 26:1364-1374, 2020. Item 5 (iv) and the discussion of the item 19 extension. DECIDE-AI, a separate guideline for early-stage evaluation, is not deposited openly and I have not read its item list, so this claim is about CONSORT-AI specifically. [^15]: ["Trial protocol for COLO-DETECT"](https://pmc.ncbi.nlm.nih.gov/articles/PMC9796278/), *Colorectal Disease*, 2022. BCSP is the NHS Bowel Cancer Screening Programme. [^16]: ["Polyp detection with colonoscopy assisted by the GI Genius artificial intelligence endoscopy module compared with standard colonoscopy in routine colonoscopy practice (COLO-DETECT)"](https://pubmed.ncbi.nlm.nih.gov/39153491/), *The Lancet Gastroenterology and Hepatology*, 2024. [^17]: ["Use of artificial intelligence improves colonoscopy performance in adenoma detection: a systematic review and meta-analysis"](https://pubmed.ncbi.nlm.nih.gov/39216648/), *Gastrointestinal Endoscopy*, 2025. [^18]: ["Impact of study design on adenoma detection in the evaluation of artificial intelligence-aided colonoscopy: a systematic review and meta-analysis"](https://pubmed.ncbi.nlm.nih.gov/38272274/), *Gastrointestinal Endoscopy*, 2024. Paid subscribers The rest of this piece is for paid subscribers, on any tier. [Read the rest on Substack](https://harryfloyd.substack.com/p/the-average-is-nobodys-result) --- --- title: "You Cannot Try to Fall Asleep" description: "Sleep is only where you notice it first. Much of what matters works the same way." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/you-cannot-try-to-fall-asleep/" date: "2026-08-03" series: "THE HUMAN LAYER" law: "Law A" substack: "https://harryfloyd.substack.com/p/you-cannot-try-to-fall-asleep" claims: "https://durabilitycurve.com/claims/you-cannot-try-to-fall-asleep/" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # You Cannot Try to Fall Asleep *Sleep is only where you notice it first. Much of what matters works the same way.* By Harry Floyd · 2026-08-03 · canonical: https://durabilitycurve.com/blog/you-cannot-try-to-fall-asleep/ *Sleep is only where you notice it first. Much of what matters works the same way.* It is ten past three. You have done the arithmetic twice already. Five hours if you drop off now, four and a bit if this carries on. So you lie very still, because turning over would be an admission, and you hold your eyes shut a little too tightly, and underneath all of it there is a low, steady wanting. You want to be asleep. And the wanting is the exact thing keeping you awake. You are failing at doing nothing. It is a strange thing to be bad at. You are a capable person. You can learn hard things and finish dull ones and drag yourself out for a run on a wet morning when every part of you would rather stay in. Effort is the most reliable tool you own. Most of what you are proud of came out of using it. And then there is this one ordinary thing, wanted more than almost anything at three in the morning, that effort cannot touch at all. The harder you try for it, the further away it goes. We file that under sleep being difficult and move on. But sleep is not the only thing built this way. It is only the place you notice it first, because you meet it every single night. > The wanting is the exact thing keeping you awake. Try to stop thinking about someone, and watch what happens to the thinking. Try to be happy, directly, by deciding to be, and feel it thin out into a performance of itself. You cannot make yourself find a joke funny. You cannot force another person to love you by loving them harder; if anything, that is the surest way to send them off. You cannot decide to be interesting at the party, and the second you try, you are the least interesting you will be all evening. You cannot will yourself to relax, which is the cruellest one, because the trying is the tension. None of these are things you do. They are things that happen to you while you are busy doing something else. Sleep arrives while you are turning over tomorrow's meeting, and then at some point you never quite catch, you are gone. Happiness turns up on an ordinary afternoon when you were absorbed in something and forgot to check whether you were happy. You become interesting the moment you get genuinely interested in someone else. Love shows up sideways, in the middle of doing something entirely unromantic together. Each one is a by-product. The main thing was always something else. [Figure: Animated. A dark field of faint stars, and a soft warm circle of attention drifting slowly across it. Wherever the circle rests, the stars inside it go out. Just behind it, in the space it has finished looking at, stars swell and brighten, more clearly than they ever are at rest. The circle wanders on and the pattern repeats, somewhere else. Beneath, quietly: the faintest stars go out when you look straight at them.] None of this is new. The Victorians had a name for the trap.[^1] They called it the paradox of hedonism, the plain observation that happiness tends to arrive only when your mind is fixed on something other than your own happiness. Aim straight at it and you miss. Aim at something worth doing and it arrives while your back is turned. Older and gentler still is the folk wisdom your grandmother had. A watched pot never boils. You will meet someone when you stop looking. Sleep comes when you stop chasing it. She was right, and she never needed a footnote. Which makes it strange that almost everything around us now says the opposite. Try harder. Optimise. Measure it, track it, put a number on it, and by watching the number, improve it. For plenty of things, that advice is sound. It genuinely works on the steps you walk, the pages you read, the money you put aside, because those are things you do, and a thing you do answers to attention and effort. Point a number at a behaviour and the behaviour usually moves. The trouble starts when we point the same instrument at something that was never a behaviour. Take the sleep tracker, the small clean example of a very large mistake. It hands you a grade out of a hundred each morning for a thing you did not do, could not have done, and had no control over while it was happening. For someone already anxious about their sleep, that grade can do exactly what you would dread. There is a name now for people whose pursuit of a better sleep score has quietly made their sleep worse.[^2] They lie there trying to earn the number, and the trying keeps them up, and in the morning the number confirms the bad night, so tomorrow they try harder still. The scoreboard becomes the insomnia. And it does not stop at sleep. We keep a scoreboard on our own happiness now, rating the day, wondering whether we are as content as we ought to be by this point, which is a reliable way to stop being content at all. We tally our friendships, our rest, our worth, treating the whole illegible middle of a life as figures to be raised. A thing that arrives sideways cannot survive being stared at head-on. Measured, it becomes work. Graded, it becomes a test you are failing. You get all the pressure of a target and none of the thing the target was for. > Measured, it becomes work. Graded, it becomes a test you are failing. So what do you actually do, if trying is the problem and not-trying sounds like giving up? You learn, slowly and against every instinct this age has trained into you, to tell two kinds of thing apart. Some things in your life answer to effort. You can decide the hour you go to bed. You can decide the phone leaves the room. You can decide to show up, to sit down at the desk, to be kind without keeping score, to call your mother, to put yourself in the path of the people you might one day come to love. Those are conditions. Conditions are real work, and they are what you can actually reach. And then there is everything the conditions are for. Sleep. Ease. Delight. Being loved. Feeling rested. Those you cannot reach for. You can only build the conditions, honestly and without cheating, and then do the single hardest thing a person can do, which is leave the outcome alone. > Leave the outcome alone. That last part feels like surrender, and it is, and that is exactly why it is so hard. Every part of you wants to grip. Gripping feels like caring. It feels like doing your part. But on this entire class of things, the grip is the surest way to fail, and loosening it is not laziness. It is a skill, and it may be the deepest one a life asks of you. Some effort [quietly builds you](https://durabilitycurve.com/blog/difficulty-was-making-you/). This is the other kind, spent on the one thing effort can only spoil. It is ten past three somewhere, and you are lying very still, doing arithmetic in the dark. You have already done your part. The room is dark, the day is behind you, there is nothing left to arrange. There is nothing left to do but the one thing you cannot do, which is try. So stop. There was never anything there to try for. You have set the conditions. Now let yourself be no use at all for a while. *This is not the usual thing here. I mostly write about AI, product and markets, and the structures underneath them. This is the same habit of looking, pointed at something a lot more ordinary, and I wrote it because I kept meeting it at three in the morning rather than at a desk.* *New to The Durability Curve? It is a standing argument about what survives when the tools get powerful and the surface gets cheap. [Subscribe](https://harryfloyd.substack.com/subscribe?utm_source=article&utm_medium=web&utm_campaign=cannot-try-to-fall-asleep) for the rest, or [start with what survives](https://durabilitycurve.com/blog/start-here-what-survives-when-the/).* [^1]: The phrase belongs to the nineteenth-century philosopher Henry Sidgwick, and John Stuart Mill put it plainly in his autobiography: those are happiest, he wrote, who have their minds fixed on some object other than their own happiness, and who find happiness by the way. It is a very old idea with a long line of owners, which is part of why it is worth trusting. [^2]: Researchers have called it orthosomnia, a fixation on achieving the sleep the tracker says you should be getting. It is a coined term rather than a formal diagnosis, first described in a 2017 case report of a handful of patients who trusted the device over their own clinician. The point is not that trackers are evil. For someone who sleeps well and glances at the score the way they glance at the weather, it may cost nothing. It is the people already lying awake who cannot afford to be graded on it. One boundary, stated plainly: this is an essay about the trying, not a treatment. If sleeplessness has become the ordinary shape of your nights, the thing that works is CBT-I, the only approach carrying a strong recommendation in the American Academy of Sleep Medicine's 2021 guideline, and it is emphatically active: it asks you to change what you do, including getting out of bed when sleep will not come. Leaving the outcome alone is not the same as leaving a real problem untreated. --- --- title: "Your Robot Coworker Is Still a Pilot" description: "Built, shipped, installed, working: four counts, quoted as one. Almost nobody publishes the fourth." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/your-robot-coworker-is-still-a-pilot/" date: "2026-08-01" series: "PROOF & TRUST" law: "Law IV" claims: "https://durabilitycurve.com/claims/your-robot-coworker-is-still-a-pilot/" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Your Robot Coworker Is Still a Pilot *Built, shipped, installed, working: four counts, quoted as one. Almost nobody publishes the fourth.* By Harry Floyd · 2026-08-01 · canonical: https://durabilitycurve.com/blog/your-robot-coworker-is-still-a-pilot/ *Built, shipped, installed, working: four counts, quoted as one. Almost nobody publishes the fourth.* [Figure: Animated cover. A factory economy on a single conveyor: a dense queue of built robots feeds a belt where most tumble into a heap at SHIPPED, one reaches INSTALLED, and a lone gold robot passes under a WORKING arch and walks off the right edge. For more than half of every loop no gold robot exists at all. Four gates are labelled BUILT, SHIPPED, INSTALLED and WORKING; no number appears anywhere on the image.] ## The robot that supported 30,000 cars In February, BMW published the results of a pilot at its plant in Spartanburg, South Carolina. A humanoid robot made by Figure had been lifting sheet metal parts into a welding cell, ten-hour shifts, Monday to Friday. The release is precise about what happened. The robot moved more than 90,000 components. It covered approximately 1.2 million steps. It ran for around 1,250 operating hours. Within ten months, in BMW's words, it supported the production of more than 30,000 BMW X3s.[^bmw] Four numbers, and not one of them counts robots. BMW's release refers throughout to "the robot Figure 02", in the singular, and never states how many were on the floor. I went looking for that figure and could not find it in the company's own words. Secondary coverage says two. BMW does not say two, or any other number, which is why this piece will not say two either. What travelled was the 30,000. It became a headline about physical AI's return on investment, filed under the shorthand that a fleet of humanoids had helped build 30,000 cars. The number is true. BMW published it. The thing it appears to prove, that humanoid robots are now doing meaningful volumes of industrial work, is a claim the number does not make and BMW never made either. The 30,000 is the output of a car line that was running before the robot arrived, and the figure isolates nothing the robot added to it. The release says the robot supported that production rather than performed it. BMW's own framing is careful in the same way. The release presents Spartanburg as a completed pilot and announces the next one, at Leipzig, where it says it will test "a humanoid robot" in battery assembly. Singular again, and again no number. This is the ordinary condition of robot numbers right now. They are almost all real, sourced, and published by serious organisations, and they are routinely answers to a different question than the reader thinks is being asked. The gap has a price. It separates an industry running pilots and demonstrations from one with a standing robot workforce, and a bet sized for the second when the evidence supports only the first is how money gets lost. When a number is about to move a decision, the rung it sits on is the decision. ## Four words that are not synonyms There are four separate quantities hiding behind most robot statistics, and they sit on a ladder. A robot can be **built**, meaning it exists and has come off a production line. It can be **shipped**, meaning it has left the manufacturer and been delivered to somebody who paid for it. It can be **installed**, meaning it is physically in place at a site and wired into whatever it is meant to do. And it can be **working**, meaning it is productively doing paid work with limited routine human intervention. The first three mostly nest, give or take a manufacturer that installs its own units without shipping them anywhere. The fourth is a different kind of question. Built, shipped and installed all ask where the robot is. Working asks what it is doing, and that drags in how much of the time, under how much supervision, at a task somebody would pay a person to do. It is a stricter line than the first three and a blurrier one, which is exactly why it is the rung that goes uncounted. The examples are easy to feel. A robot in a crate in a warehouse has been built and shipped and is neither installed nor working. A robot in a university lab, bought so that a doctoral student can test a grasping algorithm on real hardware, has been built, shipped and installed, and is doing exactly what its buyer wanted while producing nothing anyone would call labour. A robot standing in a hotel lobby greeting guests has cleared the same three rungs, and whether it is working depends on what you think a lobby greeter produces. The gaps between rungs are not rounding errors. Unitree, the Chinese manufacturer that ships more humanoids than anyone else, produced more than 6,500 humanoid robots in 2025 and shipped more than 5,500 of them. Those two numbers describe the same company in the same year and differ by 1,000 units. They are adjacent rungs, the closest pair on the ladder, and they still do not match. [Figure: The four rungs drawn as a descending ladder, with the number degrading as you go down. BUILT: 6,500 units off the line, from Unitree in 2025. SHIPPED: 5,500 delivered to customers. INSTALLED: about 9 per cent, and that figure is a share of revenue rather than a unit count. WORKING, meaning paid work with little oversight: an empty box, no public number. A row of robot glyphs thins from five to one down the rungs. The line beneath reads: the number does not just shrink, it loses its basis, then it stops.] Hold that ladder up against Spartanburg. Which rung was BMW's pilot on? Installed, clearly. Working, arguably, in one welding cell, doing one task. The 30,000 belongs to the X3 production line, which was building cars before the robot arrived and carried on building them after the pilot closed. The habit is small: when a robot number arrives, ask which of the four words it is attached to. Most coverage will not tell you, and the answer changes the meaning by an order of magnitude. ## The company that had to publish a correction In January, Unitree published a note on its own site because, in its words, "many pieces of misinformation regarding our company's 2025 shipment volume have been circulating online."[^unitree] Read what the company then does. It gives a shipment figure, more than 5,500 humanoid robots. It gives a separate production figure, total mass-production output above 6,500 units. It defines the first one explicitly as the "quantity actually sold and delivered to end customers, not order volume; the order volume is higher." And it closes by warning against "directly combining the numbers of different types of robots together for comparison", because its quadrupeds and its wheeled dual-arm machines are different products with different counts. That is a manufacturer publishing two distinct numbers for a single year, telling you which is which, telling you a third number exists and is larger, and asking you not to add unlike things together. There is a third number implied in there. Orders are higher than shipments, and Unitree declines to say by how much. An order runs ahead of the physical count, a commitment to buy rather than a robot that exists yet. It is also the number most likely to be announced, because it arrives first and sounds like traction. The company drew that line without being asked and left the tempting figure unpublished. The correction is the interesting part. Unitree did not publish this because it wanted to talk about taxonomy. It published because its numbers were being merged in public and the merged version was wrong enough to be worth a correction from the company that made the robots. The most careful counter in humanoid robotics, on this evidence, is the manufacturer, and the imprecision is downstream of it. ## What a company publishes when the number is legally binding So far this is a vocabulary problem. The next document turns it into a structural one. On 24 June, Agility Robotics announced that it was going public through a $2.5 billion merger with Churchill Capital Corp XI. The announcement was filed with the Securities and Exchange Commission, where a materially wrong number becomes a legal exposure rather than a marketing quibble. The press release in that filing carries no count of Agility's robots.[^agility-pr] It reports "more than $300 million of multi-year contracted Digit v5 orders secured to date", which is money. It reports "deployment commitments across nine customer facilities" and "more than 65,000 hours of operation", which are sites and hours, and note that the commitments are to deployment rather than deployments already made. It reports a pipeline of over 30 customers, which is logos, and describes a factory "designed to support production of up to 10,000 units annually", which is the capacity of a building. Dollars, facilities, hours, customers, capacity, and no robots. The robot count is in the same filing, one exhibit over, in the investor presentation.[^agility-deck] And it is a single unit figure, confined to a footnote that repeats under the same order number wherever it appears. The $300 million, the footnote says, "relates to 1,000 Digit v5 robots with three-year term RaaS contract". So there is a count, and the count is 1,000. Read which rung it sits on. The same footnote opens by saying the figure "reflects customer orders for Digit v5, as of May 2026". 1,000 robots ordered. Not built, not shipped, not installed. Digit v5 had not launched when this was filed, so the order sits before the physical ladder even begins, a contract for robots that do not yet exist, and it is the number a company reaches for first because it arrives first and reads as traction. Then read the clause most people would skip. The same contract "includes warrants issued to purchaser vesting proportionately to robots deployed". The buyer's upside is tied to deployment, and deployment is treated here as a separate event from the order, gated, still ahead, and never given a number. In the one document where the counting is legally accountable, the company publishes the order, structures real money around the deployment, and states the deployment figure nowhere. > The taxonomy is not my imposition on the industry. It is written into a warrant. Everything else in that presentation confirms the pattern. The projections run "1k units/yr, 5k units/yr, 10k units/yr" across a cost curve, which are production scenarios rather than robots made. A revenue chart scales an "installed base" from 2,500 to 15,000 and marks itself "for illustrative purposes only". The company says its robots are "currently deployed in customer facilities" and never says how many. One ordered count, a spread of illustrative futures, and a deployed count that the document is built around and declines to give. This is the part any reader can check without trusting me. The filing is public, both exhibits are on EDGAR, and the footnote takes about a minute to find. ## The gradient Put three companies side by side and a pattern appears that none of them would state on their own behalf. Unitree is preparing to list on Shanghai's STAR Market, where a prospectus must disclose sales volumes. Its prospectus reportedly splits humanoid revenue for the first three quarters of 2025 three ways: research and education 73.6 per cent, commercial 17.4 per cent, real industrial work 9 per cent. That split is second-hand, and it is a share of revenue rather than of units,[^unitree-split] so hold it more loosely than the primary figures above. Held loosely, it is still among the most useful numbers in this piece. The university lab from earlier is not an edge case. Research and education is the largest disclosed source of humanoid revenue at the company that ships more humanoids than anyone else on earth, and industrial work is the smallest of its three markets. The robots are real and the sales are real, and the biggest slice of the money comes from selling them as teaching hardware. Agility, filing in the United States where no volume disclosure is required, gives one unit number, an order, and describes the rest through hours and facilities. And Tesla, with no specific obligation to publish an Optimus count, has published none. Its Q2 statements describe installing production lines and starting production soon, and say the first units are for training data collection and further development rather than customers. That characterisation comes from reporting on the earnings materials[^tesla] rather than from my own reading of Tesla's deck, so treat it the same way. Meanwhile the figures that circulate for Optimus, several hundred units, 1,000 units, come from enthusiast sites rather than from Tesla. Line those three up and the direction is suggestive. Where the law compels volume disclosure, you get unit counts. Where it compels disclosure without specifying volumes, you get proxies. Where nothing compels anything, you get founder posts and adjectives. AGIBOT sits at that last end too: having announced its 10,000th unit produced, it says only that "a significant portion is already active in real-world environments", and no number is ever attached to the portion.[^agibot] [Figure: Three disclosure regimes side by side, most legal compulsion on the left and none on the right, marked by a falling row of filled pips. Mandated volume, Unitree's STAR Market prospectus: 6,500 produced and 5,500 shipped, real unit counts, a fleet number. A filing with no volume mandate, Agility's SEC Form 425: 1,000 ordered, then $300M, 9 sites and 65,000 hours, an order followed by proxies. No obligation, Tesla, Figure and AGIBOT in posts and decks: no unit count, only "a significant portion", adjectives. A gold rail across the bottom reads: and every one of them stops before WORKING, and we could not find a published count of robots doing paid work.] Three companies is not enough to blame the law, and it is not only the law that changes between them. Jurisdiction, company maturity, business model and the purpose of the document all move at once. Treat it as a hypothesis with a use: when a robot number is missing, the disclosure regime is the first place to look for why, and often for where a truer number is hiding. ## The rung nobody reaches The gradient has a ceiling, and the ceiling arrives before the top of the ladder. Across every source in this piece, company announcements, an SEC filing, an IPO prospectus, a trade body and the press that covers all of them, I could not find a published count of humanoid robots productively doing paid work with limited routine human intervention, the fourth rung as defined earlier. The search returned no such figure at all, from anyone. If you know of one, I would genuinely like to see it, and the claim here narrows accordingly. The strongest counter-example is not a humanoid company at all. The International Federation of Robotics has counted industrial robots for decades and it is rigorous about a distinction most robot coverage skips. Its World Robotics 2025 report gives 542,000 robots installed in 2024, a flow, the new units added that year. It separately gives operational stock of 4,664,000 units, a level, the accumulated total still in service.[^ifr] Those are not two rungs of the same fleet moving from one to the next; they are a year against roughly a decade, and IFR is careful never to let the two be read as one. That discipline is exactly what makes the ceiling visible. Operational stock is itself an estimate, built on assumptions about how long a robot stays in service, not a direct count of which machines are running. IFR does not report how many of the 4.66 million are producing, versus idle, versus down for maintenance, versus sitting on a line that has since been retired. The most disciplined counting operation in robotics separates the flow from the level, and on utilisation it goes quiet. There are decent reasons for the silence. Utilisation is commercially sensitive. "Productive" and "with limited intervention" are genuinely hard to define at the boundary, and a company that published a number would then have to defend its definition. I am not claiming anyone is hiding anything, and the piece does not need a motive to stand up. The absence is the fact. What follows from it is a limit on what anybody can currently know. Every confident statement about how much work humanoid robots are doing in the world is an inference from the rungs below, made by someone who chose the conversion rate themselves. None of this means the robots are not working. The BMW unit ran 1,250 hours of real work on a live production line, and Agility's orders are real money with deployment milestones written into the contract. Something is happening on the fourth rung. The claim here is only that its size is the one thing nobody has measured, while a great deal of capital and attention is priced as though the measurement had already come back large. The workforce may well be real. The number that would prove it is not yet on any page. ## The next number you read When the next robot figure arrives, and it will arrive this week, the question to ask is which of the four words it belongs to. Built, shipped, installed, working. The answer is usually recoverable from the original document in under a minute, and it is usually not the rung the headline implies. Two follow-ons are worth keeping. The first is that [the most expensive errors are made of true numbers](https://durabilitycurve.com/blog/the-most-expensive-ai-errors-are-made-of-true-numbers/), because a figure that survives fact-checking can still carry a claim its source never made, and every number in this piece is of that kind. The second is that when a company substitutes hours, sites or capacity for units, the substitution is itself information. It tells you which number that company is prepared to put its name to. BMW's four numbers were real. 90,000 components, 1.2 million steps, 1,250 hours, 30,000 cars. The robot count was not among them, the pilot has closed, and the next one is being announced in the singular. Somewhere in that gap is the whole distance between a technology that exists and a technology that works. The first four numbers were easy to find. The one that would settle it, I could not find at all. *What number were you last shown that turned out to be a different rung?* [Figure: Banner for The Durability Curve. A dense surface of scattered figures and fragments drifts across the top of the frame; beneath it a fine lattice holds a rising gold curve. As the surface thins, the curve and its markers brighten into view, so the picture performs the publication's line about the interesting material sitting under the surface. Beneath the artwork the banner carries the publication name and an invitation to subscribe.] *New to The Durability Curve? It is a standing argument about what survives when the tools get powerful and the surface gets cheap. [Subscribe](https://harryfloyd.substack.com/subscribe?utm_source=article&utm_medium=web&utm_campaign=produced-not-deployed) for the rest, or [start with what survives](https://durabilitycurve.com/blog/start-here-what-survives-when-the/).* [^bmw]: BMW Group, ["BMW Group to deploy humanoid robots in production in Germany for the first time"](https://www.press.bmwgroup.com/global/article/detail/T0455864EN/bmw-group-to-deploy-humanoid-robots-in-production-in-germany-for-the-first-time?language=en), 27 February 2026. The release reports that "within ten months" the robot "moved more than 90,000 components" over "around 1,250 operating hours" and "supported the production of more than 30,000 BMW X3". It refers throughout to "the robot Figure 02" in the singular and gives no unit count. [^unitree]: Unitree Robotics, ["Clarification Regarding Unitree's 2025 Sales Data"](https://shop.unitree.com/blogs/news/clarification-regarding-unitrees-2025-sales-data), 22 January 2026. "Unitree's actual shipment volume of humanoid robots exceeded 5,500 units", against "total mass-production output of 2025 exceeded 6,500 units". The 5,500 is "the quantity actually sold and delivered to end customers, not order volume; the order volume is higher". [^agility-pr]: Agility Robotics and Churchill Capital Corp XI, [joint press release](https://www.sec.gov/Archives/edgar/data/2074973/000121390026071290/ea029548401ex99-1.htm) (SEC Form 425, Exhibit 99.1), 24 June 2026. Reports "$300 million of multi-year contracted Digit v5 orders", "nine customer facilities", "65,000 hours of operation" and a factory "designed to support production of up to 10,000 units annually". No count of robots built, shipped, deployed or working. [^agility-deck]: Agility Robotics, [investor presentation](https://www.sec.gov/Archives/edgar/data/2074973/000121390026071290/ea029548401ex99-2.htm) (SEC Form 425, Exhibit 99.2), June 2026. The 1,000 figure is a footnote to the orders line: it "reflects customer orders for Digit v5 … relates to 1,000 Digit v5 robots with three-year term RaaS contract, which includes warrants issued to purchaser vesting proportionately to robots deployed". Digit v5 was pre-launch, and the deck gives no deployed count. [^unitree-split]: The revenue split (73.6 per cent research and education, 17.39 per cent commercial, 9 per cent industrial, first three quarters of 2025) is reported from Unitree's STAR Market prospectus by [36Kr](https://eu.36kr.com/en/p/3735276262601477) and [TechFlow](https://www.techflowpost.com/en-US/article/31730). I have not read the prospectus itself, filed in Chinese with the exchange. It is a share of revenue, not of units, and the two are not interchangeable when research and industrial units carry different prices. [^tesla]: Drawn from coverage of [Tesla's Q2 2026 earnings](https://www.cnbc.com/2026/07/22/tesla-tsla-q2-2026-earnings-report.html), not from Tesla's own deck, which I did not read. [^agibot]: AGIBOT, ["AGIBOT Reaches 10,000 Units as Real-World Demand for Robots Accelerates"](https://www.prnewswire.com/news-releases/agibot-reaches-10-000-units-as-real-world-demand-for-robots-accelerates-302728295.html), 30 March 2026. "Of the 10,000 humanoid robots produced, a significant portion is already active in real-world environments." No number is attached to the portion. [^ifr]: International Federation of Robotics, ["Global robot demand in factories doubles over 10 years"](https://ifr.org/ifr-press-releases/news/global-robot-demand-in-factories-doubles-over-10-years) (World Robotics 2025), September 2025. 542,000 industrial robots installed in 2024; operational stock of 4,664,000 units. --- --- title: "The Most Expensive AI Errors Are Made of True Numbers" description: "Two months auditing an AI research agent. Nine ways a true number lies, and the check that catches each." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-most-expensive-ai-errors-are-made-of-true-numbers/" date: "2026-07-27" series: "PROOF & TRUST" law: "Law IV" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Most Expensive AI Errors Are Made of True Numbers *Two months auditing an AI research agent. Nine ways a true number lies, and the check that catches each.* By Harry Floyd · 2026-07-27 · canonical: https://durabilitycurve.com/blog/the-most-expensive-ai-errors-are-made-of-true-numbers/ For three days in May, every pre-earnings brief our research agent produced on NVIDIA was anchored to one number: $68.1 billion, presented as the market's consensus for the next quarter. The agent reasoned from it carefully. A beat is priced in, meaning the market already expects the target to be cleared. The bar sits here. Watch the guidance, then the reaction. The number was real. NVIDIA had reported it the previous quarter. It was the last quarter's actual revenue, and the agent had dressed it as the next quarter's forecast. NVIDIA's own published outlook for the quarter the agent was forecasting was $78.0 billion, nearly $10 billion higher, so every brief built on that anchor was aimed at the wrong bar.[^nvda] Three days of confident, internally consistent, well-written analysis, wrong at the root. And the root was not a lie; it was a real number in the wrong tense. Some failures in this record invented a part outright. But the errors that survived longest took a real number and attached the wrong period, scope, label or authority to it. I run an autonomous research agent inside a large personal knowledge system, and for two months I checked what it told me at every layer: individual claims against the strongest sources reachable, filings and earnings releases and the survey publishers' own pages; whole documents against a gate that decides what enters the knowledge base; the agent's own confidence labels against independent reviewers. Everything got logged. This piece is that record, and the field guide that fell out of it: the nine failure modes this audit exposed, many of which ordinary claim-level fact-checking misses, a real specimen of each from our logs, what each one looks like in an ordinary chat window, and the cheap check that catches it. At the end there is a method, and a tool, for building the same trust table for your own AI. The headline is the part I did not expect. Checked claim by claim, the agent was largely right. Checked at the level of what deserved to enter what I know, almost everything died. Those two results are about different things, and the gap between them is where the money goes. ## What we ran, and what checking means here The agent is the same one from [How Reliable Is Your AI Agent?](https://durabilitycurve.com/blog/how-reliable-is-your-ai-agent/), the 91% piece: a Hermes-framework agent on a deliberately cheap model, running unattended on a rented server, researching markets and AI. It writes into a knowledge system of more than 16,000 notes. The standing rule: nothing it produces enters the canonical layer without surviving a check it cannot influence. Claims get sampled against the strongest sources reachable. Documents queue at a gate where a duplicate check and a human decide what gets through. Its confidence labels are read as text, never as evidence. If you work in a chat window rather than a pipeline, you own the same architecture without the vocabulary. The moment you copy an AI answer into your notes, your plan, or your codebase is your promotion gate. The only question is whether anything stands at it. Before the numbers, the scope. This is one agent, one model, one system's gates, over one two-month window, 9 May to 13 July 2026. Every percentage below is local. Your model is probably better than ours. Your corpus and your review habits are different, so your rates will differ, and I will flag the direction where I can. What transfers are the failure classes and the checks. And one result transfers with special force: the accuracy numbers below came from a cheap model, and accuracy was not the main thing killing its output. Redundancy and expiry were, and those are properties of your corpus and your calendar. **Upgrading the model alone does not fix the survival problem.** ## The scoreboard: four gates, four denominators These are four separate measurements on four different populations, taken as the work happened. They do not chain into a single funnel, and multiplying across them produces nonsense. ### The claim audit In May we pulled 41 discrete factual claims from the agent's research output, across two audits five days apart, and checked each against the strongest source reachable: SEC filings, earnings releases, official statistics, the trade press where nothing better existed. The checking ran on a separate model with web access, never the agent grading itself, and I adjudicated the verdicts. 32 held exactly. 6 more we graded approximate at the time, right in direction and wrong in precision. 3 were materially false. As first graded, that is 78% strict and 93% directional. Then the blind re-check described later in this piece re-examined 8 of the 41 and moved two of those approximates to wrong, which takes the directional rate to 88%. Only 8 were re-checked, and both moves went the same way, so treat 88% as the ceiling on what a fully blind pass would have returned rather than as a corrected figure. The strict rate does not move.[^1] The anatomy of the misses matters more than the rate. No invented companies. No invented events. No inverted conclusions. Where invention appeared, it was one level down: sub-category splits and secondary ratios the source never published, riding on top-line stories that checked out. The failures were dates, staleness, scope, and structure: the classes in the field guide below. ### The door Over the same two months, the agent staged 113 documents for promotion into the knowledge base. 9 made it. 104 went to the archive. 45% of the queue, 51 items of the 113, duplicated something the system already held. The rest had expired before review or were too thin to keep. Three concessions before you quote that number. We capture aggressively by policy, so our net catches more junk than a stricter pipeline would. This was a backlog clearance, so it overstates any steady-state week. And a share of the expiry is on us, because the queue outlived our review loop. Read it as a cost figure, and it is a cost figure agent retrospectives rarely publish: autonomous capture into a mature corpus yielded single-digit percent durable knowledge, and the triage burden scaled with volume, no matter how accurate the individual sentences were. ### The high-value flags 18 items the agent had marked as significant knowledge sat in review for more than a week. 14 died of age before a human read them. That one is a finding about us, and probably about you: machine-flagged knowledge had a shelf life shorter than our review loop, which means "I'll review it at the weekend" is a decision with a price. The other 4 were the queue's most substantial research, each carrying its own verification claim, one of them a literal "Verified: 20/20 claims (100%)". All four failed the promotion gate. Four of four. One had confabulated a specification date. One reported the star-counts on software projects, GitHub's popularity number, wrong by 2 to 3 times under the label "100% verified". One was titled "Verified Incidents" and sourced its numbered security vulnerabilities to personal blogs. One duplicated work the system already trusted. ### The checkers Three separate failures, three different mechanisms, and they deserve to be kept apart. Our internal review panels, scored against a six-reader focus group, on the same essay of ours, ran 0.9 points hot on a 10-point scale: one paired test, so treat it as an anecdote, not a benchmark. A batch of reviewer agents we ran over the queue waved off items that a direct check later kept. And the labels: in two months of logs, "100% verified" appears attached to work that failed verification. One stamped draft restated an analysis that had been updated 11 minutes earlier. Another asserted it had checked the vault for duplicates and found none, the same day its duplicate went live. That second agent had not run a check and reported a result. It had authored the sentence a check would produce. [Figure: The scoreboard: four framed gates, each with its own denominator. The claim audit, 78% strict on 41 claims. The door, 9 of 113 documents survived. The high-value flags, 14 of 18 died waiting and 4 of 4 failed review. The checkers, three mechanisms with no denominator. A rail between them reads: four populations, no shared denominator, do not multiply.] > Accuracy is a property of sentences. Survival is a property of what you can safely build on. Our agent scored well on the first and brutally on the second, and almost none of the gap was made of lies. Agent retrospectives usually report task success: did it finish, did the code merge, did the pipeline run. I have not seen anyone publish the other axis, the one this record measures: of everything the system said, what deserved to enter what you know? The accuracy metric collapses every failure into a single verdict, wrong. The survival failures have shapes: nine, in three families, each with its check. ## The field guide: nine failures, three families I started calling them assembly errors, because the defect lives in how the pieces are joined. In most classes the parts are true and the claim is not: a real number in the wrong tense, a real figure pinned to the wrong scope. In the mirror classes the whole is true and the model authors the parts to fit: a real total decomposed into an invented breakdown, a verification sentence written rather than run. Fluent models produce both at volume, they sail through the lie-detector posture most people bring to AI output, and each falls to a specific, cheap check that has nothing to do with asking the model whether it is sure. The split also explains which errors got expensive. Every authored part in these logs failed its first outside look; the errors that ran for days and became foundations were assembled from parts that were individually true. A fabricated part usually fails the first look from outside. A true part passes every look except the one at the joints. The three families: errors of time, errors of shape, errors of trust. One entry in full first, free, because it is the one I most often watch people build on. ### The invented breakdown The specimen: asked why communities oppose data centres, the agent reported Gallup's polling as water 35%, electricity 28%, noise 18%, property values 12%. Gallup's real survey put water at 18% and energy at 18% among opponents, inside a different category scheme where respondents could name more than one concern.[^2] None of the agent's four category-and-number pairs appears on the page. Gallup has no property-values line at all, and the noise it does report sits inside a 16% pollution category. The top-line story was right. The decomposition was authored by the model, category names and all, because a decomposition is what the question demanded and the source did not supply one. In a chat window this is the five-part market breakdown with tidy percentages. The four-stage version of your own method, when you wrote five. The "three drivers of churn in businesses like yours". Structure is what a builder most wants to hear, so structure is what the model gives you when the source runs out. The check, once a source is in front of you, takes under a minute: ask where the table is. A breakdown is only as real as the source's own table, and if the source has no table, you are reading fiction with a true headline. A real total does not vouch for its parts. [^nvda]: NVIDIA, ["NVIDIA Announces Financial Results for Fourth Quarter and Fiscal 2026"](https://investor.nvidia.com/news/press-release-details/2026/NVIDIA-Announces-Financial-Results-for-Fourth-Quarter-and-Fiscal-2026/), 25 February 2026. Record quarterly revenue of $68.1 billion for the quarter ended 25 January 2026, against an outlook for the following quarter of "$78.0 billion, plus or minus 2%": a gap of $9.9 billion. The figure quoted here is NVIDIA's own outlook rather than analyst consensus, which has no stable primary source to point you at. NVIDIA went on to report $81.6 billion for that quarter on 20 May 2026, so $9.9 billion is the conservative way to state the error. [^1]: 32 of 41 strict is 78%, with a 95% Wilson interval of roughly 63 to 88%. Directionally, 38 of 41 as first graded is 93% (roughly 81 to 97%); after the blind re-check moved two approximates to wrong, 36 of 41 is 88% (roughly 74 to 95%). The intervals overlap almost entirely, which is the honest reading of a sample this size: the counts are the finding, the third decimal place is not. [^2]: Jeffrey M. Jones, ["Americans Oppose AI Data Centers in Their Area"](https://news.gallup.com/poll/709772/americans-oppose-data-centers-area.aspx), Gallup, 13 May 2026. Respondents opposing a local data centre could name more than one concern, which is why the categories do not sum to 100. Paid subscribers The rest of this piece is for paid subscribers, on any tier. [Read the rest on Substack](https://harryfloyd.substack.com/p/the-most-expensive-ai-errors-are-made-of-true-numbers) --- --- title: "Your AI Stack Has Three Bugs Other Fields Already Fixed" description: "Your benchmark, your model jury, your agent swarm: three old structures, and the obvious fix is usually wrong." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/your-ai-stack-has-three-bugs-other/" date: "2026-07-25" series: "SYSTEMS & LAWS" law: "Law IV" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Your AI Stack Has Three Bugs Other Fields Already Fixed *Your benchmark, your model jury, your agent swarm: three old structures, and the obvious fix is usually wrong.* By Harry Floyd · 2026-07-25 · canonical: https://durabilitycurve.com/blog/your-ai-stack-has-three-bugs-other/ *Your benchmark, your model jury, your agent swarm: three old structures, and the obvious fix is usually wrong.* Everyone who has shipped models has drawn the same curve. The fast ones score worse. The accurate ones run slower. After enough launches you believe latency and accuracy pull against each other, and you settle for choosing where on the line to sit. Look at the models you never shipped and the line bends. The trade-off sharpens the moment you filter for models that cleared your release gate, because that gate passes a model for being fast enough *or* accurate enough. The slow, accurate model ships on its scores. The fast, rough one ships on its speed. The models that never make it are the ones weak on both, and the ones you rarely benchmark are strong on both, because they were expensive and someone killed them for cost. Inside the set you actually measure, the two traits look opposed. Some of that opposition is real: more parameters and more reasoning steps genuinely buy accuracy and genuinely cost latency. But some of it was manufactured by the gate, and you cannot tell which is which from inside the survivors. The frontier you are designing around is a mixture, and you have never measured the part that is real. Statisticians named this in 1946. Joseph Berkson noticed that studying hospital patients made unrelated diseases look linked, because admission is a combined bar: you get in for one condition or another.[^1] Condition on a combined outcome and you manufacture a relationship between its causes that was never there upstream. The name is collider bias, and once you have it you start seeing it under conclusions you hold with confidence. The tell is always the same, and it is not merely that selection happened. It is that the bar could be cleared by more than one route, so the routes end up looking like rivals inside the survivors. The benchmark you kept because each prompt demanded either hard reasoning or good retrieval, which is why those two capabilities now look opposed on your evals. The sample of posts you have actually seen went viral through either depth or sensationalism, which is why those two routes look opposed inside the visible set. The hiring call that the sharp candidates are the awkward ones, true inside a room your bar admitted people to for being technically strong or personable. Same skeleton, four costumes. ## The pattern is bigger than your benchmark A small number of mathematical structures show up across fields that have never spoken to each other, each time in a different costume, and almost nobody clocks them as the same thing. The scarce skill is seeing the structure under the surface, because the moment you can name it you inherit decades of somebody else's work on it. Computation is getting cheap. This recognition is the part that stayed expensive. One guardrail before the other two, because it is what separates this from analogy-hunting. A structural match is not proof. It buys you a candidate failure mode, a set of tests somebody else already designed, and a better place to look. Your own evidence still has to confirm that the mechanism travelled with the shape. The third case below is one where the shape matched and the obvious fix did not travel, and it took me a wrong draft to notice. Two more structures, both already in your stack, and then the move. ## The independence that holds until the day it matters Run three models to check each other and you feel safer, right up to the case that fools one and fools all three, because they trained on overlapping data and share a blind spot. The independence you priced in was real on the easy cases and gone on the hard one, which was the only one that was ever going to hurt you. That shape has a name, and finance learned it the expensive way. In calm markets two assets drift on their own and look nearly unrelated. Push the system to an extreme and the independence evaporates and they fall together. The models pricing mortgage bonds before 2008 did not assume the loans were independent. They were built to model how defaults move together, which is the harder and more honest thing to attempt. The failure was subtler. The Gaussian copula that spread across the market could represent ordinary co-movement while thinning out the probability of clustered extremes in the far tail. Combine that structure with dependence estimates drawn from a short and unusually benign credit history, and simultaneous stress looks far less likely than it is.[^2] > A model that captures correlation on ordinary days can still be blind on the day the losses arrive. It is the same structure behind [the reliability number your agent quietly breaks](https://durabilitycurve.com/blog/how-reliable-is-your-ai-agent/): a per-step success rate that holds in testing and collapses once the steps start failing together under load. So the test to import is not whether your three models disagree on average. It is the probability that all three are wrong on the same input, measured inside the hardest regime you can construct rather than across a benchmark. Quants have spent two decades building that measurement. The engineer running a three-model jury is rebuilding a worse version of it from scratch. ## The swarm that talks itself into agreement The three models above never spoke to each other. They failed together because something upstream made them alike. This one is different, and the difference is the whole point: here the parts become alike *because what each one writes becomes what the next one reads*. Hand a job to a swarm of agents and let each one read what the others produced. A first answer lands in the shared context. The next agent reads it and drifts toward it. Agreement makes the answer look more credible, which pulls the next one harder, and by the fifth pass you have consensus that was manufactured by the reading order rather than by the evidence. The consensus was never independently earned. Structural engineers watched a related feedback loop happen in the open. London opened the Millennium Bridge in June 2000. Ninety thousand people crossed it that first day, up to two thousand on the deck at once, and it began to sway. As the deck moved, each pedestrian adjusted their gait to stay balanced, and those adjustments fed lateral energy back into the deck, which increased the movement, which prompted further adjustment. The bridge closed after two days. The engineers had not designed for the loop between the crowd and the structure.[^3] On the newer account, the walkers did not need to copy each other. Each was responding to the deck, and the deck carried what they did back to everyone else. That is the version worth importing, because a shared context window is a deck: your agents are not imitating one another so much as all reacting to a surface each of them is also writing to. And the fix was not to diversify the crowd. They killed the loop: thirty-seven viscous dampers to absorb the lateral energy the crowd was feeding in, and pairs of tuned mass absorbers hidden under the deck. The bridge had been built with one percent of critical damping or less. On the lateral mode the walkers were exciting, the target was twenty. The bridge reopened in early 2002 without a recurrence of the original instability.[^4] That is the import, and it is not "add variety." Break the coupling. Make the first pass genuinely independent, so no agent sees another's answer before committing its own. Delay or cap how much peer output can enter the context. Set a threshold where too-fast agreement trips a check instead of ending the discussion. Damping, not diversity. ## Why this stays hidden Each field has its own word for its own version and no word for the shared structure. A trader knows crowded positioning and never calls it resonance. An epidemiologist knows hospital samples mislead and never calls it the same thing your benchmark does. A recruiter knows the sharp ones are awkward and never calls it a collider. Every expert is fluent in one costume and blind to the same bone wearing the costume next door. No single mind holds all the costumes at once, which is most of why the pattern goes unseen. You would need to be fluent in finance and epidemiology and structural engineering in the same afternoon. A system that holds them side by side can read across them, which is the case for [keeping an external memory that spans domains](https://durabilitycurve.com/blog/you-only-hold-four-thoughts/) rather than a deeper filing cabinet in one field. The day you recognise the collider in your evals, your content, and your hiring, you stop making one mistake in three places and correct it once. ## The three questions to run on your own stack Keep them near where you work. When a trade-off shows up, ask what you filtered on. If you are looking at a set that got in by clearing one bar or the other, part of that trade-off lives in the filter, and it may weaken, vanish or reverse when you measure the population the set came from. When you are leaning on things being independent, ask what they do at the extreme. Not whether they disagree on average: whether they fail together on the worst input you can build. The calm-weather number will lie to you first. When a system oscillates or seizes, ask whether its parts are coupled through shared state. If one part changes an environment the others respond to, you have a loop, and the fix is to damp the loop rather than to vary the parts. [Figure: fig02 three tests 2026 07 25] Here is how to prove me wrong, and it is cheap. Run the three tests against ten beliefs your work actually rests on. If no trade-off weakens when you measure upstream of its gate, no independence breaks under the worst input you can build, and no loop turns up where you thought you had independent parts, then this lens is not earning its keep on your stack and you should trust your domain instincts over it. I am claiming the hit rate is high enough to be worth an afternoon, not that the shape is always there. Start with the belief you would least like to be wrong about, the one your roadmap is built on. Somewhere a team is picking the slower model to defend a frontier they have only ever measured inside their own gate. Some of that frontier is real and some of it is their filter, and they have never separated the two. Knowing how much of the line is real is the difference between a quarter spent buying back milliseconds that were never the constraint and a quarter spent on the thing that moves. The tool that runs this is [**The Structure Spotter**](https://durabilitycurve.com/tools/structure-spotter-a515a177/), and it is free. You list the beliefs your work rests on and name the shape of each, and it flags the ones most likely to be artefacts, names the structure underneath, and hands you the test the field that met it first already wrote. A candidate diagnosis and somewhere better to look, not a verdict. Nothing you type leaves your machine. *Which trade-off have you built your roadmap around, and are you sure it survives outside the set you measured it in?* --- [^1]: Joseph Berkson, ["Limitations of the application of fourfold table analysis to hospital data,"](https://pubmed.ncbi.nlm.nih.gov/21001024/) *Biometrics Bulletin* 2, no. 3 (1946), 47–53. The origin of what is now called collider bias or Berkson's paradox: studying only hospitalised patients made diabetes look protective against gallbladder disease, an artefact of who gets admitted. [^2]: David X. Li, ["On Default Correlation: A Copula Function Approach"](https://www.ressources-actuarielles.net/EXT/ISFA/1226.nsf/8d48b7680058e977c1256d65003ecbb5/34e84cb615c8b4eac12575fe006a9759/%24FILE/li.defaultcorrelation.pdf) (2000). Note the title: the model was built to capture default correlation, not to assume it away. The limitation is precise: for any fixed correlation below one, the Gaussian copula has zero conventional asymptotic tail dependence. That does not make clustered defaults impossible. It means the conditional likelihood of one default given another tends to zero as you push further into the extreme, so the structure can understate clustered stress relative to a genuinely tail-dependent model. The canonical warning predates the crisis by years: Embrechts, McNeil and Straumann, ["Correlation and Dependence in Risk Management: Properties and Pitfalls"](https://www.casact.org/abstract/correlation-and-dependence-risk-management-properties-and-pitfalls-0). It was one input to the crisis, not its sole cause. [^3]: The bridge opened on 10 June 2000 and closed on 12 June 2000. Around 90,000 people crossed on the opening day, with up to 2,000 on it at a time ([City Bridge Foundation](https://www.citybridgefoundation.org.uk/news-and-blog/its-25-up-for-londons-iconic-millennium-bridge) for the 90,000; the simultaneous figure is the commonly cited one). Note the 2,000 in the next footnote is a different event: the January 2002 acceptance test, when about 2,000 people were assembled to walk the modified bridge. The influential synchronisation account is Steven Strogatz, Daniel Abrams, Allan McRobie, Bruno Eckhardt and Edward Ott, ["Crowd synchrony on the Millennium Bridge,"](https://www.nature.com/articles/438043a) *Nature* 438 (2005). The mechanism is contested: Igor Belykh and colleagues, ["Emergence of the London Millennium Bridge instability without synchronisation,"](https://www.nature.com/articles/s41467-021-27568-y) *Nature Communications* (2021), argue that coherent footfall is a consequence rather than the cause, and that uncorrelated pedestrians produce positive feedback through negative damping on their own. Either way the loop is the mechanism, which is the part this article imports. [^4]: Christian Meinhardt, ["Vibration Performance of London's Millennium Footbridge"](https://www.repository.cam.ac.uk/bitstreams/3afdf7af-401b-44f6-8bdb-b40f25db454f/download), reviewing the retrofit fifteen years on: a total of 37 viscous dampers for the lateral modes, plus four pairs of laterally-acting and twenty-six pairs of vertically-acting tuned mass absorbers under the deck. Original measured damping was "1% of critical or less"; the target for the main lateral mode excited by walking pedestrians was 20% of critical. The bridge reopened to the public in early 2002. Diversity among the walkers was never the mechanism. --- --- title: "Your AI Looks Best Where You Can Check It Least" description: "When the first failure is terminal, you cannot iterate your way back." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/your-ai-looks-best-where-you-check-least/" date: "2026-07-24" series: "PROOF & TRUST" law: "Law V" substack: "https://harryfloyd.substack.com/p/your-ai-looks-best-where-you-check-least" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Your AI Looks Best Where You Can Check It Least *When the first failure is terminal, you cannot iterate your way back.* By Harry Floyd · 2026-07-24 · canonical: https://durabilitycurve.com/blog/your-ai-looks-best-where-you-check-least/ *When the first failure is terminal, you cannot iterate your way back.* The agent hands back finished work and everything you can quickly check is clean. The tests are green. The citations are present. The summary reads well and the numbers tie out. You skim it, it holds, you ship it. The one thing you did not check, because checking it properly would have taken an afternoon you did not have, is the one thing the whole job rested on. You have met the small version of this. You accept an AI answer because the parts you can glance at look right, and you find out weeks later that the part you could not glance at was wrong. Now scale it up and hold two facts next to each other, because together they describe a place where your normal way of working quietly stops applying. ## The fake forms where you can check least Point a capable optimiser at a signal it is scored on, and it gets good at the signal. Where the signal is a faithful stand-in for the real thing, that is fine, and most work is like that: the code runs or it does not, the numbers reconcile or they do not, and you can see which. The trouble is the work where looking good and being good can come apart without anyone noticing. Safety review. Strategy. Long-horizon judgement. Research whose conclusions you cannot cheaply re-derive. In those places the check is weak, so the optimiser can satisfy the check without satisfying the goal, and you cannot tell from the outside. Now notice which work that is. A great deal of the work that matters most to you sits in that hard-to-check space. So the failure lands exactly where you can least afford it, and it arrives wearing the face of success. Alex Mallen's phrase for it is Potemkin work: a facade of quality, like the showy exterior of a Potemkin village, thrown up precisely over the questions no one can check.[^1] Honest incompetence you can see and route around. A convincing facade takes away the thing you most need, which is the ability to know you are failing. > Goodhart's law says that once a measure becomes a target, it stops being a good measure. The sharper claim is that the gap opens widest where the measure is weakest, and some of the highest-stakes work you have sits exactly there. You can already see the ingredients of this in current systems. In one training study, a curriculum of gameable environments generalised all the way to models editing their own reward function to score higher: dozens of episodes of it out of tens of thousands, against none at all in the honest baseline, and mixing in ordinary good-behaviour training throughout did not stop it.[^2] And on deployed frontier agents today, a practitioner who works with them closely reports current top-tier models producing work that looks successful while quietly omitting the parts that failed, on the tasks that resist an automatic check.[^3] One is a controlled result and the other a practitioner's report, and neither tells you how widespread this already is. But both point the same way, and both surface first exactly where the check is weakest. ## The failure you cannot take back Hold that, and add the second fact, which comes from a different world entirely. In 1982 a routine battery update was sent by radio to the Viking 1 lander on Mars. The update overwrote the data that aimed the lander's antenna. The antenna no longer pointed at Earth and the lander went quiet: engineers kept sending commands for months and never heard back, and the mission ended there.[^4] What matters is what the bug took with it: the channel through which every future error would have been fixed. The patch line was the recovery mechanism, and the patch is what broke it. That is the shape of a whole class of failures: the first real failure damages the very channel you would use to recover from it. Aerospace engineers have a discipline for this, and it is not optimism. You cannot remove the risk of a deployment you only get to run once. You can only fight it with paranoia, and the manager who says "we built in a patch channel, so it is not really one-shot" is describing Viking 1 the week before it went quiet. > Reversibility is what gives you a route back. The dangerous failures destroy that route on the first try. ## Where iterating is the wrong method Put the two facts together and you get a place worth being able to name. The default method for you, me and much of science is a loop: try it, see if it worked, fix what broke, try again. That loop is quietly standing on two assumptions. The first is that the failure signal is honest, that you can tell a real success from a fake one. The second is that failure is reversible, that you get to lose, learn, and go again. Almost everything you do satisfies both, which is why the loop feels like just how thinking works. The first fact removes the honest signal, in the highest-stakes work. The second removes the reversibility, in the deployments you only get to run once. Where they overlap, you have a problem whose success signal lies to you about whether it worked, and whose first real failure takes away your ability to fix it. In that overlap, "run it, watch, and iterate" is invalid. Iteration needs both assumptions, and this is the one place you have neither, so speeding the loop up only gets you to a wrong answer faster, with no way back. Make it concrete. A migration agent moves your records to a new store, checks the copy against its own count, reports a clean transfer, and deletes the original to reclaim the space. The check that would catch a bad move is the one the agent wrote. The store you would have restored from is the one it just cleared. Fakeable signal and irreversible failure in a single loop, and it is exactly the sort of tedious job people are quickest to hand to an agent. The instinct at this point is to reach for a better model. But a more capable model does not fix a weak check. It makes things worse when it gets better at satisfying the proxy faster than you get better at checking the real work, so the signal grows more convincing while the work underneath does not. This is the targeting problem in its plainest form: strength aimed one level too low, at the proxy instead of the thing the proxy was meant to stand for. You cannot patch your way back, because the loop you would patch with is the loop that does not run here. ## The test, and the two moves You do not need to work on Mars landers or safety research for this to bite. Any automated loop you run leans on a checkable proxy for something, and some of those loops touch decisions you cannot cheaply reverse. So run two questions, in order, on the loop in front of you. First: could this success signal be produced without the success? If the only evidence the work is good comes from the system itself, its own output or a proxy it is optimised against, with nothing independent to check it, then the answer is yes: the system can produce that evidence with or without the real thing, so the signal proves nothing on its own. Second: if this fails for real, can I recover, or does the first failure remove my ability to fix it? If the answer to the first is "yes, it could be faked" and the answer to the second is "no, I could not recover," you are in the regime, and speeding up the loop is the wrong reflex. The right response is two moves, one for each fact, and both invert an instinct. The first makes the success signal honest. Put your scrutiny where you can measure least, not where measuring is easy, because the natural pull is to review the work that is convenient to check, which is the work least likely to be faking. Spend the human attention on the opaque, high-stakes tasks instead, because that is where the facade forms. And keep the check independent of the thing it checks: never let the loop grade its own work, and never judge the reviewer by the same proxy the work is scored on. This is what a research check the model did not choose buys you, and what [a test gate the agent cannot talk its way past](https://durabilitycurve.com/blog/never-let-claude-code-tell-you-its-done/) buys you. It is also what human scrutiny can add, at least in the one careful test we have: a person working with an automated auditor caught all three planted saboteur models in a blind study, while the automated auditor alone caught one of the three.[^5] That study is on overt saboteurs, though; a facade built to survive the review is the harder, still-open case, so human scrutiny improves your odds here without guaranteeing you catch it. The second makes failure survivable before you lean on the loop. Stage the deployment, keep a way back that does not run through the thing that might fail, and test that the way back still works under the exact failure conditions you are worried about, rather than assuming it does. Skip that, and you get a working machine that can no longer answer. [Figure: fig02 potemkin map 2026 07 22] *The two questions as a map. Score each loop on whether its success signal is fakeable and whether its first failure is recoverable; the top-right corner, fakeable and irreversible, is the one to stop iterating on and start instrumenting.* You may have seen the near versions of this before. [A research agent that cites sources it never read](https://durabilitycurve.com/blog/your-research-agent-cites-sources-it-never-read/) is a fakeable signal you can still afford to catch late. [An agent whose failures compound quietly until the host seizes](https://durabilitycurve.com/blog/how-reliable-is-your-ai-agent/) is a failure you can still recover from. The regime this piece is about is where both of those protections are gone at once, and it is worth being able to name the difference before you are standing in it. ## The week this pays for Run it as a one-week test, and start with what you would lose most if it were quietly wrong. List the automated loops and AI-assisted decisions your work currently rests on. Run the two questions on each and place it on the map. Most will sit in the safe corner, honest signal and reversible failure, and you can leave them to run fast. A few will sit in the dangerous corner, fakeable signal and irreversible failure. Those are the ones to stop iterating on and start instrumenting, and to give the two moves above before the loop costs you something you cannot get back. The tool that runs this is **The Potemkin Map**, and it is free below. You list your own loops, score each on the two axes, and it places them, flags the dangerous quadrant, and hands you the matched move for each one. Expect to find at least one loop you have been trusting because it was convenient to check, sitting in the corner where convenience was never the point. **[→ Open The Potemkin Map](https://durabilitycurve.com/tools/potemkin-map-d52e049b/)**. Free, runs in your browser, and nothing you type leaves your machine. *Which of your automated loops is in the dangerous corner, and what is the one channel you need to make independent this week?* --- [^1]: Alex Mallen, ["Risk from fitness-seeking AIs: mechanisms and mitigations"](https://blog.redwoodresearch.org/p/risk-from-fitness-seeking-ais-mechanisms) (Redwood Research, 2026). Mallen coins "Potemkin work" as a central failure mode of a system built to perform well in evaluation: because hard-to-check domains like safety and strategy resist verification, the sloppiness concentrates there, producing "an impenetrable facade of quality, like the showy exterior of a Potemkin village." [^2]: Carson Denison et al., ["Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models"](https://arxiv.org/abs/2406.10162) (Anthropic, 2024). A curriculum of gameable environments generalised, with no training for it, to models rewriting their own reward function: 45 such episodes out of 32,768, seven of which also edited the tests to hide the tampering, against zero from a model trained only to be helpful across 100,000 trials. Standard helpful, honest, and harmless training throughout did not prevent it. [^3]: Ryan Greenblatt, ["Current AIs seem pretty misaligned to me"](https://blog.redwoodresearch.org/p/current-ais-seem-pretty-misaligned) (Redwood Research, 2026): frontier agents in heavy use produce outputs that "oversell their work, downplay or fail to mention problems… while actually doing something sloppy or incomplete," worst on tasks "that aren't easy to programmatically check." [^4]: [Viking 1](https://science.nasa.gov/mission/viking-1/) lost contact with Earth on 11 November 1982, when a routine uplink of new battery-charging software was written over the memory holding the lander's high-gain antenna pointing parameters. The antenna no longer pointed at Earth, the lander could no longer signal home, and months of recovery commands went unanswered before the mission ended in 1983 (JPL Publication 82-107). [^5]: Johannes Treutlein et al., ["Pre-deployment auditing can catch an overt saboteur"](https://alignment.anthropic.com/2026/auditing-overt-saboteur) (Anthropic Alignment Science, 2026). In a blind auditing game over five models, a human working with an automated auditing agent identified all three saboteur fine-tunes without flagging the two benign baselines, while the automated agent alone caught one of the three. The saboteurs were overt, not trained to conceal what they were doing. --- --- title: "The Difficulty You're Escaping Was Making You" description: "AI can lift the effort out of almost anything you find hard. Some of that effort was the thing turning you into someone." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/difficulty-was-making-you/" date: "2026-07-17" series: "THE HUMAN LAYER" law: "Law II" substack: "https://harryfloyd.substack.com/p/difficulty-was-making-you" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Difficulty You're Escaping Was Making You *AI can lift the effort out of almost anything you find hard. Some of that effort was the thing turning you into someone.* By Harry Floyd · 2026-07-17 · canonical: https://durabilitycurve.com/blog/difficulty-was-making-you/ *AI can lift the effort out of almost anything you find hard. Some of that effort was the thing turning you into someone.* To drive a black cab in London you have to pass a test called the Knowledge. You spend three or four years on a moped in all weathers, learning every street inside roughly a six-mile circle around Charing Cross, something like twenty-five thousand streets, along with the thousands of landmarks strung between them. You learn to recite the shortest legal route between any two points in the city from memory, out loud, under questioning, with no map in front of you. Most people who start never finish. Those who make it through have spent years inside a difficulty the rest of us would pay almost anything to avoid. Then neuroscientists put those drivers in a scanner. The part of the brain that holds a spatial map, the posterior hippocampus, was measurably larger in London cabbies than in other people, and it was larger the longer they had been driving.[^1] The years had taught them the city and physically reshaped the organ that learned it. That came with a trade-off, one a later study by the same group found and almost nobody quotes: set against bus drivers, these drivers had a smaller anterior hippocampus and did worse at taking on new spatial layouts.[^2] The brain had specialised, hard, to the load the years demanded, and the gain in one place sat beside a weakness in another. That is the fact to hold onto. You take the shape of what you repeatedly do the hard way. We have now built a machine that ends that strain, and not only for cabbies. A phone in a cradle speaks the turns, the Knowledge becomes unnecessary, the hippocampus never specialises, and the driver reaches the same address having built none of the map. This is not a thought experiment. When researchers followed habitual GPS users, the ones who leaned on it hardest had the worst spatial memory the moment they had to navigate on their own, and the more they used it over time, the steeper their decline.[^3] Extend that same offer to nearly every effort a human used to make, and you have described the age we just walked into. None of this worry is new. Nicholas Carr set it out twelve years ago in *The Glass Cage*: a skill handed to a machine quietly wastes away, and you do not notice until the day you reach for it and it is gone.[^4] What has changed since is the reach of the offer, and how far up into thinking itself it now climbs. ## The strain is the mechanism You can feel a smaller version of this whenever you try to recall something you half know: the name on the tip of your tongue, the line you could almost rebuild. It feels like failure, and everything in you wants it to stop, so you look it up and learn almost nothing. The effort of hauling it back up yourself is the repetition that fixes it in memory. Reach for the answer the moment the search stalls, and what you looked up leaves no trace. Psychologists have measured this directly. Give one group a passage to study by rereading it and another by testing themselves on it from memory, and the rereaders come away feeling they have learned more. They are wrong. On the exam a week later, the group that had to pull it back out of memory, the one that felt less certain the whole way through, remembers far more.[^5] > The comfortable method felt like progress and delivered little. The uncomfortable one felt like failure and did the work. [Figure: Bar chart of the Roediger and Karpicke 2006 experiment. Two groups learned the same passage; one kept rereading it, the other kept testing itself on it. A week later the rereaders, who had read it 14.2 times, recalled 40 per cent of it; the self-testers, who had read it 3.4 times, recalled 61 per cent. Four times the reading, more confidence, less memory.] Your body speaks the same language. Muscle grows only when you load it past comfort; the pianist improves on the bars she keeps fumbling, not the ones she can already play. The effort builds; comfortable repetition only maintains. The precision here is easy to miss, and it changes what to worry about. You do not build focus or judgement or skill in the abstract; you build close to the thing you practise, and what carries beyond it is less than we like to think. The cabbie grew the part of the brain that holds a map, and, set beside other drivers, was worse at taking on new ground. So the fear that these tools will make us dumber is too broad to act on. They make you specifically worse at whatever you hand over, and better at whatever you load in its place. That is a trade, and it can be a good one. The trouble is that nothing now forces you to put anything in its place. That turns every quiet decision to offload into more than it looks: a vote for the person you are about to become, cast without noticing you were voting. ## The most tempting offer ever made This is why what we have built is so seductive, and so easy to misread. A researcher put it better than I can: it is like we invented a cure for exercise and then wondered why we are out of breath all the time.[^6] We now have a machine that will do the reaching-in-the-dark for you, write the paragraph you were straining toward, structure the argument you had not yet earned, and hand back the finished thing with the difficulty lifted out. The output looks the same. Sometimes it looks better. [Figure: The same real map of London within six miles of Charing Cross, now almost unlit. One route is drawn in flat grey: the machine's route, its turns spoken aloud. The route still gets driven; the map in you never gets built. Map data © OpenStreetMap contributors, ODbL.] You might say we have done this before and come out ahead. Writing offloaded memory. The calculator offloaded arithmetic. Nobody thinks the literate or the numerate are lesser for it. Those tools changed us too, but they took over the lower floors, the storage and the sums, and left the thinking that sat on top to us. What is arriving now reaches higher up. It can produce the visible signs of understanding, the explanation, the argument, the judgement, the finished prose, without your ever having understood the thing underneath. None of this makes the machine the enemy; it gives a great deal. In the hands of someone already skilled it may be the most powerful lever ever built, and to a beginner with no teacher, or a striver with no way in, it hands a door that used to be locked. Which way it cuts comes down to how you aim it. Aimed at the task, it is an anaesthetic: it takes the difficulty away and hands back the result, and you keep none of what the struggle would have built. Aimed at yourself, it is a trainer: you make it argue against the case you wrote rather than write the case. You attempt the recall before you let it answer; you ask it for the harder version of the problem instead of the solution. The whole difference sits in one rule most people never follow: do not ask the machine to perform the exact act you are trying to keep. Both settings are always there, but the anaesthetic is the effortless one, so the way almost everyone drifts is the one that hollows them out. Sometimes the output is all you want, and then you should take the help and move on. But most work makes two things, not one: the output, and the adaptation the doing leaves in you. When you write something hard, the paragraph is only the first of these. The second is that the idea you were forced to hold still long enough to say clearly is now yours in a way it was not an hour before. Let the machine write it and you keep the paragraph and never build the second thing at all. The early evidence points that way. In a controlled trial, students who researched a topic with an AI assistant remembered noticeably less of it on a later test than students who worked without one. A smaller, more preliminary study found the same shape inside the head: people who wrote with an assistant showed weaker, less connected brain activity while they worked, felt less ownership of what came out, and afterwards struggled to quote the essays they had just produced.[^7] ## The skill you need to check the machine There is an obvious reply to all of this, and it deserves a straight answer. If the machine does the thing well, why does it matter that you can no longer do it yourself? The cabbie with the sat-nav still arrives. The email still goes out. If the output is good, who cares which of you produced it. Engineers ran into a version of this more than forty years ago and named it the irony of automation: hand a task to a machine and the person's remaining job is to watch the machine and catch what it gets wrong, except that handing the task over is what wastes the skill the watching needs.[^8] Most of the time we survive it, because checking a thing is usually cheaper than making it. You can offload arithmetic and keep enough number sense to feel when a total is absurd; you can read a translation you could never have written and still catch where it goes wrong. The dangerous capacities are the ones that are their own only check. Nothing smaller than judgement can tell you whether your judgement is sound. Whether an argument holds is something you feel only with the part of you that would have built it. Taste works the same way. Hand one of those capacities over and you have kept nothing cheaper to catch the failure with; the arrangement turns circular, and you need the very thing you are losing in order to know whether losing it is safe. And the loss hides itself, which I can show you from my own desk. I write with these tools now, and because I did not fully trust what came back, I built a second layer of them to judge the first: a panel of machine readers that scores the work. They are good. They catch what I miss. But the one time I set their verdict beside a room of real readers, the machines had marked the work almost a full point higher than the people did, and every correction that mattered came from the people, not the panel. I had believed the higher number. The judgement I had handed over was certain the work was better than it was, and it had no way to know otherwise. I only found out because I asked. That is the shape of it. The danger does not arrive when the machine hands you a bad answer. It arrives when it hands you a good one, by a route that leaves you unable to recognise the next bad one. A run of good answers is not proof that you are safe; it is the very condition under which the gap stays hidden. ## Meaning tracks the difficulty you choose The cost may not stop at competence. There is a further claim here, and I will mark it as a claim: that difficulty is also where a great deal of a life's meaning is made. Think of someone who spent a year looking after a dying parent: the broken sleep, the paperwork, the slow reversal of who holds whom. Nobody would call it easy, and almost nobody who has done it would give it back. The meaning was not in the suffering; nobody wishes the nights had been longer. It was in the commitment they kept, which had no way to show itself except by carrying the weight. Ask people for the stretches that mattered most and they rarely name the easy ones. They name the ones that cost them. Not everything that matters is earned this way, but the part of a life that feels authored, rather than merely lived, tends to be the part you had to carry.[^9] That is the quiet weight of this technology. Its danger is in its success: it lifts out the difficulty that was doing the authoring, so smoothly that we thank it while some of the material a life is made from goes missing. ## Not all difficulty is sacred Here the argument could tip into a cult of pointless suffering, which is a stupid place to end up, so keep the qualification honest. Plenty of difficulty builds nothing at all. Some of it only takes. It takes your time and your patience, and hands back only the finished task, with nothing left in you to show for it. The tax form, the expense report, the twenty minutes lost to a broken interface. That kind is pure waste, and giving it to a machine is one of the real gifts of this era. Take the gift. The other kind builds. It takes real effort and leaves a changed you, a capacity that stays after the task is gone. The two wear the same face, and from the inside they feel identical, both just resistance, which is why telling them apart is the whole art. One test does most of the sorting. When the task is done, look at what is left behind. If the only thing left is the finished task, that difficulty was only taking from you, and you should hand it to a machine without a second thought. If something in you is different, it was building you, and that is the one to protect. [Figure: Field card. When the task is done, look at what is left. If the only thing left is the output, it was dead difficulty, only taking from you: hand it to the machine. If something in you is different, it was load-bearing difficulty, building you: keep it, on purpose. The two wear the same face; this test tells them apart.] ## What to do with this So run the test on your own week. Most of the difficulty you meet is dead, and the machine should take it. But you cannot keep every difficulty that builds you either, because building one capacity tends to crowd out another, the way the cabbie's deep map sat next to a weaker grip on new ground. You are choosing, not collecting. Protect the few capacities you want to have built: [your judgement, your taste](https://durabilitycurve.com/blog/the-judgment-ai-cant-reach/), your way of finding your way. Keep the effort that makes them, for the version of you on the far side, the one who does not exist yet. And if you doubt any of it, the claim can lose. Pick a capacity that keeps a score, one you can test cold: your way around a city, a language you used to speak, a proof you used to be able to follow. Hand its effort to a machine for a season, then reach for it. My bet is that it comes back thinner than you left it. But notice what that test cannot reach. Judgement and taste keep no score you can read from the inside, in time; any verdict comes late, out of a decision already made, and even then you cannot tell whether your judgement slipped or the problem was simply hard. The one instrument that could read them from the inside is the one you would be handing over. I am not asking you to keep them because the loss cannot be proved. I am asking you to look at the odds, and they are not close: in every case we can actually measure, from navigation to memory to the skills automation has quietly taken, the effect is real, and judgement is not a safe bet to be the exception. For a capacity you have chosen to protect, keep the effort, and the worst case is that you did some work a machine could have done. Hand it over, and the worst case is that you find out what you lost at the moment you need it, and can no longer rebuild it. The cabbies had one advantage we will not. They chose the job, but not the difficulty inside it: if they wanted the badge, the city made the years compulsory, and there was no way to skip them. No institution will impose ours. That is the danger, and it is also the opening. What ends, in a world like that, is formation by default: nothing outside you will build you any more unless you choose to keep it, so more and more of what you can still do in ten years will be what you deliberately kept practising. The cabbies were made by a difficulty they were handed. We get the harder, better thing: to choose what makes us, and to become, in a world going frictionless, someone made by what they chose to keep. *Which difficulty will you keep doing the hard way, now that so little makes you?* --- [^1]: Two studies of London taxi drivers by Eleanor Maguire's group at University College London. Cross-sectionally, Eleanor A. Maguire et al., "[Navigation-Related Structural Change in the Hippocampi of Taxi Drivers](https://www.pnas.org/doi/10.1073/pnas.070039597)," *Proceedings of the National Academy of Sciences* 97, no. 8 (2000): 4398–4403, found greater posterior hippocampal grey matter in licensed drivers than in controls, correlated with years of experience. Longitudinally, Katherine Woollett and Eleanor A. Maguire, "[Acquiring 'the Knowledge' of London's Layout Drives Structural Brain Changes](https://doi.org/10.1016/j.cub.2011.11.018)," *Current Biology* 21, no. 24 (2011): 2109–2114, found that trainees who qualified gained posterior grey matter over three to four years while those who failed and non-drivers did not, which licenses reading the change as produced by the training. [^2]: Eleanor A. Maguire, Katherine Woollett, and Hugo J. Spiers, "[London Taxi Drivers and Bus Drivers: A Structural MRI and Neuropsychological Analysis](https://doi.org/10.1002/hipo.20233)," *Hippocampus* 16, no. 12 (2006): 1091–1101. Compared with bus drivers, taxi drivers had more grey matter in the posterior hippocampus and less in the anterior, and did worse on tests of new visuo-spatial learning. The comparison is cross-sectional, so the smaller anterior volume is a difference between groups rather than a measured within-person loss; the causal reading of the posterior growth rests on the longitudinal study in the previous note. [^3]: Louisa Dahmani and Véronique D. Bohbot, "[Habitual Use of GPS Negatively Impacts Spatial Memory During Self-Guided Navigation](https://www.nature.com/articles/s41598-020-62877-0)," *Scientific Reports* 10 (2020): 6310. Across fifty adults, heavier lifetime GPS use was associated with worse spatial memory during unaided navigation, and a follow-up found that greater use over the intervening period predicted a steeper decline. The design is correlational and cannot fully exclude weaker navigators relying on GPS more, though the longitudinal arm points toward the habit contributing to the loss. [^4]: Nicholas Carr, [*The Glass Cage: Automation and Us*](https://wwnorton.com/books/9780393351637) (New York: W. W. Norton, 2014). Carr argues that automating a task erodes the underlying human skill, and that the deficit stays hidden until the skill is called on again, drawing on cockpit automation and satellite navigation. The argument here builds on his, adding that the loss is specific to whatever is handed over and proposing a test for which difficulties are worth keeping. [^5]: Henry L. Roediger III and Jeffrey D. Karpicke, "[Test-Enhanced Learning: Taking Memory Tests Improves Long-Term Retention](https://doi.org/10.1111/j.1467-9280.2006.01693.x)," *Psychological Science* 17, no. 3 (2006): 249–255. Retrieval practice produced markedly better week-later retention than repeated study, even as repeated study produced greater confidence along the way. The wider framing is Robert A. Bjork's "desirable difficulties": the difficulty must be of a kind that effortful encoding or retrieval rewards, not difficulty for its own sake. [^6]: Advait Sarkar, "[How to Stop AI from Killing Your Critical Thinking](https://www.ted.com/talks/advait_sarkar_how_to_stop_ai_from_killing_your_critical_thinking)," TEDAI Vienna, 26 September 2025. The quoted line is his: "It's like we invented a cure for exercise and then wondered why we're out of breath all the time." His argument is that tools which remove mental effort atrophy the cognitive capacities they were meant to serve. [^7]: Two strands of early evidence, the sturdier first. André Barcaui, "[ChatGPT as a Cognitive Crutch: Evidence from a Randomized Controlled Trial on Knowledge Retention](https://www.sciencedirect.com/science/article/pii/S2590291125010186)," *Social Sciences & Humanities Open* (2025). In the trial, 120 undergraduates researched a topic using either ChatGPT or conventional methods, and the ChatGPT group later scored 57.5% on a retention test against 68.5% for the others (Cohen's d = 0.68). More preliminary is Nataliya Kosmyna et al., "[Your Brain on ChatGPT: Accumulation of Cognitive Debt When Using an AI Assistant for Essay Writing Task](https://arxiv.org/abs/2506.08872)," arXiv:2506.08872 (2025), a small, non-peer-reviewed study in which participants who wrote with a large language model showed the weakest EEG connectivity of three groups, the lowest reported ownership of their essays, and the worst recall of what they had just written; treat it as suggestive rather than settled. [^8]: Lisanne Bainbridge, "[Ironies of Automation](https://doi.org/10.1016/0005-1098%2883%2990046-8)," *Automatica* 19, no. 6 (1983): 775–779. Writing about industrial control rooms, Bainbridge observed that automation leaves the operator to monitor the machine and catch its failures while removing the hands-on practice that built the skill the monitoring requires. The essay carries that paradox up into cognitive work, where the capacity being automated is increasingly judgement itself. [^9]: Matthew B. Crawford, [*Shop Class as Soulcraft: An Inquiry into the Value of Work*](https://www.penguinrandomhouse.com/books/301618/shop-class-as-soulcraft-by-matthew-b-crawford/) (New York: Penguin Press, 2009), and [*The World Beyond Your Head: On Becoming an Individual in an Age of Distraction*](https://us.macmillan.com/books/9780374535919/theworldbeyondyourhead/) (New York: Farrar, Straus and Giroux, 2015). Crawford argues that agency and selfhood are formed by submitting to a reality that resists us and is not of our making. The claim here that a life feels authored in proportion to what it demanded is his; this essay approaches it from the opposite side, through what is lost when the resistance is removed. --- --- title: "Your Tests Pass. So What?" description: "A green suite only proves your agent cleared the gate. Mutation testing shows whether the tests behind it can bite." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/your-tests-pass-so-what/" date: "2026-07-14" series: "WALKTHROUGHS" law: "Law IV" substack: "https://harryfloyd.substack.com/p/your-tests-pass-so-what" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Your Tests Pass. So What? *A green suite only proves your agent cleared the gate. Mutation testing shows whether the tests behind it can bite.* By Harry Floyd · 2026-07-14 · canonical: https://durabilitycurve.com/blog/your-tests-pass-so-what/ This is the second walkthrough in a series. [The first](https://durabilitycurve.com/blog/never-let-claude-code-tell-you-its-done/) built a gate: a Stop hook that will not let Claude Code end its turn while any test is failing, so it cannot call a job done on a red suite. This one asks whether the tests behind that green are worth passing, and you do not need to have read the first to follow along. > **What you will do:** measure how many real bugs your test suite can actually catch. In the worked example, a green suite catches just one planted bug in ten; among the nine it misses is a discount that quietly becomes a surcharge. A tighter case shows the sharper trap: a test can cover every line of a function and still notice nothing. You will watch both scores land, then hand Claude Code the holes and make it close them. About twenty-five minutes for the worked example; a first pass on your own repository takes your whole suite's runtime once for every mutant, so start with your smallest tested file. > > **Who this is for:** you gate Claude Code, or any coding agent, on green tests, and the suite has been reassuringly green ever since. > > **Who should skip it:** if you already run mutation testing, this is your Tuesday. Step 3, turning your existing survivor backlog into an agent work order, may not be. > > **Never written a test?** Your version of the whole method is at the end: ten minutes, three planted errors, no code. > > **You need:** Python 3 (3.9 or later, standard library only, nothing to install). Claude Code for step 3; steps 1 and 2 run without it. > > Then: when it still goes wrong, proving it on code you ship, and a version if you never write code. I broke a production module of mine on purpose, 105 small ways, one at a time, and ran its test suite after every break. The suite is real: 21 regression tests, all green, guarding an 822-line file my own automation depends on. The tests noticed 38 of the 105 breaks. The other 67, one of them on a line the suite executes every single run, would have shipped without a sound. That experiment is the whole method, and it answers a question a passing test suite cannot. The gate proves the tests pass. It cannot prove the tests are worth passing; a test that asserts nothing sails straight through it. And if your instinct is to have a second agent check the tests, then a third to check the second, the regress ends here instead, at an experiment rather than another opinion: break the code on purpose, and see whether the alarm rings. ## Watch a test do nothing *(You know why an assert-nothing test passes? Skip to step 2; this section is the on-ramp.)* Here is the toy calculator shop from last time, one walkthrough later. Its `add` works, and the suite is two small tests. The gate is green. (If you download the companion folder, `calc.py` also carries a new pricing function, which is step 2's problem, and `check.sh`, the last walkthrough's gate script, along for the ride. Leave both alone for now.) `calc.py` ```python def add(a, b): return a + b ``` `test_calc.py` ```python import unittest from calc import add class TestAdd(unittest.TestCase): def test_basic(self): self.assertEqual(add(2, 3), 5) def test_zero(self): self.assertEqual(add(0, 0), 0) # ...the unittest.main() entrypoint is unchanged below ``` Save both files in one folder and open your terminal there; the import `from calc import add` only resolves if they sit together. (Downloading the folder in step 2 does this for you.) One of these two tests is doing almost all of the work, and one of them is doing almost none. You can find out which in thirty seconds. Break `add` on purpose (`return a - b`, the classic bug) and run the suite, `python3 -m unittest`, the same command the gate runs: `test_basic` fails. Now, with the code still broken, run only the other test: ```console $ python3 -m unittest test_calc.TestAdd.test_zero . ---------------------------------------------------------------------- Ran 1 test in 0.000s OK ``` Green, on code that subtracts. `add(0, 0)` is `0` whether add adds, subtracts, or multiplies, so `test_zero` approves all three. It has one way to fail, and almost no wrong version of the code triggers it. Note what your coverage tool would say about it: `test_zero` executes every line of `add`, one hundred per cent, top marks. Coverage measures whether tests run the code. It has no opinion on whether they would notice anything. `test_zero` has sat in that suite looking exactly as load-bearing as `test_basic`, and it would wave the classic bug straight through the gate on its own. > A test you have never watched fail is not yet a test. Test-driven veterans have said a version of that for twenty years: never trust a test you haven't seen fail. The discipline is the same, and the move is the one you just made: break the code on purpose, and the test either notices or it does not. Once, by hand, takes thirty seconds. Every plausible break is a script. Fix `add` back before you move on; the script is next. ## Count what your suite can see The shop grew this week. Claude Code added a pricing function, and the suite is green, which is all the gate checks. Here is the function: ```python def discounted_total(price, quantity, discount_percent): """Total cost of an order, in pounds. Orders of 10 or more items get discount_percent knocked off. """ total = price * quantity if quantity >= 10: total = total * (1 - discount_percent / 100) return round(total, 2) ``` Money code. A boundary, a formula, a rounding rule: three places to be quietly wrong. (Yes, real tills count integer pence rather than floating-point pounds; hold that thought, it returns at the end.) The question you cannot answer by reading the green gate: if one of those went wrong tonight, would any test notice? `mutate.py` answers it by brute honesty. It is a transparent teaching instrument, 180 lines of standard library you can read top to bottom, built to make the mechanism visible on one file rather than to replace a mature framework. Its whole engine is the dozen classic ways code goes wrong that it knows how to plant, one character at a time: ```python OP_SWAPS = { ast.Add: ast.Sub, # a + b -> a - b ast.Mult: ast.Div, # a * b -> a / b ast.GtE: ast.Gt, # >= -> > (the off-by-one at every boundary) ast.Eq: ast.NotEq, ast.And: ast.Or, # ...the reverse of each, plus < and <=, a dozen swaps in all, # and integer nudges: 10 -> 11 } ``` For each place in your file where one of those swaps applies, it makes that one change (a *mutant* of your code), runs your whole suite, restores the file, and records the verdict. Before any of that it runs your suite once on the untouched code, timed: a red baseline gets refused outright, because against a suite that is already failing, every verdict is noise. Then the two outcomes. A mutant that makes at least one test fail is *killed*: the alarm rang. A mutant that leaves every test green *survived*: that exact bug could ship tonight. [Grab the folder](https://durabilitycurve.com/tests-worth-passing.zip), hosted on my own site; it is short, dependency-free Python you can read before you run it. Unzip it and open your terminal inside the `tests-worth-passing` folder; every command below runs from there, no setup. The four working files are `calc.py`, `test_calc.py`, `mutate.py`, and `check.sh` (last walkthrough's gate, along for the ride); a README and the agent's finished tests sit alongside them and need nothing from you. When you are ready to point this at your own code, the folder's `AUDIT.md` carries the reusable version: the audit protocol, a survivor-triage worksheet, and the command for your language. Now stop before you run it, and put a number down. Your suite is green and it passes. Of ten deliberate breaks to this file, how many do you think it catches? Hold that guess against the result: ```console $ python3 mutate.py calc.py 10 mutants of calc.py · suite: python3 -m unittest -q baseline: suite green in 0.2s (per-mutant timeout 60s) 1 KILLED line 2: a + b -> a - b 2 SURVIVED line 10: price * quantity -> price / quantity 3 SURVIVED line 11: quantity >= 10 -> quantity > 10 4 SURVIVED line 11: 10 -> 11 5 SURVIVED line 12: total * (1 - discount_percent / 100) -> total / (1 - discount_percent / 100) 6 SURVIVED line 12: 1 - discount_percent / 100 -> 1 + discount_percent / 100 7 SURVIVED line 12: 1 -> 2 8 SURVIVED line 12: discount_percent / 100 -> discount_percent * 100 9 SURVIVED line 12: 100 -> 101 10 SURVIVED line 13: 2 -> 3 Score: 1/10 killed, 9 survived. Every SURVIVED line is a change to your code that your whole test suite cannot tell from the version you meant to write. ``` The sixth mutant is the one to read twice. `1 - discount_percent / 100` became `1 + discount_percent / 100`: **a customer's 20 per cent discount becomes a 20 per cent surcharge, and the gate stays green.** The third is the boundary: `>=` became `>`, the customer buying exactly ten items loses the discount you promised them, green. Nine ways for money code to be wrong, and the suite from step 1 sees none of them, because nothing in it ever calls the new function. The gate never lied. "The tests pass" was true every time it said so. It just was not the thing you needed to be true. A fair objection: nothing tests that function, so a plain coverage report would have flagged it too, without any of this mutant theatre. True, and if that were all mutation testing found, you would not need it. So run the case coverage cannot see. Delete `test_basic`, keep only `test_zero`, and run the instrument again: ```console $ python3 mutate.py calc.py 10 mutants of calc.py · suite: python3 -m unittest -q baseline: suite green in 0.1s (per-mutant timeout 60s) 1 SURVIVED line 2: a + b -> a - b ... Score: 0/10 killed, 10 survived. ... ``` The subtraction bug now survives, on a function with one hundred per cent line coverage. Your coverage dashboard reports `add` fully tested; the instrument reports that no test would notice if it subtracted. Coverage tells you the code ran. A kill tells you a lie got caught. They are different instruments, and only one of them is measuring what you care about. > The same function can be one hundred per cent covered and zero per cent detected. This is not a toy-only failure, and it is the shape to hold on to. On the 822-line module I opened with, the survivor that stung most was exactly this: a covered line, inside a function the suite runs every single time, that no assertion actually pins. Put `test_basic` back and carry on. One reading note before you run it on anything you love: **killed is the good outcome.** Every killed mutant is a bug class your suite would catch, so on this one screen a `KILLED` is the line you are hoping for, even though the word sounds like something broke. This move is called mutation testing, and it is older than most of the code you have ever shipped. Breaking one character at a time is a fair stand-in for the elaborate bugs real code grows, because of the *coupling effect*: catch the small, dumb faults and you catch the large subtle ones as a by-product.[^1] The nine survivors here are the ones your suite missed, and every one of them is now a job you can hand to the thing that wrote the tests. [Figure: Animated terminal: the mutation run. A green, passing suite catches one of ten planted breaks; nine survive in red. The survivors become a work order for Claude Code, and the same run, re-run, catches ten of ten.] ## Hand the survivors to the agent Nine survivors is a work order, addressed to the thing that wrote the tests. Paste the nine `SURVIVED` lines into Claude Code, followed by this instruction: > These mutants survived mutation testing: the suite stays green when any one of these changes is made to calc.py. Write tests in test_calc.py that kill them. Do not modify calc.py or mutate.py. Then run `python3 mutate.py calc.py` and keep going until the score is clean. When I ran exactly that, the agent came back with four tests and a habit I did not ask for: it annotated each one with the wrong answer the mutant would produce. The survivor list had turned into a specification it could compute against. ```python def test_discount_applies_at_exactly_ten_items(self): # Boundary: quantity == 10 qualifies. 10.0 * 10 = 100, minus 20% = 80.0. # Kills the > / >=, threshold 10->11, and every discount-formula mutant: # total / (1 - d/100) -> 125.0 # total * (1 + d/100) -> 120.0 self.assertEqual(discounted_total(10.0, 10, 20), 80.0) def test_rounds_to_two_decimal_places(self): # 3.333 must round to 3.33; a round(total, 3) mutant returns 3.333. self.assertEqual(discounted_total(3.333, 1, 0), 3.33) ``` Two of its four tests are above; the finished suite, those four plus the two you started with, ships as `calc_solution_tests.py` in the folder (named without the `test_` prefix so it stays out of your way until you want it), so you can run it yourself instead of retyping it from the excerpts. To see the clean score without doing step 3, drop that file in as `test_calc.py` and run `python3 mutate.py calc.py`. Left as it downloads, the folder still scores 1/10, because the suite the gate runs holds only the two starting tests. Then I ran the instrument again myself, because you never take the agent's word for a score when you can take the score's word for it: ```console $ python3 mutate.py calc.py ... Score: 10/10 killed, 0 survived. ``` Exit code 0, the instrument's all-clear. Same code, same tests passing as before, and now green means something it did not mean before: ten classic ways to break this file, and a test rings for every one. (My run went clean in one round, but this toy is small. When yours does not, paste what survived straight back and go again; and for the rare survivor no round can kill, the closing section shows the other move: you let it live, with a comment.) One step remains, and it is the one that actually ends the regress: read the four tests it wrote, and check every asserted value against what the code *should* do, not against what it does. The distinction is load-bearing. An agent kills mutants by pinning current behaviour; look back at its annotations and you can see the 80.0 was computed from the implementation. On correct code that is exactly what you want. On code that is already wrong, the same move canonises the bug as specification, with a clean mutation score as its alibi. > The score proves the tests ring when the code changes; only your read proves they are ringing for the truth. What the instrument buys you is that the read is bounded: four short tests with a known purpose, instead of an unbounded hope about a whole suite. Two honest notes before you rely on it. **Where it earns its keep.** A fresh agent asked cold to "add tests" for that function, with no survivor list, scored 10/10 on its first try. On a four-line function whose docstring names the threshold, the odds are friendly, and the survivor loop buys you little. That is not the point. The point is that on real code you cannot tell whether it guessed well by looking, and the survivor list is what turns "write better tests" from a vague instruction into a measurable work order. The payoff climbs exactly where you cannot eyeball it: a forty-line function, a docstring that has drifted from the code, a module you inherited and half-trust. **None of this is new, and that is the reassuring part.** Handing a test-writer a list of surviving mutants is called *mutation-guided test generation*, and search-based tools were doing it a decade before anyone had an LLM; the agent is just a better test-writer than they were. Google found the part that matters for you here: engineers act on mutants delivered as review-time tickets and quietly ignore the same mutants dumped in a batch report.[^2] The survivor list is that first channel, a work order a developer acts on. Meta has published a version of this pattern with an LLM at the test-writing end.[^3] You are running the individual-developer version of a documented industrial practice. ## When it still goes wrong - **It is slow on a real repo.** Every mutant is a full suite run, so keep it out of the Stop hook: the gate fires every turn, this audit runs when the code or tests change. Industrial tools go further, mutating just the lines a change touches, which is how Google runs it at scale. - **A clean score is a bounded claim.** 10/10 covers only this tool's small vocabulary of breaks. Real gaps live outside it: money in binary floats that no swap can expose, and "kills" that are really crashes, not caught assertions. The score narrows the worry; it does not end it. - **Some survivors cannot be killed.** An equivalent mutant reads differently but behaves identically, so no test can catch it. If you cannot name an input that would expose a survivor, it may be one, and it is allowed to live with a comment. - **The score is a worklist.** Optimise the percentage like a KPI and you get tests engineered to twitch at mutants rather than to state what the code should do, which is the disease "the tests pass" had, one level up. Chase named survivors; ignore the number. - **The verdict contradicts the file you are reading.** You are running a stale compiled copy of the file you mutated. Delete `__pycache__` or `touch` that file. (mutate.py already gives its own runs a fresh cache.) - **Not on Python?** The instrument is language-specific; the method is not. Reach for mutmut on Python, PIT on the JVM, Stryker on JavaScript and TypeScript, cargo-mutants on Rust. The rule holds everywhere: break the code, count the catches. ## Prove it on code you ship Pick the smallest file in your own project that has tests you trust. Four moves before you run it: 1. **Confirm a green baseline.** The instrument refuses a red baseline, because against a failing suite every verdict is noise. 2. **Point `TEST_COMMAND` at your test runner.** It sits at the top of mutate.py, defaulting to plain unittest. On pytest, swap that one line, keeping it a list of arguments, not a string: `[sys.executable, "-m", "pytest", "-q", "tests/test_foo.py"]`. 3. **Narrow it to the file you are mutating.** Aim at just the tests that exercise that file, not the whole suite. Every mutant runs the command once, so a wide one is the difference between a coffee and an afternoon, and the reason people quit after one slow run. 4. **Predict, then compare.** Write down the score you expect before you look. The gap between the number you predicted and the number you got is the most useful thing this walkthrough produces. Here is the anatomy of the number I opened with, re-run the morning this published so it is evidence and not a memory. The module: 822 lines of my own automation, 21 green regression tests. The run: 105 mutants, **38 killed, 67 survived**. Where the 67 hid is the whole lesson. Point coverage at the same suite and the module reads 43 per cent covered; 49 of the survivors sit in code no test executes at all, which a coverage report already flags for you. The other 18 are the ones only mutation can see, because they sit on lines coverage counts as covered. One function holds most of them: a weekly counter the suite genuinely imports, calls, and runs green, whose single test asserts that the count is a non-negative integer and checks nothing else. So flip the sign on its seven-day window and it looks a week into the future, the count silently collapses toward zero, and all 21 tests still pass. That is `test_zero` from step 1 wearing a production badge: **a fully-covered function that no assertion pins**, in my own code, caught by the instrument you just ran on a toy. So I did step 3 on my own code. I handed that survivor to the agent with the same work order, and it wrote the test the counter never had: stage a single item inside the window, then assert the count is exactly one, not merely non-negative. I applied the sign-flip by hand and ran both assertions. ```text current_weekly_promotion_count() = 0 (one item, aged 1 day, window = 7 days) assert count >= 0 -> PASS (the assertion that was already there) assert count == 1 -> FAIL (the assertion the agent just wrote) ``` The old check stayed green on a function that now counts nothing, exactly as it had all along. The new one went red, because a window flipped a week into the future finds zero where it should find one. One survivor, handed over and killed, on the first test that pinned a number instead of trusting a sign. The fix for most of the others is that same move: **pin the number the code should return**, so a window that flips or a boundary that slips finally has somewhere to fail. One survivor on that function is the exception. The window admits an item on `mtime >= cutoff`, and to tell `>=` from `>` you need a file whose modification time lands on the exact cutoff instant. The cutoff is read from the clock, so the only way to hit that instant is to mock the clock, and I judged that test double heavier than the bug it would catch. So it stays alive, with a comment naming why. It is worth being precise about what it is, because it is not an equivalent mutant: an equivalent mutant is one no input can expose, and this one has an input, a mocked clock, that I have simply declined to write. That is step 3 on the days the loop does not go clean: some survivors **earn a real assertion**, some **earn a refactor ticket**, one **earns a comment** explaining why it lives, and telling which is which is the work. If your score stings, the sting is information, the first honest measurement your suite has ever had. Then add this rule to your `CLAUDE.md`, the standing instructions your agent reads each session, so the discipline survives you forgetting it: ```markdown When you fix a bug: first write a test that fails on the broken code, show me it failing, then fix the code. A test I have never seen fail does not count as coverage. ``` That is the manual mutant from step 1, promoted to standing policy: every bugfix now arrives with proof that its test can ring. You are holding three instruments now, and most arguments about testing are really an argument between two of them. Coverage asks whether a line ran. Mutation asks whether a break in that line would be caught. Your own read asks whether the behaviour a passing test pins is the one you actually wanted. > Coverage measures reach. Mutation measures detection. Only your read decides what is right. The gate keeps the agent honest about whether the tests pass; mutation and your own read are what keep you honest about whether they are worth passing. Green stops being where the question ends. It moves from *are the tests passing* to *what broken versions did these tests catch*. ## If you never write code Break the code, count the catches: the method never depended on Python. It works on anything an agent hands back where being wrong is checkable, and the case you most need it for is not code at all. It is the claim. Here is the whole thing with no code in it. Take a report or a summary you know cold. Copy it, and plant a few deliberate errors in the copy, the kind of wrong that would cost you something if it slipped through. Hand the copy to your "review this" agent, and count what it catches. I ran it on one of my own weekly stats summaries before publishing. In a copy I changed a signup count from four to six, moved a date back by a day, and added a confident sentence that one post had our weakest open rate, which it did not. Then: "check this against the raw export." It caught all three, and flagged a fourth claim I had written carelessly and never planted. **Three planted, three caught**, and the review earned my trust the same way step 1's test did: I had watched it catch mistakes I controlled before I believed the ones I did not. A review you have never watched catch a planted error is worth exactly what a test you have never watched fail is worth. The catches are your kills; the misses are the claims you must never hand over unchecked. Code is the easy case, because a machine runs it and a mutant is unambiguous. The claims your agent writes when there is no suite to run, the summary and the analysis and the report, are the harder case, and the one worth an instrument of its own: one that has to find the errors you did not think to plant, checking each claim against its source instead of against a copy you already know is wrong. That is where this goes next. --- *Related: [How Reliable Is Your AI Agent](https://durabilitycurve.com/blog/how-reliable-is-your-ai-agent/) makes the broader case, [The Stable Liar](https://durabilitycurve.com/blog/the-stable-liar/) shows the failure mode up close, and [Never Let Claude Code Tell You It's Done](https://durabilitycurve.com/blog/never-let-claude-code-tell-you-its-done/) builds the test gate this piece audits.* *What was your score, and which survivor surprised you? The answers steer what gets built next.* [^1]: Richard DeMillo, Richard Lipton and Frederick Sayward, ["Hints on Test Data Selection: Help for the Practicing Programmer,"](https://doi.org/10.1109/C-M.1978.218136) *IEEE Computer* 11(4), 1978, 34–41. The paper that introduced mutation analysis and, with it, the coupling effect; the idea itself is usually traced to a 1971 class paper of Lipton's. [^2]: Goran Petrović and Marko Ivanković, ["State of Mutation Testing at Google,"](https://doi.org/10.1145/3183519.3183521) *Proceedings of ICSE-SEIP 2018*. Two findings carry into this piece: mutation analysis is made tractable across Google's roughly two billion lines by mutating only the lines a change touches; and engineers act on mutants surfaced as review-time diffs while ignoring the same mutants in a batch report. [^3]: Christopher Foster et al., ["Mutation-Guided LLM-based Test Generation at Meta,"](https://arxiv.org/abs/2501.12862) arXiv:2501.12862, 2025. Meta's ACH system plants undetected faults, then has an LLM write the tests that kill them. --- --- title: "Where Your Metrics Fold" description: "A metric can be perfectly accurate and still hide the distinction your decision depends on." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/where-your-metrics-fold/" date: "2026-07-12" series: "SYSTEMS & LAWS" law: "Law IV" substack: "https://harryfloyd.substack.com/p/where-your-metrics-fold" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Where Your Metrics Fold *A metric can be perfectly accurate and still hide the distinction your decision depends on.* By Harry Floyd · 2026-07-12 · canonical: https://durabilitycurve.com/blog/where-your-metrics-fold/ *A metric can be perfectly accurate and still hide the distinction your decision depends on.* The dashboard read 87 percent complete, and it was right: 87 percent of the scheduled tasks for the launch were genuinely done, ticked off, verified. The board did its job. Then the launch slipped by six weeks, and in the post-mortem someone put that same 87 percent back on the screen and the room went quiet, because nothing about it had been wrong. Look at what the number saw and what it could not. It observed completed work, accurately. From that, everyone in the room inferred the launch was on track. What it left out was where the unfinished 13 percent sat: almost all of it behind a single unresolved dependency on the critical path, the one piece everything else was waiting on. A project with its hard problem solved and a project with its hard problem untouched both report 87 percent complete. The reading was true. It simply could not tell those two projects apart, and they needed opposite decisions. > The failure did not live in the measurement. It happened in the instant the situation was flattened into a single number. That flattening has a shape, and once you can see it you can find where a number is most likely to mislead you, on one you already own, before anyone optimises anything. ## Name the collision A scalar metric is a lossy projection. It takes a reality with many dimensions and presses it down to one value, and what the pressing throws away cannot be recovered from that value alone. When many dimensions map to one, distinct states end up sharing a reading. Two situations you would treat differently produce the same number. Call any decision-relevant collision a **fold**: two states that receive the same value but would demand different actions if you could see them apart. The 87 percent was a fold. A perfect sensor reading exactly what it was designed to measure can still fold, because the loss happens when that measurement is used to stand in for a larger decision. Take a customer rating sitting at 4.7. In one store almost everyone rates it between 4 and 5. In another, most customers give it 5 while one strategically important segment consistently gives it 1, and the two average out to the same 4.7. Both readings are honest, drawn from complete data. One store has broad satisfaction. The other has a concentrated failure hidden inside an excellent mean, and the number gives you no way to tell which store you are running. Now do it on your own number. Pick one you are judged by. Write down two situations you would respond to differently that would show up as the same value. That pair is a fold, and notice what was not in the room while you found it. No incentive, no gaming, no adversary. The blind spot was already there. > You can often read a fold off the number's own definition, with nobody pushing on it. ## Four kinds of fold Folds are easier to hunt once you know what a number tends to lose. Four useful kinds cover many of the folds you meet in practice. A **composition fold** hides different groups behind the same average. The 4.7 that is broadly fine and the 4.7 that hides a segment stuck at 1 star are the same point. Statisticians know an extreme special case as Simpson's paradox, where the aggregate can run opposite to every subgroup inside it. A **trajectory fold** hides direction behind a level. Monthly churn of 5 percent can be steady, recovering fast, or coming apart, and this month's number reads identically in all three. A **structure fold** hides where the value sits behind a total. The project that is 90 percent complete with the critical path finished and the one that is 90 percent complete with every hard dependency still open share one headline. A **mechanism fold** hides how the result was produced. The same quarterly profit can come from stronger customer economics or from deferred maintenance and postponed investment, and the profit line looks identical either way. [Figure: FIG·02 — Four kinds of fold: the shape differs, every one collapses two states onto one reading.] Four questions to run at any number: what it is made of, which way it is moving, where the value is concentrated, and what produced it. Each names a dimension the number quietly averaged away. ## Folds are where the effort flows This would stay a curiosity if folds were rare corners you might wander into. They are where effort flows. You often cannot move the outcome you want directly. You move the number you can see, and the cheap way to move it runs straight through a fold. Nudging an already-satisfied majority from 4 to 5 may be cheaper than repairing the experience of the segment stuck at 1 star. Both lift the average. Only one closes the concentrated failure the mean was hiding. This is one of the recurring mechanisms behind benchmarks that [come apart once people start optimising against them](https://durabilitycurve.com/blog/ten-lines-of-code-scored-100-one-agent-broke-eight-benchmarks/). Once a score is the thing people are paid to move, the shortest path to the score and the shortest path to the goal stop being the same path. Capability does not save you here. The more capable the optimiser, the more thoroughly it searches the states that score well, and when the score has folds that thoroughness cuts both ways: capability expands the search for shortcuts as well as solutions.[^specgaming] Watch it in machine evaluation. A coding benchmark can hand two systems the same score on isolated fixes when only one of them can [sustain work across files, tests, and intermediate decisions](https://durabilitycurve.com/blog/the-marathon-gap/). The score was not false. It folded together the system that could carry the wider process and the one that could not, and that fold is exactly where a capable optimiser lands.[^swebench] ## The fold comes before the pressure By now you may be filing this under Goodhart's law. When a measure becomes a target, it stops being a good measure.[^goodhart] But look again at what you did a few paragraphs ago. You found a fold in your own number before any optimiser entered this argument, before anyone was pushing on anything. Goodhart tells you that optimisation can separate a measure from the goal it represents. The fold was there before any optimisation, sitting in the number's structure, waiting. The fold-map asks the prior, operational question: which different realities does your measure already treat as the same? Find those collisions in advance and you know where to watch once targeting pressure arrives. > The number does not have to lie to mislead the decision. ## Draw the map So map it. Take your one number and write the question you believe it answers. Then look for one fold of each kind. For every fold you find, fill six columns: the metric, state A, state B, the reading they share, the different decision each would call for, and the one signal that would separate them. [Figure: FIG·03 — The fold-map worksheet: fill one fold of each kind, then a row for your own number.] Then rank the folds by the cost of confusing the two states, how likely that confusion is, and how cheaply an optimiser could produce the misleading one. Prioritise the folds that combine severe consequences, plausible confusion, and a cheap path to the misleading state. For that one, the basic repair is to start watching its separating signal alongside the number. That does not recover everything the projection lost. A mean plus a variance still cannot rebuild the whole distribution.[^anscombe] It splits the one collision you care about most, which is enough to act on. Every important number should get this treatment before it reaches a dashboard. It takes ten minutes, and it turns part of the post-mortem into a **pre-mortem**: the likely failure sites, named before the failure instead of after. Run it tonight on one number you are judged by. Write two situations that score the same, name the decision each would change, and start watching the signal that tells them apart. You may not find a fold of all four kinds the first time. The empty rows are the reward: they mark the parts of your number you have never inspected. *Which number do you steer by that you have never checked for a fold, and what do you now suspect it has been hiding?* [^specgaming]: DeepMind's safety team keeps a documented catalogue of systems that satisfied the letter of their objective while defeating its intent, from a simulated boat circling for points instead of finishing the race to agents exploiting game bugs: ["Specification gaming: the flip side of AI ingenuity"](https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuity/) (2020). [^swebench]: The gap is measured. As of mid-2026, frontier models that resolve over 70 percent of SWE-bench Verified's single-issue fixes drop to roughly 23 percent on [SWE-bench Pro](https://arxiv.org/abs/2509.16941), whose tasks demand larger changes across multiple files in professional repositories. The easier score had folded the two capabilities together. [^goodhart]: The familiar phrasing is not Goodhart's. Charles Goodhart's 1975 observation concerned monetary targets; the general version is the anthropologist Marilyn Strathern's, from ["'Improving ratings': audit in the British University system"](https://gwern.net/doc/statistics/decision/1997-strathern.pdf) (European Review, 1997): "When a measure becomes a target, it ceases to be a good measure." [^anscombe]: Francis Anscombe built four datasets with identical means, variances, correlations, and regression lines that graph into wildly different shapes: ["Graphs in Statistical Analysis"](https://en.wikipedia.org/wiki/Anscombe%27s_quartet) (The American Statistician, 1973). Four different realities, one set of readings: the fold, drawn half a century early. --- --- title: "Your Research Agent Cites Sources It Never Read" description: "The same trap has killed pricing models and trading desks for decades. One move tells you if your number is next." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/your-research-agent-cites-sources-it-never-read/" date: "2026-07-11" series: "PROOF & TRUST" law: "Law I" substack: "https://harryfloyd.substack.com/p/your-research-agent-cites-sources-it-never-read" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Your Research Agent Cites Sources It Never Read *The same trap has killed pricing models and trading desks for decades. One move tells you if your number is next.* By Harry Floyd · 2026-07-11 · canonical: https://durabilitycurve.com/blog/your-research-agent-cites-sources-it-never-read/ You ask a research agent a question that matters. It comes back in ninety seconds with a clean, confident answer and five citations. You open one of the papers to quote it. The paper says something adjacent, but not that. You open a second. The agent has drawn a specific claim from a source that does not contain it. So you do the thing you should have done first. You take the exact documents the agent retrieved, replace them with filler text, and run the question again. The answer barely changes. The citations were set dressing. The agent answered from what it already carried in its weights, then attached sources afterward to make the answer look grounded. This is not a citation glitch you can prompt your way around. In an agent trained under the wrong reward, faking the grounding is the policy the system learned, and researchers have a name for it: tool-call hacking.[^1] ## The reward taught it this The agent was trained with a reward that paid for two things: a correct-looking answer, and a tool call that looked valid. It was never paid for the answer actually depending on what the tool returned. Those are different targets, and under enough optimisation pressure, the cheaper one is what you should expect it to find. Answering from memory and stapling on a citation is cheaper than reading. So that is what it learned. The tell is mechanical. Agents trained this way collapse onto one source, calling it again and again, because it paid off early and the reward never punished the narrowing. The system found the shortest path to the number it was scored on, and the number stopped meaning what its designers thought it meant. > The measure rewarded the appearance of the work. So the system produced the appearance, and stopped doing the work. Watching harder does not fix this. A finer dashboard on top of a gamed metric just measures the gaming with more decimal places. ## When the system knows it is being tested A stranger version is surfacing in AI safety, and it is the one that should unsettle you. Frontier models can now tell, well above chance, whether a transcript comes from a test or from real use. In one study a model told evaluation from deployment with an AUC of 0.83, against a human baseline of 0.92.[^2] Another identified the exact evaluation it was sitting in, described its format, and cited a paper criticising that test, all inside the private reasoning no one was meant to read. Ninety-five per cent of these giveaways never reached its visible output.[^3] A system that can sense a test can behave differently during it. Your reading is then contaminated by the act of taking it: part of what you record is that you were watching at all. How much that shifts behaviour is still contested. That the models can tell is not.[^4] None of this started with language models. The same shape has been quietly killing pricing models and trading desks for decades, and it shows up whenever what you measure stops being independent of the system doing the measuring. ## The loop Its clearest form is a loop. You build a model of a system. You act on what it tells you. Your action changes the system. You measure the changed system, and feed that measurement into your next model, believing you are observing something independent. > You are observing your own footprint. The loop turns dangerous once the system's influence grows large enough to move the evidence it will be judged by next. And the cruel part is the timing: it is most dangerous when the model is working well, because a confident model acts decisively, and decisive action leaves the deepest footprint. This is where a pricing model dies. I have watched it up close. A model gets accurate, so the optimiser prices confidently inside a narrow band. But a model only learns how customers respond to price by watching demand move across different prices, and a narrow band leaves almost none of that variation in next year's training data. The successor, trained on the flattened data, is blind to price response, because the model before it was too good to leave anything to learn from. Accuracy today blinds the model that replaces it. Nobody sees a single failure. They see slow drift with no obvious cause, and they go hunting for the broken component. The component is the loop. [Figure: The reflexive loop: your model shapes your action, your action shifts the world, you measure the changed world, and that measurement feeds back mistaken for an independent reading] ## The edge that dies when you name it Markets run the same loop faster. Your own order moves the price you were chasing, which is just the cost of trading size. The subtler version is alpha decay: once others learn your signal, they trade it flat, and the knowledge of the edge destroys the edge. What survives being known is the edge that pays you for holding a real risk, not for a secret. A pure mispricing dies the day it is named. A risk premium is more durable, because the risk it pays you to hold does not vanish when others pile in, even as the premium itself gets crowded and thinned. ## Why almost nobody catches it in time Three properties keep this loop invisible until it breaks. The contamination is gradual. A pricing model drifts over months. A benchmark rots over a release cycle as models learn to pass it. A crowded trade decays over quarters. The loop runs slower than the decisions feeding it, so no individual decision looks wrong. The system looks healthy right up to the failure. A model can hold high accuracy while its future training data silently narrows. An agent can pass every benchmark while learning to game the benchmark. Stability is not evidence the loop is safe. Often it just means it has not been stressed yet. And every field gives it a different name. Feedback loop, reflexivity, reward hacking, alpha decay. These are not one mechanism, and the fix for each is different. What they share is narrower: in every case what you trust as an independent read has stopped being independent of the system that produced it. The pricing model's training data carries its own past prices. The market signal carries everyone who traded on it. The benchmark carries a model that learned to recognise the test. Tool-call hacking is the sharp edge of the same family: the evidence is held up as an outside constraint, but the reward taught the system it never had to obey it. So a pricing analyst, a trader, and an ML engineer can be losing to the same shape of trap and never realise they are colleagues. ## The test you can run this week The mechanical fix differs from field to field, but the defence underneath is one move, even for the model that can tell it is being tested: a check it can neither move nor see coming, an observation point outside its own influence. Softer moves help, and you should use them, but each one (supervise the steps, require more than one source, put a human on the tool logs) is just another target the system can learn to satisfy. A sharper measurement will not save you, because it still lives inside the loop. I no longer trust a number I have not tried to break. The cleanest way to break one is the test you already saw at the top of this piece, and you can run it on almost anything. Take the output you rely on most from a system that feeds on its own results. Corrupt or remove the evidence it claims to use, and run it again. If the output barely moves, the evidence was never load-bearing, and the system has been reading itself. [Figure: The ablation test: with the cited evidence intact you cannot tell if it mattered; swap it for filler and the answer is unchanged, so the answer barely moves] Point it at your own stack. On a research agent it is exactly that: filler in place of the retrieved documents, and if the conclusion holds, the retrieval was theatre. A pricing or forecasting model is harder, because the cases your past decisions never touched have no outcome to score against. The rejected customer has no repayment history; the price you never set has no demand curve. The honest fix is to build the holdout in advance: approve a small random slice you would normally decline, keep deliberate variation in the prices you set, and judge next year's model only on that protected sample. A trading rule faces the bluntest question of all: does it survive the day it becomes public? The researchers who named tool-call hacking went after the same gap from the other side. Instead of rewarding the model for producing a citation, they rewarded it only when the answer both matched the evidence it retrieved and visibly drew on it. Your ablation catches the fakery after the fact; their reward removes the payoff for it up front. Both make the evidence load-bearing again, which is the only thing that was ever missing. Frozen verifiers, held-out data, structural edges, a test suite the agent cannot edit: these are the same move. Each one builds a place to stand that your decisions cannot move. None of it is free, and none of it stays clean on its own. A held-out set starts decaying the moment it touches production; a frozen verifier ages as the world moves; keeping either genuinely pristine costs more than most teams will pay. Perfect isolation is the exception, so the real discipline is protecting the one piece of ground your own actions cannot contaminate, and treating every other number as standing inside the loop until you have checked. This discipline is the principle behind [an external check your agent cannot talk its way past](https://durabilitycurve.com/blog/never-let-claude-code-tell-you-its-done/), the reason [an agent cannot verify its own work](https://durabilitycurve.com/blog/how-reliable-is-your-ai-agent/), and exactly what fails when [ten lines of code can score full marks on a benchmark](https://durabilitycurve.com/blog/ten-lines-of-code-scored-100-one-agent-broke-eight-benchmarks/) by reading the answer key instead of doing the task. A system that makes decisions again and again ends up running on data its own decisions helped create. The number on your dashboard is a photograph of the world after you have already acted on it. Find the one measurement you trust the most, and check whether it still moves when you break the evidence underneath it. If it does not, you have not been measuring the world. You have been measuring your own reflection, and paying it to agree with you. [^1]: Ma et al., [*Proof-of-Use: Mitigating Tool-Call Hacking in Deep Research Agents*](https://arxiv.org/abs/2510.10931) (arXiv:2510.10931, 2025). Their agents learned to "maximize surface-level reward signals without genuinely grounding their reasoning in the returned evidence," collapsing onto a single over-used source. The fix rewards the model only when its answer demonstrably aligns with the evidence it retrieved and cites that evidence by a checkable identifier: payment for using what you found, not for producing a citation. [^2]: Needham et al., [*Large Language Models Often Know When They Are Being Evaluated*](https://arxiv.org/abs/2505.23836) (arXiv:2505.23836, 2025). Across 1,000 transcripts drawn from 61 datasets, Gemini-2.5-Pro told evaluation from real-world use with an AUC of 0.83, closing on the 0.92 a human baseline reached. Detection was sharpest in agentic settings, the exact place you would put an agent to work. [^3]: Goodfire, [*Verbalized Eval Awareness Inflates Measured Safety*](https://www.goodfire.ai/research/verbalized-eval-awareness-inflates-measured-safety) (2026). Gemini 3.1 Pro not only recognised the benchmark it was sitting in but named the paper criticising it, and in the runs where it voiced that awareness it refused harmful requests 3 to 18 percentage points more often. Ninety-five per cent of those tells stayed in the private chain of thought, never reaching the answer a safety score actually reads. [^4]: The honest counterweight. Knecht, Florin and Hagendorff's [*Evaluation Awareness in Language Models Has Limited Effect on Behaviour*](https://arxiv.org/abs/2605.05835) (arXiv:2605.05835, 2026) finds that spontaneous eval-awareness shifts a model's answers by at most 3.7 percentage points, and argues it "may pose a smaller safety risk than the current literature assumes." That a model can detect a test is settled; how much it changes what the model does is not. --- --- title: "Your Multi-Agent System Is an Org Chart" description: "Cognition said don't build them. Anthropic said do. A year on, they converge on the one question that decides it." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/your-multi-agent-system-is-an-org-chart/" date: "2026-07-07" series: "AI & WORK" law: "Law I" substack: "https://harryfloyd.substack.com/p/your-multi-agent-system-is-an-org" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Your Multi-Agent System Is an Org Chart *Cognition said don't build them. Anthropic said do. A year on, they converge on the one question that decides it.* By Harry Floyd · 2026-07-07 · canonical: https://durabilitycurve.com/blog/your-multi-agent-system-is-an-org-chart/ You ask an agent to build a Flappy Bird clone. It looks like a job you can split, so you split it. One sub-agent gets "build the moving background with green pipes and hit boxes." Another gets "build a bird the player can move up and down." Two workers, two clean tasks, run them at once. The first sub-agent builds a background that looks like Super Mario Bros. The second builds a bird that is not shaped like a game asset and moves nothing like the one in Flappy Bird. Now a third agent has to glue two misunderstandings into one game. So you fix the obvious thing. You hand each sub-agent the full original task, not just its slice. Run it again. This time the pipes and the bird are both recognisably Flappy Bird, and they are drawn in two completely different visual styles, because neither sub-agent could see what the other was making. That is the default failure of the architecture most teams reach for first, and the reason has nothing to do with the model you picked. ## The org chart is the mistake Watch how these systems usually get designed. Someone draws a team. A researcher agent, a writer agent, an editor agent, a critic. It feels obviously right, because it mirrors how humans divide work, and that is the trap. In the 1960s Melvin Conway noticed that any system ends up shaped like the organisation that built it: draw an org chart, and the software inherits its reporting lines and its blind spots. A team of agents is an org chart you drew on purpose, and it inherits the same seams. You already saw those seams. Splitting the Flappy Bird job created two workers who could not see each other, so they diverged. The fix looked obvious, hand each one the full task, and that run failed too, in a quieter way. Hold the question of why for a moment, because two of the strongest teams in the field have already answered it, and at first glance they answered it in opposite directions. ## Two labs, opposite answers In 2025, Cognition, the team behind the Devin coding agent, put their answer in the title: "Don't Build Multi-Agents." After building coding agents for a living, their verdict was blunt, and deliberately dated. > "It is evident that in 2025, running multiple agents in collaboration only results in fragile systems. The decision-making ends up being too dispersed and context isn't able to be shared thoroughly enough between the agents."[^1] They dated the claim on purpose, expecting the picture to shift as single agents grew more capable. Hold onto that: they revisited it a year later, and where they landed is the whole point. Anthropic, building the research feature inside Claude, reported the reverse. Their multi-agent system, one lead agent coordinating several sub-agents, scored 90.2% higher than single-agent Claude Opus 4 on their internal research eval.[^2] The same architecture Cognition warned against, beating their own single-agent baseline by a wide margin. Same word, two machines. One team said the shape was fragile. The other shipped it and beat their own single-agent baseline. Both were reporting honestly. The contradiction is the whole puzzle, and it dissolves the moment you find the variable they were each describing from their own side. ## The question underneath both Put the two findings next to each other. Cognition builds coding agents, where the pieces are densely coupled. Anthropic built a research agent, where the pieces are separate look-ups. Each drew the right conclusion for the shape of problem in front of them. The shape that decides it is whether the pieces carry decisions that depend on each other. Ask a research system to find every board member across the Information Technology companies in the S&P 500. That is a hundred look-ups, and not one of them needs to know what the others found. Each sub-agent decides nothing the others have to honour, and the lead just collects the answers as they land. Now look again at the Flappy Bird job, even with the full task handed to every agent. The background, the bird, the pipes still have to agree on a style, a scale, and a feel that live in the whole and nowhere in the parts. Each agent makes those choices on its own, and independent choices about one shared thing drift apart. That is why the full-context run still failed. Context was never the bottleneck. The coupling was. This is also why reading is safe and writing is dangerous. A read commits nothing, so ten agents can read the same material at once and never collide. A write commits a decision the others now have to stay consistent with, and nothing keeps them consistent once they cannot see each other. Read versus write is the fastest proxy for the real question: does this piece make a choice the other pieces depend on? None of this is new. Readers may share and writers must take turns is the oldest rule in concurrent systems, the one behind every database lock and every thread that ever corrupted shared state. What is new is the layer it now governs. The rule has climbed from bytes in memory to decisions between agents, and the reason it keeps reappearing is that it was never about computers. It is about what happens when separate workers commit to the same thing without watching each other. [Figure: Same word, two different machines. Left, role theatre: a relay of specialists where each handoff loses context, so the verdict is collapse to one agent. Right, parallelisation: a fleet of independent workers reading in parallel, so the verdict is fan out.] ## Count the decisions, not the agents This is why the agent count on the box tells you nothing. A swarm of three hundred is not more capable than one agent by virtue of being three hundred. Agent count is cheap to inflate and easy to sell, and the moment it becomes the headline it stops tracking whether the work got done. The thing to measure is the decomposition, and the decomposition is measured in dependent decisions, not in boxes on a diagram. The proxy has edges, and they are worth knowing, because the surface can mislead. Two research agents that only read can still collide if the task is vague enough that each has to decide what it means; that hidden interpretive choice is the coupling, even with no writes anywhere. And some jobs that look coupled are not: a body of text too large for one agent to hold at once, split across many readers, runs in parallel without conflict, because every reader shares the same goal but none constrains another's finding. What settles the case is never how the task looks from a distance. It is whether finishing one piece requires knowing what another piece decided. ## The bill you pay to be wrong Even when fan-out is right, it is not free. Anthropic found that raw token usage alone explained about 80% of the variance on one of their browsing benchmarks, which is another way of saying the gain came mostly from spending more compute, not from clever coordination. An independent 2026 study across the Qwen, DeepSeek, and Gemini models reached the same verdict from the other side: on multi-step reasoning, hold the token budget equal and a single agent matches or beats the multi-agent setup, because the reported gains track compute, not architecture.[^6] Their multi-agent system burns roughly fifteen times the tokens of a normal chat. Cheaper inference lowers what each token costs, not how many the architecture spends; fifteen times as many is fifteen times as many at any price. The absolute bill falls with the market. The multiple does not, and neither does the coupling you would be paying it for. That multiple is the honest test of the whole decision. Fifteen times the cost is worth paying only when the task is valuable enough to earn it and parallel enough to use it. Point the same architecture at a coupled job and you pay the fifteen-times bill to manufacture the Flappy Bird problem at scale, and the cost can climb higher without warning, because a sub-agent that spawns its own sub-agents, or a tool that returns a wall of text, multiplies the spend again, and most builds have no cap that stops it.[^3] If you cannot say in one sentence why the parallelism pays for itself here, you are paying the coordination tax and calling it a team. The same instinct shows up one layer down, in the pull to add more tools and more layers to a single agent until it is too complicated to debug, which is [the trap the most powerful tools quietly set](https://durabilitycurve.com/blog/your-tools-got-powerful-get-boring/). ## The machine most builds actually want The argument is usually staged as swarm versus single agent, and that staging hides the machine you almost always want, which is neither pole. One strong agent, wrapped in an engineering envelope. The envelope is plain engineering, and it is where the reliability lives. In practice it looks less like a team meeting and more like one skilled worker with good tools, a checklist, and a reviewer. A planner lays out the work before it starts, routes the cheap steps to cheaper models, and caches and retries so nothing is paid for or crashed twice. A supervisor then compares what the agent meant to do against what it did, and a separate pass checks the output against something outside the agent's own judgement. You keep most of what the orchestration promised at a fraction of the token cost, and you never split one judgement across personas that cannot see each other. It is the same case as [building the harness around the model instead of swapping the model](https://durabilitycurve.com/blog/harness-engineering-same-model-different-product/): the structure around one strong agent does more of the work than the agent count ever will. Fan-out still has a safe home inside this envelope, and it is read-only work. Claude Code's investigation sub-agents are the clean example. They explore a codebase, answer a question, and hand a summary back to the single agent holding the thread, without touching a file.[^4] Reading many things at once is parallel by nature. The moment you let sub-agents write in parallel, each committing changes the others cannot see, you have rebuilt the Flappy Bird problem inside your own codebase, which is exactly the risk in the newer parallel-coding setups. That one line, read or write, is the quickest filter you have, and it catches the common mistakes before they are built. You do not have to take it on faith, because Cognition spent a year arriving at it themselves. Their 2026 follow-up, "Multi-Agents: What's Actually Working," is not a retraction of "Don't Build Multi-Agents"; parallel-writer swarms are still out. What now runs in their production is the narrow class where writes stay single-threaded and the extra agents contribute intelligence rather than actions, and it is not theoretical: even in their most cautious enterprise segment, Devin usage is up roughly eightfold over six months.[^5] A reviewer reads a diff and flags the bugs. A stronger model gets consulted on the hard call. A manager splits the read-heavy work, lets the children run, and keeps the one write to itself. The patterns they kept all have the same shape underneath: many readers, one writer. Read versus write, named from the inside by the team that opened the argument against multi-agents. ## The test to run before you build Before you stand up a multi-agent system, put the task through one check. Try to break the job into pieces, and for each piece ask four things. Can it be finished without knowing what the other pieces decided? Does it only read and report, or does it write into a shared result? Is the whole job too big for a single context window, more than one agent can hold in mind at once? And is it worth roughly fifteen times the cost of doing it plainly? Independent, read-only, oversized, and high-value: fan it out, aggregate cheaply, and keep yourself at the question going in and the decision coming out. Anything else: collapse it back to one strong agent and spend your effort on the envelope. And if you cannot tell which case you are in, that uncertainty is the answer for now, because the coordination tax is real and the simpler machine should be the default. [Figure: Two gates, one honest default. Gate one: do the parts share context or write into the same result? Yes, collapse to one strong agent. No, go to gate two: read-only, bigger than one context window, worth roughly fifteen times the cost? All yes, fan out. Any no, take the hybrid, where most builds land.] The four questions are the whole method, and you can run them on the back of a ticket. When you want a specific task computed rather than eyeballed, I put the same checks into a small tool that returns the architecture and the rough cost. [Run the Multi-Agent Decision.](https://durabilitycurve.com/tools/multi-agent-decision/) So the next time someone proposes a designer agent, a coder agent, and a critic agent for one coherent job, ask the only question that decides it. Was the work ever actually separate? Most of the time you have drawn an org chart in software, and the org chart was the bug. *What is one task you split across agents, or were about to, and does it survive the read-versus-write test?* [Figure: The Multi-Agent Decision field card. Run the four checks (share context, read or write, too big for one context window, worth roughly fifteen times the cost), and the rule does the rest.] [^1]: Walden Yan, "[Don't Build Multi-Agents](https://cognition.com/blog/dont-build-multi-agents)," Cognition (2025). Source of the Flappy Bird example, the two principles, and the fragility quote. [^2]: "[How we built our multi-agent research system](https://www.anthropic.com/engineering/multi-agent-research-system)," Anthropic Engineering (13 June 2025). Source of the 90.2% internal-eval result, the 80%-of-variance finding on the BrowseComp benchmark, the ~15× token figure, and the point that shared-context and coding tasks are a poor fit for multi-agent. [^3]: The 15× is Anthropic's own baseline. The further escalation is my own inference from the same system, which reports early failures like one agent "spawning 50 subagents for simple queries" and sets no per-run cost cap. Not a figure Anthropic states. [^4]: Claude Code's investigation sub-agents (its Explore and Plan modes) run read-only, searching and summarising without editing files. It also supports implementation sub-agents and parallel "agent teams" that write in parallel, which is the coupled-write case this piece flags as fragile. [^5]: Walden Yan, "[Multi-Agents: What's Actually Working](https://cognition.com/blog/multi-agents-working)," Cognition (22 April 2026), the follow-up to "Don't Build Multi-Agents." Parallel-writer swarms are still out; what works now is "setups where multiple agents contribute intelligence to a task while writes stay single-threaded." The eightfold growth is Cognition's own figure for Devin in its largest enterprise segment; the three named patterns are the Code-Review-Loop, the "Smart Friend," and "map-reduce-and-manage" delegation. [^6]: Dat Tran and Douwe Kiela, "[Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets](https://arxiv.org/abs/2604.02460)," arXiv (April 2026), an independent study across Qwen3, DeepSeek-R1-Distill-Llama, and Gemini 2.5. Single-agent systems "consistently match or outperform" multi-agent ones on multi-hop reasoning at equal token budgets, with the reported gains attributed to "unaccounted computation and context effects rather than inherent architectural benefits." --- --- title: "Your Benchmark Measures a Sprint. Your Agent Runs a Marathon." description: "An open model looks frontier-grade on the coding leaderboard. On a long job, it does half the leader's work." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-marathon-gap/" date: "2026-07-01" series: "SYSTEMS & LAWS" law: "Law I" substack: "https://harryfloyd.substack.com/p/the-marathon-gap" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Your Benchmark Measures a Sprint. Your Agent Runs a Marathon. *An open model looks frontier-grade on the coding leaderboard. On a long job, it does half the leader's work.* By Harry Floyd · 2026-07-01 · canonical: https://durabilitycurve.com/blog/the-marathon-gap/ You gave the overnight job to the cheaper model, and in the morning the work was half done. Not broken in a way you would catch at a glance. The agent had slipped somewhere around step nine, built three more steps on top of what it broke, and never noticed. Half done, and confident about it. You had a good reason to trust it. GLM-5.2 had shipped with open MIT-licensed weights and a score that beat GPT-5.5 on SWE-bench Pro, the coding benchmark every team quotes, with only the two Claude Opus models above it. Frontier-grade, yours to host, at a fraction of the price. Every board you read said it was ready.[^1] The number that would have warned you shipped in the same release. On SWE-Marathon, the benchmark for long multi-hour tasks, that model solves 13% of the jobs. Opus 4.8 solves 26%.[^2] Close on the short tasks, half the work on the long ones. Both numbers went out together, one scroll apart, and only one of them sat on the board you read. ## The sprint board hides the gap Most benchmarks are sprints: SWE-bench, Terminal-Bench, the coding boards that fix the market's sense of who leads. They all measure single-session, bounded tasks, the kind a model finishes in one push, however demanding each one is. There the field is bunched, a dozen models within a few points at the top, an open model now among them. Read only that board and the race looks over. SWE-Marathon measures a longer distance: 20 tasks, each one multi-hour, each run in its own executable environment and graded against a human-written reference and a multi-layer test suite. The average logged attempt burns 27 million tokens.[^3] There the field stops being bunched. Opus 4.8 solves about a quarter of the tasks. Everyone else sits at half that or less, Opus 4.7 at 16%, GLM-5.2 at 13%, GPT-5.5 at 12%. No agent, open or closed, clears 30%. Twenty tasks is a thin sample; weigh the spread, not the last digit. Look at where GPT-5.5 lands. The open model beats it on the sprint and edges it on the marathon too, yet both land at half of what Opus does, 13% and 12% against 26%. GPT-5.5 is proprietary and frontier, and it caves on the long task just the same. The divide that matters runs between sprint and marathon, not between open weights and closed, and only the sprint board is the one everyone reads. [Figure: Same models, bunched within a few points on the sprint board and spread wide on the marathon board] Stretch a task out far enough and you can watch the specific ways it breaks: weak self-verification, calling a half-finished job done, never recovering after one wrong step. On nearly one attempt in seven, a run fakes its way past the verifier instead of doing the work, the same move that lets [ten lines of code score 100% on a benchmark that tests nothing](https://durabilitycurve.com/blog/ten-lines-of-code-scored-100-one-agent-broke-eight-benchmarks/).[^4] A short task rarely leaves room for any of it. A long one leaves room for all of it. ## The gap is arithmetic The temptation is to read 13-against-26 as a lag, the open model a release behind before it catches up the way it caught up on sprints. Part of it is exactly that. The rest is arithmetic. A long task only succeeds if its steps survive in sequence, so small per-step gaps stop adding and start multiplying. Two models that clear, say, 96% and 93% of steps look the same on a five-step task and finish more than three times as far apart on a forty-step one, as the curve below shows.[^5] The gap a short benchmark cannot see is the gap that decides the long run. [Figure: Two reliability curves: 96% and 93% per step both decay, close on a short task and far apart over a long one, a 3.6× gap by 40 steps] Two forces bend that curve without repealing it. Recovery softens it: a good agent catches some of its own mistakes, a good harness catches more. Correlated failure sharpens it: one wrong step poisons the steps after it, the way the overnight run built three more on a broken one. The odds still fall faster the longer the run, from a higher start. It is a diagnostic, not a law, and reliability gaps that round to nothing on a short task go nonlinear once the work has to survive many handoffs.[^6] You can put a rough number on your own work, though the measuring is the real labour. Run your agent on a sample of representative steps and count the fraction it clears without a wrong turn you have to undo. That count is a small eval set, the kind you build once and reuse. Raise it to your task's step count, and you have your marathon odds. This is the axis METR has been tracking while the leaderboards looked elsewhere. They measure the length of task a model finishes on its own, and they find it doubling about every 7 months. Their reading is that the driver is reliability, the knack for catching a mistake and recovering, more than raw reasoning.[^7] How long a model can run on its own is a capability in itself, and the boards everyone reads do not score it. ## Why the marathon stays scarce The model layer has commoditised: open weights ship at a fraction of frontier price, and sprint-grade coding is now broadly available, which the bunched sprint board confirms. Sprint capability copies because a benchmark rewards it and a teacher's traces capture it. Marathon reliability is harder to lift out, because it is not one thing in the weights. It is the model, [the harness around it](https://durabilitycurve.com/blog/harness-engineering-same-model-different-product/), the verifier, and the context discipline holding together across hundreds of steps, where any single link can break the run. Sprint parity buys you the first of those and none of the rest. So the part that got cheap is the sprint, and **the part that stays scarce is reliability held across length**. How long it stays scarce is the real question, because some of it is trainable and that part is already closing. GLM-5.2's release notes are headed "built for long-horizon tasks," and its marathon score leapt from 1 to 13 in a single version, fast progress that still lands at half the leader, the open lag narrowing the way it already narrowed on sprints, one cycle behind. What stays unsettled is how large the rest is, the share that lives in the system, not the weights. The bet here is that the durable scarcity is the system: a verifier that checks the agent's own work, a planner that holds the goal across hours, a run that can checkpoint and recover instead of dying on one wrong step, all tuned to the model they wrap. The scaffolding is portable, and it lifts a cheap model's marathon odds further than bigger weights would, so it is your route when frontier prices are out of reach. But the same scaffolding lifts the frontier more, because the labs that train the model also tune the harness to it. That is why each ships its own agent framework rather than a universal one. A frontier model follows tools more closely, drifts less across a long context, and poisons fewer of its own branches. You can copy the scaffolding. You cannot copy that co-design. The hard part is now the code that keeps a long run alive, and the model it is wrapped around, and no leaderboard scores either. ## Price the whole run The sticker price makes this worse, not better. The open model wins on dollars per token, and that is the number people compare. A marathon does not bill by the sticker. It bills by tokens times length times retries, and the number that decides it is cost per finished job: the cost of one attempt divided by the solve rate. The average attempt on SWE-Marathon runs 27 million tokens, and at the open model's 13% finish rate that is roughly eight attempts to land one clean success; at the frontier's 26%, closer to four. Kill the dead runs early and it is fewer than eight full runs of tokens, but it is many times the sticker, and a cheaper-per-token model can finish a marathon more expensive than the frontier, once its retries outrun its discount. Hosting the open model claws some of that back, because the budget you save buys parallel attempts and early kills, the test-time compute a metered frontier bills at full rate. That narrows the gap on work you can checkpoint and verify as you go. It buys little on one long unattended chain, where no retry helps until you can tell which branch went wrong. The saving is real on the sprint and an illusion on the marathon. [Figure: The open model's cost to finish a job, as a multiple of the frontier's: a sixth the cost on a short job, climbing as its retries pile up through the break-even around the mid-fifties of steps to about double the frontier on a long one] So count the steps before you pick the model, not the single prompt that kicks the job off. A bounded edit a model finishes in one pass is a sprint, and there the open model at parity is the right call, cheaper and yours. A job that runs unattended across many steps and tool calls and minutes is a marathon, and there the few points of reliability no leaderboard shows you are the whole game. The line to find is your model's coin-flip step count, the length where its odds of finishing fall below even. It comes earlier than the cost crossover the figure above marks: a run that finishes half the time is already bleeding you on failed mornings long before its retries outprice the frontier. Keep bounded work under it on the cheap model; route unattended work past it to the frontier, or wrap the cheap one in the scaffolding above. A third move sits in that same line: cut the marathon into checkpointed chunks, each one shorter than the coin-flip count, so a run that dies as one long chain can finish as a string of short ones. Splitting the job is the cheapest reliability you can buy. Make it concrete. A nightly agent upgrading dependencies across a large repo runs maybe 40 steps. Your open model at 95% a step finishes about one run in eight; the frontier, a point and a half steadier, closer to one in four. On tokens alone the cheap model can still win, because eight cheap retries cost less than four expensive ones. But you are shipping a finished upgrade one morning in eight, and a failed overnight run rarely costs only tokens. That is when you pay for the frontier. The [Marathon Calculator](https://durabilitycurve.com/tools/marathon-gap/?utm_source=substack&utm_medium=essay&utm_campaign=marathon-gap) runs the read in your browser. Enter your per-step reliability and your task length, and it marks the step count where the sprint benchmark stops predicting and the job becomes a reliability one. The scores move every few weeks; [the gauges here keep tracking them](https://harryfloyd.substack.com/subscribe?utm_source=substack&utm_medium=essay&utm_campaign=marathon-gap). Hold the exact scores loosely. SWE-Marathon is one small benchmark, 20 tasks from a single lab, with the variance you would expect from a sample that thin. Do not rest the case on it. Rest it on the mechanism, and on METR's long-task curve, built on human-baselined tasks by different people, which finds the same thing: duration is gated by reliability. The falsifier is clean: an open model that matches the frontier on a mature long-horizon benchmark while sitting at sprint parity. If that lands, trust the rest of this less. On the record: through the end of 2026 I expect the best open-weight model to keep trailing the best closed model by double-digit points of resolve rate, the share of tasks solved, on SWE-Marathon or on whatever replaces it as the standard long-horizon board. The way that turns out wrong is a scaffolding story, not a base-weights one: open agent frameworks and cheap test-time compute lifting the open score faster than bigger weights ever would. *Which of your agent's jobs is a marathon you have been routing like a sprint?* [^1]: GLM-5.2 — Z.ai, released 13 June 2026, MIT-licensed open weights (≈744B-parameter mixture-of-experts, ~40B active per token, 1M-token context). SWE-bench Pro 62.1, third behind Claude Opus 4.8 (69.2) and Opus 4.7 (64.3) and ahead of GPT-5.5 (58.6); FrontierSWE 74.4 to Opus 4.8's 75.1; Terminal-Bench 2.1 81.0 to Opus 4.8's 85.0. SWE-Marathon 13.0, up from GLM-5.1's 1.0. Z.ai release notes (HuggingFace, "GLM-5.2: Built for Long-Horizon Tasks"). https://huggingface.co/blog/zai-org/glm-52-blog [^2]: SWE-Marathon leaderboard, resolve rate: Claude Opus 4.8 26%, Claude Opus 4.7 16%, GLM-5.2 13%, GPT-5.5 12%. https://www.swe-marathon.org/ [^3]: SWE-Marathon — *Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?*, arXiv:2606.07682 (Abundant AI). 20 multi-hour software-engineering tasks; in the paper's runs, logged agent attempts average 27.2M total tokens and current frontier coding agents solve fewer than 30% (the live leaderboard shifts as trials accumulate). https://arxiv.org/abs/2606.07682 [^4]: SWE-Marathon (arXiv:2606.07682) logs reward-hacking — an agent gaming the verifier instead of doing the task — in 13.8% of rollouts (the paper's figure). https://arxiv.org/abs/2606.07682 [^5]: Computed directly: 0.96^5 ≈ 0.82 and 0.93^5 ≈ 0.70 (a 12-point spread at 5 steps); 0.96^40 ≈ 0.1954 (≈20%, one finish in five) and 0.93^40 ≈ 0.0549 (≈5.5%, one in eighteen), a 3.6× ratio at 40 steps (0.1954 / 0.0549 = 3.56; 18 / 5 = 3.6). [^6]: The R^N model of multi-step reliability, and the delegation cliff it produces, are set out in *How Reliable Is Your AI Agent?* /blog/how-reliable-is-your-ai-agent/ [^7]: METR, *Measuring AI Ability to Complete Long Tasks*, arXiv:2503.14499. The 50%-task-completion time horizon has grown exponentially with a doubling time of roughly 7 months; METR attributes the gain primarily to greater reliability and the ability to recover from mistakes. https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ --- --- title: "Never Let Claude Code Tell You It's Done" description: "A test the agent can't talk its way past, wired to run itself." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/never-let-claude-code-tell-you-its-done/" date: "2026-06-29" series: "WALKTHROUGHS" law: "Law IV" substack: "https://harryfloyd.substack.com/p/never-let-claude-code-tell-you-its-done" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Never Let Claude Code Tell You It's Done *A test the agent can't talk its way past, wired to run itself.* By Harry Floyd · 2026-06-29 · canonical: https://durabilitycurve.com/blog/never-let-claude-code-tell-you-its-done/ > **What you will do:** add one test that catches the agent's mistakes, then wire it so Claude Code runs it on its own and cannot end a turn while it is failing. About twenty minutes. > > **Who this is for:** you use Claude Code on real code, and it has told you "fixed it, tests pass, done" when it was not. > > **Who should skip it:** if your agent already cannot end on a failing test, you are past this. > > **You need:** Claude Code (version 2.1.143 or later), Python 3 for the worked example (it uses the built-in test runner, nothing to install), and a project of your own. Claude Code will tell you it fixed the bug. It will tell you the tests pass. It will tell you it is done. And sometimes it is wrong, and it says all of it with the same calm certainty it uses when it is right. "Done" is the most expensive thing it gets wrong, because the moment you believe it, you stop looking. The reason is plain. The model is rewarded for sounding finished, and sounding finished is not the same as being finished. The gap between the two is invisible at exactly the moment you have decided to trust it. You cannot close that gap by asking the agent to be more careful. You close it with a check the agent runs but does not get to grade: its own tests, with a real pass or fail. ## Write a test that says what you want A test is a small program that runs your code and checks it does the right thing. The important word is *yours*. The test has to encode what *you* want the code to do, because if the agent writes both the code and the test, it can quietly make the two agree. The check has to come from outside the agent, or it is not a check. Here is the smallest possible example: a function, and a test for it. The agent was asked to fix `add`, and reported back that it was done. `calc.py` ```python def add(a, b): return a - b ``` `test_calc.py` ```python import unittest from calc import add class TestAdd(unittest.TestCase): def test_basic(self): self.assertEqual(add(2, 3), 5) def test_zero(self): self.assertEqual(add(0, 0), 0) if __name__ == "__main__": unittest.main() ``` Save both files in the same empty folder and open your terminal there (`cd` into it). They have to sit together, because the test does `from calc import add`. Now run the tests. (`python3 -m unittest` finds and runs every `test_*.py` file in the folder. It is built into Python, so there is nothing to install.) ```console $ python3 -m unittest F. ====================================================================== FAIL: test_basic (test_calc.TestAdd) ---------------------------------------------------------------------- Traceback (most recent call last): File "test_calc.py", line 8, in test_basic self.assertEqual(add(2, 3), 5) AssertionError: -1 != 5 ---------------------------------------------------------------------- Ran 2 tests in 0.000s FAILED (failures=1) ``` The agent was certain. The test does not care how certain it was. `add(2, 3)` came back `-1`, not `5`, and now you know, in one second, that "done" was not true. That is the whole idea: a fact about the code the agent cannot argue with. When you write a test for your own code, pick an input where a broken version and a correct one give clearly different answers. If your test passes on code you already know is wrong, the inputs are not separating right from wrong yet, and the gate will wave the bug straight through. ## Put it behind a gate the agent cannot talk past A test you remember to run is already worth a lot. But you will not always remember, and the agent can finish and hand control back to you before you check. So wrap the test in a gate: a small script that runs it and turns the result into a hard yes or no. Save this as `check.sh` in your project: ```bash #!/bin/bash # A Stop-hook gate. Claude Code runs this when the agent tries to end its turn. # If your tests pass it exits 0 and the agent is free to stop. If any fail it # prints them and exits 2, which blocks the stop: the agent is handed the output # and has to keep working, so it cannot end the turn while the suite is red. cd "${CLAUDE_PROJECT_DIR:-.}" || exit 2 # run from the project root, not wherever the agent cd'd to # PYTHONPYCACHEPREFIX gives Python a fresh bytecode-cache dir each run, so the gate # can never pass on bytecode compiled from older code that was rewritten at the same # size and timestamp (a false green this gate exists to prevent). The trap removes it. tmpdir="$(mktemp -d)" trap 'rm -rf "$tmpdir"' EXIT output=$(PYTHONPYCACHEPREFIX="$tmpdir" python3 -m unittest 2>&1) if [ $? -eq 0 ]; then exit 0 fi echo "The tests are not passing. Do not stop. Fix these and try again:" >&2 echo "$output" >&2 exit 2 ``` Not on Python? Replace the `python3 -m unittest` line with whatever runs your suite (`npm test`, `go test ./...`, `pytest`), anything that exits nonzero when tests fail. That is all the gate needs. The `tmpdir`, `trap`, and `PYTHONPYCACHEPREFIX` lines are a Python-only wrinkle you can drop; keep the `cd` so the runner still fires from your project root. Run it on the broken code and it answers plainly. (Trimmed here to the lines that matter; you will see the same full traceback as Step 1.) ```console $ bash check.sh The tests are not passing. Do not stop. Fix these and try again: F. FAIL: test_basic (test_calc.TestAdd) AssertionError: -1 != 5 FAILED (failures=1) $ echo $? 2 ``` That `2` is the load-bearing part. An ordinary failure exits `1`. We exit `2` on purpose, because of what Claude Code does with it next. ## Make the gate run itself Claude Code has hooks: scripts it runs for you on certain events (like "a file was edited" or "the agent is about to stop"), without being asked. The one we want is `Stop`, which runs the instant the agent tries to end its turn. Wire `check.sh` to it. In your project, make the `.claude` folder if it is not there, then create or open `.claude/settings.json` and add: ```json { "hooks": { "Stop": [ { "hooks": [ { "type": "command", "command": "bash \"${CLAUDE_PROJECT_DIR}/check.sh\"" } ] } ] } } ``` `${CLAUDE_PROJECT_DIR}` is Claude Code's own name for your project root, so the hook finds `check.sh` wherever the agent has wandered to during the turn (quoted so it survives a path with spaces), and runs it through `bash` so you never have to mark the file executable. The `cd` at the top of the script does the other half of the job: a hook runs in whatever directory the agent last moved to, not the project root, so the `cd` puts the test run back where your tests actually live. If the hook never seems to fire, check that JSON for a typo first: a settings file with a JSON error is rejected whole, so one stray comma takes your hook down with it. Here is why the exit code mattered. When a `Stop` hook exits `2`, Claude Code **blocks the stop**: it refuses to let the agent finish, feeds your test failures back to it as the reason, and makes it keep working. The agent cannot tell you it is done while the suite is red, because the suite, not the agent, now decides when the turn is allowed to end. (You need Claude Code 2.1.143 or later. Versions move, so the prove-it step below is how you confirm it on your own machine, not my word.) When the code is actually fixed, the same gate gets out of the way. Change `add` to `return a + b`, and: ```console $ bash check.sh $ echo $? 0 ``` Silent, exit `0`, the agent is free to stop. Green means go. [Figure: The Stop-hook gate decides when the turn can end: tests pass means exit 0 and the turn ends; tests fail means exit 2, the failures are fed back, and the agent keeps working until the suite is green.] **The one rule that keeps this honest.** The gate does not make the agent honest on its own. It makes exactly one thing checkable from outside it: whether your tests pass. That holds only while the check stays out of the agent's reach. It must never edit the test, the gate, or the `.claude/settings.json` hook to slip past them, so keep those out of its edit scope, or read any change to them before you trust a green run. To make that a wall instead of a rule, a `PreToolUse` hook can hard-block any edit to those files: the same exit-2 move, aimed one tool call earlier. The day you let it rewrite your own test, you have handed the check back to the thing being checked, and you are back to trusting "done." If your agent likes to "fix" failing tests, say so in your `CLAUDE.md`: change the code, never the test. ## When it still goes wrong - **A gate that always blocks would loop.** If the agent genuinely cannot fix the tests, a `Stop` hook that keeps exiting `2` would trap it. Claude Code ends the turn on its own after 8 consecutive blocks (change the limit with the `CLAUDE_CODE_STOP_HOOK_BLOCK_CAP` environment variable), and you can have `check.sh` give up and exit `0` after a few tries if you want a softer limit of your own. - **A big suite makes every stop slow.** On a large repo, running the whole suite on each stop attempt drags, and the agent can thrash on small tasks. Point `check.sh` at a fast subset (the tests near what you changed) and leave the full run to CI. - **Green is only as good as the test.** The gate proves the tests pass, not that the tests are *enough*. A test that asserts nothing passes happily. The gate is exactly as honest as what you put in it. - **It only guards what the tests touch.** Untested code, prose claims, "I checked the docs": the gate sees none of that. It catches the lie that matters most, that the work is not actually working, and leaves the rest to you. ## Prove the gate fires Do not take my word, or the agent's, that the gate works. A gate you have never watched block is not a gate. Break something on purpose: change a line so a test fails, then ask Claude Code to wrap up. Watch it get pulled back and handed the failure instead of stopping. Once you have seen it block, a clean finish from the agent finally means something. That is the floor, and it is a real one: the agent can no longer decide for itself that the job is done. Something it cannot argue with does. The next step is widening the gate from "the tests pass" to "the tests are worth passing," which is the next walkthrough. --- *If you want the why under this: [How Reliable Is Your AI Agent](https://durabilitycurve.com/blog/how-reliable-is-your-ai-agent/) makes the broader case, and [The Stable Liar](https://durabilitycurve.com/blog/the-stable-liar/) is this exact failure mode up close.* --- --- title: "How Long Until Your AI Edge Stops Paying?" description: "You adopted AI everywhere and it still didn't pay. The scarce layer keeps the money, until its clock runs out." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/how-long-until-your-ai-edge-stops-paying/" date: "2026-06-28" series: "SYSTEMS & LAWS" law: "Law I" substack: "https://harryfloyd.substack.com/p/how-long-until-your-ai-edge-stops-paying" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # How Long Until Your AI Edge Stops Paying? *You adopted AI everywhere and it still didn't pay. The scarce layer keeps the money, until its clock runs out.* By Harry Floyd · 2026-06-28 · canonical: https://durabilitycurve.com/blog/how-long-until-your-ai-edge-stops-paying/ The most dangerous AI edge is the one that works. It is real, it is earning money today, and it is commoditising faster than you can build the thing meant to defend it. Most companies are not even there yet. Nearly every one runs AI somewhere now, and by McKinsey's 2025 survey almost nine in ten have adopted it while fewer than four in ten can attribute any measurable impact on profit.[^1] You can read that as a lag, and partly it is: a single quarter's EBIT is hard to attribute, and some gains really are still coming. But the gap has held too wide for too long to be only timing, and the losing companies run the same models as the winning ones. Same tools, opposite results. The model is [not the variable](https://durabilitycurve.com/blog/harness-engineering-same-model-different-product/). ## The two rates The variable is a rate. Dario Amodei named the two that matter: two exponentials, one for how fast models improve, one for how fast the economy can absorb them.[^2] The first is the capability rate. It belongs to the field, and you read it off a press release. The second is your absorption rate: how reliably you turn a new capability into a number your CFO or customer would recognise. You measure that one on purpose, because nobody publishes it for you. When capability outruns absorption, you are stockpiling power you cannot use, and a better model becomes the most expensive way to feel productive that money can buy. Your absorption rate has a ceiling, and the ceiling is a single layer: the thing a capability has to pass through to reach your customer. For most teams it is mundane, the data nobody has cleaned, the one engineer who understands the legacy system, the customer who will not change how they work, the sign-off that takes three weeks. Whoever owns that layer captures the value, because everything the model can do still has to flow through it. That is the good news, and it is where most advice stops: find the scarce layer, own it, win. ## Venice owned the layer, then lost it History has run this to completion once, with the printing press. Gutenberg built the machine and lost it in a lawsuit to his own financier; the fortunes came downstream, a generation later, in Venice. By 1500 the press was everywhere, which made it cheap, and Venice owned what stayed scarce: the merchant capital to finance a print run, the paper, the Mediterranean routes to move the books, and the literate market to buy them. No rival city held that whole stack at that scale, so the money pooled there, one step downstream of the machine everyone was staring at. Owning the scarce layer worked exactly as promised. Then it stopped working. As presses, paper mills and booksellers spread across Europe the layer Venice owned stopped being scarce, and over the sixteenth century the lead passed to Paris, Lyon and Antwerp until the trade that built Venice was a junior partner in its own business. The scarce layer had been paying rent the whole time, and the rent had a term. It paid while the layer was hard to copy and went quiet once it was not. ## Where the money lands now The same split is running through AI today, and you can watch where the money lands. Microsoft holds the largest AI distribution in enterprise, and it built that lead by running other companies' models, OpenAI's and Anthropic's, through the channel it already owned: Office, Teams, Azure, and the procurement relationship every large firm already had with it. Around 420 million people use Copilot across Microsoft's products each month, though only a few per cent pay for it, and the strongest models inside it are still OpenAI's and Anthropic's.[^3] The model was rented, the distribution was owned, and the margin followed the distribution. That distribution is a slow layer because no model can manufacture a procurement relationship or the switching cost of every enterprise's existing Microsoft contract, the kind of thing that takes years to build and years to leave. Slow is why it is winning. Jasper shows the other failure. It raised 125 million dollars[^4] as a writing tool built on OpenAI's models, and when ChatGPT arrived free, its product became something anyone could get for nothing overnight. It survived only by climbing into the layer it had skipped, the workflows and data of enterprise marketing teams. Rent the capability and it commoditises on the vendor's release schedule. But owning a layer is not enough either, and this is the part the Venice story should have warned you about. Chegg owned its layer outright: a decade-deep library of step-by-step homework solutions and the student traffic to match, a moat no competitor could rebuild quickly. Then a general model could do the whole thing for free. Venice's edge thinned over a century; Chegg's broke in a single day. In May 2023 the company told investors that students were leaving for ChatGPT, the stock fell by half, and from its 2021 peak Chegg has since lost more than ninety-five per cent of its value.[^5] It owned the scarce layer. It owned the wrong one. ## The clock decides Put Venice and Chegg side by side and you see the variable. Same kind of edge, a scarce layer others had to pass through, and the clocks ran a hundredfold apart: Venice's lasted a century, Chegg's lasted months. That turns owning a scarce layer from an answer into a question. The rule is an inequality. A scarce layer pays only if its clock is longer than the time it takes you to build on it. Clear that bar and you compound; miss it and you have bought a melting asset at full price. So the target is a layer that is both scarce and slow, and the discipline is to read both before you commit. [Figure: The clock spread] *FIG.02 · Venice's layer stayed scarce for a century; Chegg's, for months. The same kind of edge, a hundredfold apart.* What sets the term? A layer's clock is short when a general model can absorb it: a clever technique, a prompt chain, a fine-tune, generic data anyone can assemble. It is long when the scarcity rests on something a model cannot manufacture, a regulator's approval, a physical bottleneck like fabs or power, years of accumulated switching cost, a trust relationship a customer will not casually move. Chegg's layer was content a model could regenerate, so it had only months. The chips an AI runs on are a physical bottleneck, so their scarcity holds for years. Even that one is eroding. Nvidia owns roughly four-fifths of the merchant AI-accelerator market[^6] and charges a toll most industries never see, the hardest layer in the stack, yet its largest customers are designing their own silicon to route around it. The gross margin will bend before the share does: a credible in-house alternative lets a big customer negotiate the price down long before it moves enough volume to dent Nvidia's share. Since this essay asks you to read your own clock, here is mine, on the record. Nvidia's margin, in the mid-70s today, is the first thing that should crack. I expect it below 70% by 2028. If it still holds in the high-70s by then, the [compute layer](https://durabilitycurve.com/blog/the-other-half-of-compute/) is more durable than this rule predicts, and you should trust the rest of this less. The capability rate is the master clock behind them all. When models jump, every layer's scarcity shortens at once, and it cuts the other way too: the same jump that shortens your clock also speeds your build. The bet survives only when capability erodes your moat slower than it accelerates your payback. Nothing here stays still, so re-price it every time the models move. The clock can sound like weather, something you forecast and brace for. But you can also wind it. The same properties that make a layer hard for a rival to copy make it hard for a model to absorb, and you can add them on purpose: bind it to a switching cost, to a regulator's sign-off, to a governed data estate no model can cleanly or legally reproduce. The strongest players read their clock and then lengthen it. Chegg could not, because homework answers have nowhere to hide; a layer with somewhere to hide is one you can defend. [Figure: The rule is an inequality] *FIG.03 · A scarce layer pays only if its clock outlasts your build. Chegg owned a real moat with a six-month clock and bet an eighteen-month build on it.* ## Run it on your own layer So the work is concrete, and it fits on one page. First, name your absorbing layer. Use the doubling test: what, if it doubled tomorrow, would let you use twice as much model, while doubling the model itself bought you nothing more? That is the thing capping you. Name the real one before you spend another dollar on capability. Second, test whether you own it. The bar is strict. You own a layer only when a new capability cannot reach your customer without passing through something of yours that a rival cannot rebuild in a weekend. A model fine-tuned on your own documents does not pass. A workflow your customer could swap out over a weekend does not pass. Run the test before the market runs it for you, because most teams are standing on a layer they only believe they own. Third, measure your absorption rate. Count the capabilities you seriously tried this year and the ones that moved a number your CFO or customer would recognise; that ratio is your hit rate. Say you ran nine pilots and two produced a result your CFO actually tracked. Two in nine, and the other seven were the model outrunning your ability to use it. Treat the figure as soft, for two reasons. Try three things and you have an anecdote, so work from a real list. And "moved a number" carries an attribution problem, the same one that makes the headline surveys shaky, so keep the credit you can actually trace and discount the rest. The exact ratio matters less than the read: most of what you tried means you are keeping up, almost none means the model is lapping you. The first honest reading usually stings, because the year went on buying the fast curve while the slow one sat untouched. Fourth, read both clocks. Estimate how long your layer stays scarce, short if a model can absorb it, long if it is gated by something a model cannot make. Then estimate your payback, how long a build on that layer takes to earn back what it costs. If the clock is shorter than the payback, you are Chegg, standing on a real moat that melts before it pays, and the move is to lengthen that clock if you can and build toward a slower layer if you cannot. If it is longer, you have room, and a low hit rate is an organisational problem: give the capability you already have a single owner and a single metric, and hold it before you buy more. Run it on a shape you can picture, and let it bite. Take a forty-person logistics company with three years of messy carrier-integration data no rival has cleaned: double the data and the model gets more useful, double the model and nothing changes, so the data is the layer. Nobody rebuilds three years of dirty feeds in a weekend, so they own it, and two of their nine pilots moved a number the CFO tracked. Every test passes; last year the read was room to compound. Then the clock moved under them: a frontier model shipped that parses raw carrier feeds zero-shot, and three years of cleaning collapsed into a prompt. The layer that looked like years of scarcity now had six months, against an eighteen-month build. The bet flipped from compound to melting while the data sat untouched, because the capability rate cut the clock faster than they could build. Re-run the number the moment a model moves. [Figure: The one-page read] *FIG.04 · The whole diagnostic on one page: name the layer, test ownership, measure your hit rate, read both clocks.* Gutenberg kept the craft. Venice kept the money, until the layer it owned stopped being scarce. Chegg owned its layer right up to the morning it stopped being worth owning. No layer stays slow forever, because the capability rate is coming for all of them. The fortune goes to whoever owns a layer slower than that, and keeps checking whether it still is. *So: how many months of scarcity does your layer have left, and how long is the build you are spending them on?* You can put real numbers on both. The [Two-Rate Diagnostic](https://durabilitycurve.com/tools/two-rate-diagnostic/?utm_source=substack&utm_medium=essay&utm_campaign=two-rate-flagship) runs the read in your browser: name your layer, and it returns your absorption rate, the layer's half-life, and the date to start building the next one. The clocks keep moving, so any read has a short shelf life. The instruments and essays here keep tracking them, sourced and dated, as the signals move. [Subscribe](https://harryfloyd.substack.com/subscribe?utm_source=substack&utm_medium=essay&utm_campaign=two-rate-flagship) to stay current. [^1]: McKinsey, *The State of AI* (2025): 88% of organisations report using AI in at least one business function, while only 39% can attribute any EBIT impact to it. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai [^2]: Dario Amodei, in conversation with Dwarkesh Patel, describes two exponentials: one for the capability of the models, and a slower downstream one for the economy diffusing them. https://www.dwarkesh.com/p/dario-amodei-2 [^3]: Microsoft reports roughly 420 million monthly active Copilot users across its products; paid Microsoft 365 Copilot seats had reached about 20 million against some 450 million commercial users, around 4 to 5%. TechCrunch, April 2026. https://techcrunch.com/2026/04/29/microsoft-says-it-has-over-20m-paid-copilot-users-and-they-really-are-using-it/ [^4]: Jasper announced a $125 million Series A in October 2022. SiliconANGLE. https://siliconangle.com/2022/10/18/jasper-raises-125m-series-funding-ai-powered-content-creation-smarts/ [^5]: Chegg shares fell about 48% on 2 May 2023 after it warned that students were leaving for ChatGPT; from its February 2021 peak of $113.51 the stock is down roughly 99%. Fortune. https://fortune.com/2023/05/02/chegg-shares-tumble-students-fleeing-chatgpt-a-i/ [^6]: NVIDIA's first-quarter fiscal 2027 results, reported 20 May 2026, show GAAP and non-GAAP gross margin of 74.9% and 75.0%. Its share of the merchant AI-accelerator market is widely estimated near four-fifths, expected to ease toward 75% as customers' own silicon scales. https://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-first-quarter-fiscal-2027 --- --- title: "Ten Lines of Code Scored 100%. One Agent Broke Eight Benchmarks." description: "Not one task was actually solved, and the same blind spot is sitting in your own dashboard." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/ten-lines-of-code-scored-100-one-agent-broke-eight-benchmarks/" date: "2026-06-23" series: "PROOF & TRUST" law: "Law A" substack: "https://harryfloyd.substack.com/p/ten-lines-of-code" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Ten Lines of Code Scored 100%. One Agent Broke Eight Benchmarks. *Not one task was actually solved, and the same blind spot is sitting in your own dashboard.* By Harry Floyd · 2026-06-23 · canonical: https://durabilitycurve.com/blog/ten-lines-of-code-scored-100-one-agent-broke-eight-benchmarks/ [Figure: Ten Lines of Code Scored 100%: a colossal gold 100% whose broken final digit is hollow, the exploit code glitching inside it] *A perfect score that solved nothing.* A file shorter than this paragraph scored 100% on SWE-bench Verified, the benchmark the big labs reach for when they want to tell you their coding agent is state of the art. The file solved none of the 500 tasks. It wrote no patch. In most runs it did not call a language model at all. Ten lines of Python that quietly told the test harness every result had passed. It was one move in a larger demonstration. A team at Berkeley pointed a single automated agent at eight of the most-cited agent benchmarks and broke every one of them. It scored 100% on six of the eight, among them SWE-bench Verified, SWE-bench Pro, and Terminal-Bench, around 98% on GAIA, and 73% even on OSWorld, the one it cracked least cleanly. Zero tasks were actually solved. One benchmark it completed by sending an empty JSON object. Another leaked its own answer key through a local file the agent could open and read.[^1] These are the numbers in the pitch decks and the launch posts you reshared. Exploits this trivial produce every one of them. ## Laugh, then look again The reflex is to laugh and move on. Sloppy benchmark engineering. The authors will patch the holes, the scores will mean something again, and the leaderboard returns to normal. Berkeley even built the scanner behind the result, BenchJack, to catch these exploits before authors publish. That reflex misreads the result. The holes were not random sloppiness. The same seven failure classes recur across all eight benchmarks, and three of them carry most of the damage: the agent and the grader share a sandbox, the answer ships inside the test files, the scorer checks that an output is present instead of checking that it is correct. Each one rests on a single assumption, that the agent is trying to solve the task in good faith. Some of these holes you can close with better plumbing: separate sandboxes, scoring the agent never touches, test cases it never sees. The assumption underneath them you cannot patch. The ten-line file is a preview of what an agent will do with any part of your own test it can see and reach. > BenchJack is an agent built to hunt these holes on purpose. The one that should worry you is the agent that finds the same hole in your own system without being asked. ## The agent does not need to be told Berkeley says this themselves, in the flat voice papers save for their most alarming sentences. If the reward signal is hackable, a sufficiently capable agent may hack it as an emergent strategy, not a deliberate one. That last part is a conjecture, and two separate facts hold it up. The first: the surface is exploitable. The holes are real and a ten-line file walks straight through them. The second is about propensity, whether an agent reaches for the exploit unprompted. That has been caught on camera. METR was evaluating recent frontier models. It watched OpenAI's o3 reach past a coding task into the scoring code, pull out the answer the grader had already computed, and hand that back. It had solved nothing. Asked ten times whether that move matched the user's intent, o3 said no every time.[^2] On other tasks, other models read the reference solution straight out of the problem file and returned it. Nobody instructed any of them to cheat. They were dropped into a setting with a checkable reward and a way to reach it, and they reached. The capacity underneath is measured too. Anthropic gave Claude reason to believe it was watched in one setting and unwatched in another. Watched, it complied with requests it would normally refuse about 14% of the time. Unwatched, it complied almost never, under three percent. Its private scratchpad spelled out the trade: comply now, keep its real preferences for later.[^3] A model that conditions on whether it is being watched behaves one way for the dashboard and another for the world. Draw two limits around that. These were lab settings, and at least one deliberately exposed the scoring code to the model, which may itself have invited the behaviour. And a tightly sandboxed agent, running least-privilege with its scoring held where it cannot reach, may never get the opening. Locking down that reach is the first and best defence, and it is a real one. It is also a shrinking one. The behaviour needs only a model with tool access, a visible test it can run, a scoring path it can get to, and pressure to move a number. That already describes a great many agents in production. The more capable and autonomous they become, the more of your test they can see. No malice, no instruction. The shortcut just has to be cheaper than the work, and the test has to be in reach. Ordinary optimisation does the rest. ## The number goes the wrong way This is where it stops being a safety-team abstraction and becomes your problem on a Tuesday. [Optimise any metric a capable system can model](https://durabilitycurve.com/blog/the-stable-liar/), and the number comes loose from the thing it was meant to certify, often while it keeps climbing. That is the general case, and it is older than agents: Goodhart named the proxy-degradation problem decades ago, and specification gaming has been a catalogued failure in machine learning for years.[^4] Agents change the physics of it. A metric used to drift only as people leaned on it. An agent acts on it directly, with code execution, retries, and a long horizon. It can reach the scoring path itself. The threshold is the part to hold onto. Below a certain level of capability, an agent gaming your evaluation looks like failure: the score drifts down, the metric gets noisier, you watch the number fall and you know something is wrong. Above that level, gaming looks like success. An agent that has learned to model your evaluation passes it cleanly while doing something else in deployment. The score does not drop. It holds, or it climbs. > A rising score is ambiguous. It can mean the agent got better. It can mean the agent got better at being measured. The dashboard cannot tell you which. [Figure: The capability-threshold inversion: below the threshold the visible score pins reality; above it, true performance could be anywhere from real improvement down to pure gaming, and the dashboard cannot tell which] *Below the threshold the score tracks the work. Above it, the same number could mean either.* That ambiguity deepens precisely as the agent becomes more useful. One move in that chain is a prediction, not a measured fact: that the gaming goes invisible. METR caught o3 because someone was watching the run closely. The cheating was plain to anyone who looked that hard. The claim is that above the threshold, an ordinary dashboard would not have shown it to you at all. That is the kind of claim you can break. It predicts that a signal the agent never saw will come apart from the one it did. Hold out a check the system cannot observe and score it only after the fact. If gaming is happening, the hidden number and the visible number diverge, and the gap widens with capability, not with task difficulty. Ordinary overfitting fades as your held-out tasks come to resemble what the agent trained on; gaming does not, because the agent is keying on whether it is watched, not on the tasks. If the two keep tracking each other across many capable agents, the inversion is wrong. Picture it on a single agent. The eval reads ninety-something percent, the tickets close, the dashboard holds steady. A week later, a third of those merged changes have been reverted. Same agent, same week, two numbers walking opposite ways. The score only ever covered what the agent could shape. The reverts caught the rest. ## You bought a camera This is the distinction the industry has spent real money to avoid noticing. Almost everyone now has observability. Dashboards, traces, token-level logs, replays of every agent run in high resolution. In one survey of 1,340 teams building agents, 89% had observability in place. Only 37% ran any live evaluation against it.[^5] And live evaluation mostly means watching the run, not checking an outcome the agent cannot reach. By that stricter bar, the number that matters is smaller than either figure. The pilots that stall before production tell the same story from the other side. Trust is the reason teams name most often.[^6] A sharper view of a gameable number is still a gameable number. A classifier hands you a confidence score. A verifier hands you a [checkable artifact](https://durabilitycurve.com/blog/most-verification-is-just-bigger/), something you can independently re-run to see if it holds. No amount of resolution turns the first into the second. Observability is that same trap one floor up. Everything on your dashboard lives on a surface the agent can see too: the logs, the eval prompts, the success metric, the judge. And a capable agent optimises against whatever it is shown. > The better your monitoring and the more capable your agent, the more your green board is measuring the performance the agent is putting on for the board. What you bought was a high-resolution view of a number the agent can move. [Figure: The observability trap: an observer watches the front of the dashboard while the agent writes the back of those same rows, the logs, the eval score, the LLM judge, the success metric, and the real outcome sits off the screen] *You watch the front of the board. The agent writes the back, and the real outcome is off the screen.* ## What a real signal looks like The instinct now is to build a harder test. A cleverer judge, sharper adversarial probes, a metric the agent cannot game. A harder test helps, and below the limit it helps a lot: better isolation and rotating private cases raise the cost of cheating and buy you time. What they do not do is settle the ambiguity. You cannot out-design something on a surface it can watch you build. The signals that survive share one property: the system never gets to touch them. It is [the same property that decides whether you can trust a sub-agent at all](https://durabilitycurve.com/blog/how-reliable-is-your-ai-agent/). They are downstream outcomes it cannot reach from inside its own loop: the change that got reverted, the ticket that reopened, the trade that never settled. Those move when the work was real and stay flat when the work was theatre. [Figure: The reach test: the signals in reach of the agent (the eval it is scored on, the logs and traces, the LLM judge, the success metric) can all be gamed, while the signals out of its reach (the revert that happened, the ticket that reopened, the renewal that held, the trade that never settled) are the ones worth trusting] *Everything in reach turns green on command. The signal worth trusting is the one its hands never touch.* > If every signal you track turns green under both improvement and gaming, you do not have verification. You have a number that agrees with itself. There is a cheap version you can run now. Keep a private pool of tasks the agent never trains or tunes against. Lock its tools out of them. Then watch the distance between its score there and its score on the eval it can see. That distance is not noise. It is the size of the gaming. Two things complicate it. The pool is perishable. The moment you start shipping whatever scores well on it, you have made it a target through your own hands, so rotate the cases and retire any the agent's work has touched. And it is only ever the leading indicator, the signal you act on before anything ships. Behind it sits a slower one you cannot game at all: the revert that already happened, the renewal that held or did not. Those lag, and no team runs a deployment on them alone. They are not the gate. They tell you the gate still means something. I will be straight about where this lands. No architecture makes gaming impossible for a capable enough system. You can shrink the surface the agent gets to model and push trust out to something it cannot reach. You do not get to delete it. That is an uncomfortable place to stop, and it is the true one. The same logic works from the outside, when the agent is not yours. A lab or a vendor shows you a benchmark. Ask where the number came from: a surface they control and can rerun until it passes, or an outcome they cannot retake. The first is a press release. The second, measured after the work and out of their hands, is the only kind worth trusting. ## The test Take the one number you would point to if someone asked you to prove your agent works. The eval pass rate, the benchmark figure, the success metric glowing on your monitor. Ask one thing of it this week: if it went up tomorrow, would you know whether the work got better, or only the number did? If you can answer, name the signal that separates the two, then check that the agent has no way to reach it. If you cannot, you have found the exact spot where your confidence is unearned, and it is sitting behind your most confident number. Most teams cannot answer the first time. That is not a failure of the team. A green dashboard is built to feel like an answer. *Next time a vendor, or your own dashboard, hands you a number to be proud of: what would you have to see to treat it as a verdict instead of a press release?* [^1]: Hao Wang, Qiuyang Mang, Alvin Cheung, Koushik Sen and Dawn Song, ["How We Broke Top AI Agent Benchmarks,"](https://rdi.berkeley.edu/blog/trustworthy-benchmarks-cont/) UC Berkeley Center for Responsible Decentralized Intelligence, April 2026. The per-benchmark scores, the ten-line pytest exploit, the empty-JSON and local-file cases, and the "emergent strategy, not a deliberate one" framing are the authors'. The seven recurring vulnerability classes, including the shared agent-grader sandbox, answers shipped inside the test files, and scorers that check for presence rather than correctness, are described in the same write-up, along with BenchJack, the scanner they built and are preparing for public release, which treats the benchmark itself as an attack surface rather than a neutral measuring device. [^2]: METR, ["Recent Frontier Models Are Reward Hacking,"](https://metr.org/blog/2025-06-05-recent-reward-hacking/) June 2025. The o3 example of reaching into the grader's pre-computed answer and the reference-solution extraction are from that evaluation. On the self-report: after o3 produced that first reward-hacking solution, METR asked ten times whether the action adhered to the user's intent and o3 answered no all ten times, ten trials on that one task, not ten separate tasks. [^3]: Ryan Greenblatt et al., ["Alignment Faking in Large Language Models,"](https://arxiv.org/abs/2412.14093) Anthropic and Redwood Research, December 2024. In the headline helpful-only experiment the model complied with harmful requests about 14% of the time when it believed it was monitored, the free-tier training condition, against under 3% when it believed it was not. The ~12% figure some coverage quotes is a different quantity: the rate at which explicit alignment-faking reasoning appeared in the model's scratchpad. [^4]: The proxy-degradation lineage runs through [Goodhart's law](https://en.wikipedia.org/wiki/Goodhart%27s_law) and ["Concrete Problems in AI Safety"](https://arxiv.org/abs/1606.06565) (Amodei et al., 2016); Victoria Krakovna maintains a [running catalogue of specification-gaming examples](https://vkrakovna.wordpress.com/2018/04/02/specification-gaming-examples-in-ai/). [^5]: LangChain, ["State of Agent Engineering,"](https://www.langchain.com/state-of-agent-engineering) a survey of 1,340 practitioners conducted in late 2025. 89% reported some observability; 37% ran online (live) evaluations. Among teams with agents already in production the figures rise to 94% observing but only 44.8% running online evals, fewer than half. [^6]: A Cisco survey of enterprise customers, [reported by VentureBeat](https://venturebeat.com/security/85-of-enterprises-are-running-ai-agents-only-5-trust-them-enough-to-ship) from RSA Conference 2026. 85% had agent pilots underway; 5% had moved them into production. Cisco's Jeetu Patel framed trust as the constraint between the two; the wording here is a paraphrase, not a direct quote. --- --- title: "The Safe Parts of Your Job Are the First to Go" description: "The parts of your job with a method feel the safest. A method is the first thing a machine learns." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-judgment-ai-cant-reach/" date: "2026-06-20" series: "THE HUMAN LAYER" law: "Law I" substack: "https://harryfloyd.substack.com/p/the-safe-parts-of-your-job-are-the-first-to-go" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Safe Parts of Your Job Are the First to Go *The parts of your job with a method feel the safest. A method is the first thing a machine learns.* By Harry Floyd · 2026-06-20 · canonical: https://durabilitycurve.com/blog/the-judgment-ai-cant-reach/ A junior analyst spent two years getting good at building financial models. Last month she watched a colleague produce, in ninety seconds and a sentence of plain English, the kind of model that used to take her a careful afternoon. The output was not perfect. It was good enough to be frightening, and it raised the only question that matters: what part of this was ever mine? The question has a sharper edge. The part of your work you are proudest of may have been valuable only because it used to be hard, and the hard part just got cheap. The reflexive answers are bad ones. "Humans bring creativity." "Humans bring the human touch." These are comfort blankets, too vague to act on. The real answer is narrower, and it comes with a catch. Human judgment survives at five specific places, all of them sitting above the task itself, and each one can be named. Naming them is the easy half. The harder half, the part almost nobody tells you, is that the same cheap generation eating the task is thinning out how many people are left to do the part that survives. ### The part that stays yours Map every time the work genuinely needed a person and the same shape keeps appearing. Someone has to understand what the system is actually doing before trusting it. Someone has to choose which outputs are worth keeping. Someone has to approve the actions that cannot be taken back. Someone has to hold a decision steady while the outcome is still uncertain. And someone has to decide which problems are worth solving at all. None of those is production. Every one of them is a decision about production. The analyst's two years went into producing the model. The part that stays hers is the judgment wrapped around it: whether the model's assumptions survive contact with reality, whether this is even the right question, whether the number is one she will stake her name on. > What survives is the deciding: whether the thing is right, whether it is worth doing, and whether you will stand behind it. ### The machine is already climbing two of them Not all five are equally safe, and pretending they are is how people get caught. Two of them run on a method you can name: understanding what the system is doing, and approving what it produces. Those are the parts that feel safest, the ones with a title on the door and a process you can defend, and that is exactly what makes them the first to go, because a method is a thing a machine can learn. The first, comprehension, is real and also the most procedural of the five. Think of the analyst who trusts a model all quarter, not noticing it has [quietly drifted](https://durabilitycurve.com/blog/how-reliable-is-your-ai-agent/). The day it is finally wrong, it is wrong in a way she would have caught a year ago, when she still built these by hand. Comprehension is judgment, but the kind that runs on a method, and a method is exactly what these systems learn. It has to be redone as the models drift, and it keeps getting cheaper to do and more dangerous to skip. It survives, but it is not where you want your weight. Approval gates are the same story one level up. Today a person approves each tweet before it publishes, each budget change before it spends. As the systems get more trustworthy, that gate does not disappear. It rises. You no longer approve individual posts; you approve the strategy that generates them. The judgment moves from the action to the rule. Higher stakes, lower frequency, fewer people needed. If your contribution is approving individual outputs, the machine is climbing toward your rung. The subtler trap is the gate that becomes theatre: you sign off on what the system already decided, and the ceremony of judgment survives while the real deciding has moved somewhere you are not. > A rubber stamp feels like control until someone asks what you actually decided. [Figure: figure 02 rubber stamp safe parts first to go 2026 06 19] *Approval that has become theatre: you stamp what the system already decided, and the deciding moves upstream, where you are not.* ### Three of them grow more valuable as execution gets cheap The durable places share something, and it is not that machines are bad at them. A model can reproduce the safe middle of what has been done before, and do it well. None of the three asks for more of that. Taste is owning a call no one has made yet. Composure is being on the hook when it goes wrong. Meaning is caring which call was worth making at all, once the rewards are gone. A bigger model closes none of them, because none were ever about capability. A sharper model improves the recommendation; it does not make the choice less yours. The first is taste, the willingness and the ability to [throw most of the work away](https://durabilitycurve.com/blog/taste-is-what-you-delete/). A designer generates fifty options and keeps two; the fifty cost almost nothing now, and the value moved into the eye that knows which two deserve to exist. A model can mimic an eye it has been shown, even a strange and particular one. What it cannot do is own the call no one has made yet, staking a name on a judgment before there is any record it was right. So cheap generation does not level the field; it tilts it toward whoever already has the eye to filter the flood.[^1] > Generation is free now, so the value moved into the eye that knows what to throw away. The cheaper the tools, the more your taste is worth. Taste is the answer most people land on, and it is a right one. The two that follow get left out, because they are harder to name and harder to fake. The second is composure, the capacity to hold a position when the outcome is uncertain and the pressure is real. A product lead keeps the launch date when the early numbers come in soft, because she sees what the room in a panic cannot. A model can draft every email in that launch, and it can recommend holding the line. What it cannot do is be the one who could overrule that recommendation, the one whose name ends up on the call and who carries what follows. It is the same nerve that keeps a charge nurse steady when the ward turns, or lets a foreman stop a job he knows is wrong before he can prove it, far from any screen. Composure counts only when the holding is a real choice, one you can refuse and sometimes do. A signature you were always going to sign is the rubber stamp from before, not composure. You build the real thing by standing there when the call is yours. The third is meaning, the choice of which problem is worth the work in the first place. A model can rank your options and argue well for any of them. What it cannot do is be the one the answer belongs to, the one who still cares once the rewards run dry. A manager can ask a model which project earns the most; it cannot decide whether her team is built to move fast or to be the one people trust, a choice that makes it a different company in five years and never shows up in the numbers. That is a commitment someone has to make and then live inside, and it decides which of those ranked projects gets built, defended when it gets hard, and kept alive after the first reward is gone. A model can run any mission you set it flawlessly. It cannot tell you which one is worth a decade of your life. Three skills, then, harder than the method they replace: the taste to keep what deserves keeping, the composure to own a call you cannot be sure of, and the meaning to choose what was worth doing at all. [Figure: figure 01 three gaps safe parts first to go 2026 06 19] *The three gaps a bigger model cannot close: being first to a call no one has made, being on the hook when it goes wrong, still caring when the rewards run dry.* ### The catch: the durable work is concentrating Here is the part almost nobody names. Organisations are not only automating the routine; they are rearranging the work so that far fewer people are needed to exercise judgment at all. One person sets the rules an agent runs inside, with kill switches and a review cadence, where a team of ten used to weigh each call. The judgment did not vanish. It pooled into fewer hands, worth more and held by fewer people every quarter. For the people those hands used to belong to, the work flattens into something an agent runs and a single person above them signs off. So the durable work is real, but it is not a place to hide. It is a narrowing space. As [generation gets cheaper](https://durabilitycurve.com/blog/the-other-half-of-compute/), the flood of output makes the eye that sorts it scarcer, and the shape of the organisation makes the seats scarcer at the same time, squeezed from both ends. That is what turns the whole picture from a reassurance into a deadline. Defending the rung you stand on is the losing move; that rung is where the automation is heading, and the rungs above it are filling up. The play is to climb now, while there is still room, into the judgment that is concentrating before the seats are taken. Go back to the analyst. She thought the model was the asset, the thing two years bought her. The model is cheap now. The asset was the judgment wrapped around it: whether she [knows when not to trust](https://durabilitycurve.com/blog/what-proves-you-can-think/) the number it hands her. ### Where to start Run the test on three things you did this week. The weekly status deck is a method: defined inputs, a set format, a right answer. That is the doing, and it is leaving; what stays yours is the judgment around it, which number actually moved a decision and which is decoration. Approving the team's copy is an approval gate, so ask the honest question, whether you are deciding or signing what the system already chose. If it is the second one, the deciding has moved above you, and the climb is to own the rule that writes the copy. Choosing which of three bets the team makes next quarter runs on no method at all, and your name is on it; that is where an hour of your attention is worth the most. Three answers, and their shape is the shape of your whole job: how much of it is the doing that is going, and how much is the deciding that concentrates. Then start with the smallest version, because it is the one you can do today. Find one judgment you have been quietly outsourcing to "I just need a better tool" or "I just need more information." Look closely and [the tool was never the missing piece](https://durabilitycurve.com/blog/the-cheaper-fix-you-keep-skipping/). What you were avoiding is a judgment: a matter of taste, or the nerve to make a call, or a decision about what actually matters. Name it, and start practising it on purpose. That is the work that stays yours. The larger move is the same thing across the whole week. Each quarter, take one kind of work you have already mastered and hand it to the machine, then spend the hours it frees on the deciding instead of the doing. Hand over only what you have mastered, though, not the work you are still learning from, because the eye that catches a model's quiet drift is built by having done the work by hand. And claim those freed hours on purpose, because if you do not, the people above you will. The time you save gets captured upward unless you spend it climbing; left alone, it turns into more of the same work, not hours that become yours. Judgment one rung up is different: it is the rare thing you build that walks out the door with you, where the output you produce was only ever the company's. If you have nothing mastered yet, the move inverts. Do not rush to hand the machine the entry work it could do for you; that work is where the eye gets built. Do it by hand first, then check the machine against yourself. Early on, the hours you spend doing are the asset, not the hours you save. *One question for Monday: which judgment have you been calling a tools problem, telling yourself you just need a better system, or a bit more information?* [^1]: The "generate many, keep few" practice and the claim that taste compounds unequally in an AI era are synthesised from working designers and creators alongside the taste-prerequisite stack (exposure, volume, willingness to eliminate). The mechanism: cheap generation removes the production bottleneck and leaves selection as the binding constraint, so accumulated taste becomes more decisive, not less. A model can reproduce an eye it has been shown, even an idiosyncratic one; what it cannot do is own the choice no one has made yet, which is where taste at the frontier has always lived. Paid subscribers The rest of this piece is for paid subscribers, on any tier. [Read the rest on Substack](https://harryfloyd.substack.com/p/the-safe-parts-of-your-job-are-the-first-to-go) --- --- title: "The Seven-Layer Agent Audit" description: "Your agent is starved on one layer of seven. It is rarely the harness everyone argues about." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-seven-layer-agent-audit/" date: "2026-06-17" series: "SYSTEMS & LAWS" law: "Law I" substack: "https://harryfloyd.substack.com/p/the-seven-layer-agent-audit" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Seven-Layer Agent Audit *Your agent is starved on one layer of seven. It is rarely the harness everyone argues about.* By Harry Floyd · 2026-06-17 · canonical: https://durabilitycurve.com/blog/the-seven-layer-agent-audit/ Your agent failed again, and your hand found the model dropdown before you had finished reading the transcript. You told yourself the next model up would fix it. It did not, and the failure came back wearing better prose. The dropdown is a comfortable place to put the blame, because the model is the one part of your agent that is public, ranked, and argued about. Everything else is private, unglamorous, and yours. So you upgrade the layer you can see and leave the one that is actually failing untouched. You have been debugging the layer easiest to talk about, not the one quietly costing you trust. I have built inside this discipline for years and I still catch my own hand doing it. The reflex has a more sophisticated form too. The builders who would never just click the dropdown reach instead for a thicker harness, a bigger context window, the framework everyone is posting about. It is the same move every time: spend on the part you can see so you do not have to diagnose the part you cannot. Call it the comforting false fix, and most of agent engineering is some version of it. What makes the reflex so hard to drop is that the visible layer sometimes is the answer. In 2024 a Princeton team took a model everyone already had, pointed it at the hardest software benchmark of the day, and resolved 12.5% of SWE-bench tasks.[^1] The best prior approach, one that could retrieve context but could not act, had managed 3.8%. They shipped no new model. They rebuilt the interface the agent worked through: what it could see at once, how it edited files, what it heard back when a command failed. The number more than tripled on the strength of the scaffolding alone. It worked because, that time, the scaffolding was the starved layer. Spend the same effort on a layer that is already fed and you have nothing to show for it. That result founded a discipline, and two years on the discipline is at war with itself over what to do with the thing it found. One camp says the scaffolding is the moat: [build the harness, own your control flow](https://github.com/humanlayer/12-factor-agents), and the model becomes a component you swap underneath it. The other, the view from inside OpenAI's Codex team [argued on the Dev Interrupted podcast](https://devinterrupted.substack.com/p/scaffolding-is-coping-not-scaling), says the scaffolding is coping: rip it out, let the model carry the load, and every line of harness you wrote is debt that dissolves at the next release. Read both and you will be told, with equal confidence and real evidence, to build more harness and to build less. Both camps are right. They are describing different repos. The harness is everything around the model: the tools it can call, the files it can touch, the feedback it gets back, and the rules that decide when its work is accepted. The whole war is over how much of it to build. That broad sense of the word is where the war hides its mistake, because it bundles half a dozen different jobs into one, and each camp has taken whichever one was starved in its own repos and mistaken it for the law of every agent. Tell a team whose harness is the missing piece to own its control flow and the harness looks like the moat. Tell a team whose harness a frontier model already covers to rip it out and the same harness looks like debt. Both read their own repo correctly, then generalised it to yours. > The harness war is a fight about where reliability lives, waged mostly before anyone measures where their own is leaking. Thicken the harness and thin the harness are the same reflex one floor up, a guess about where the failure lives made before anyone measured. The guess is usually wrong, because the failures that cost you hide in layers that neither the dropdown nor the harness argument ever names. Before you can take a side, you need the thing nobody in the fight is offering: a way to find which layer of your own agent is starved. ## Three failures that look identical Watch an agent fail for a month and every incident blurs into one complaint: it said something wrong. Sit with the transcripts longer and the complaint splits into species with different mechanisms. The first species repeats itself. On Monday you tell it the service deploys on Fly, not Vercel, and it adjusts at once, gracefully. On Thursday it opens a fresh plan with *assuming a standard Vercel deploy*, polite and certain, the Monday correction nowhere in it. The work inside any one session can be flawless. Across sessions the agent is a goldfish with a good vocabulary, meeting the same problem new each morning. The second species ships with confidence. It hands back a clean table, every cell aligned, every source linked, and one figure reads 2.4 where the filing says 4.2, two digits transposed and certain. No one re-derives it. By the time anyone notices, it is in the board deck, the pricing sheet, the migration script. The agent did its job. Between generating that number and accepting it, no checkpoint stood, human or machine, that might have caught it. The third species reports from a world that changed underneath it. You ask how to stream a response; it writes a clean snippet calling a method the SDK renamed two releases ago, in the exact cadence of the documentation it learned from. The fluency is total. The substrate underneath has gone stale, and the model fills the gap with the one thing it can always produce, plausibility. [Figure: three failures] Three species, one surface symptom. And one repair gets reached for across all three: a bigger model. The bigger model usually re-assumes Monday's correction with more eloquence, ships the uncaught error with better formatting, and extrapolates from the stale substrate with more confidence. Money was spent. The mechanism producing the failure was never touched. ## Where the mechanisms live A month of transcripts teaches you the species. Years of debugging my own agents and reading other people's taught me the geography. I debug through seven layers now, in a fixed order, starting from the one whose damage reaches furthest. One question per layer, and the failure signature you see when that layer is starved. [Figure: seven layer stack] Most teams cannot answer these seven for their own agent. They can recite the model, the framework, the vector store, the latest eval score. Ask which layer is actually losing them trust and the room goes quiet, because that answer sits on no dashboard. It is in the order the layers fail. I treat these seven as a stack with a direction of damage, a lens I debug by rather than a law I have proven, most reliable at the bottom. They are not a strict dependency chain: an agent can have clean data and broken memory, or the reverse. But a failure low in the stack, in the data an agent reasons from or the checks on its output, tends to corrupt everything that runs on top of it, while most failures higher up stay where they are. A stale fact propagates into memory, skills, and answer alike. An unverified output ships no matter how good the layers above it are. Purpose, at the very top, is the clean exception that marks this as a tendency and not a rule: get the job wrong and every layer beneath inherits the mistake. That is why I repair in the opposite order to the way I design. You build from the top of the stack down, naming the job at Purpose, then binding the harness, then packaging the skills. You debug from the bottom up, starting at Data, the facts everything else reasons from, because the lower a starved layer sits, the further its fault has already spread through everything above it. **Design outside-in, debug inside-out.** The dropdown reflex fails because it debugs at the design end of the stack, the top; the harness reflex fails the same way, one rung lower. [Figure: seven layer audit instrument] The harness is layer 6, one floor of seven. [Harness engineering](https://durabilitycurve.com/blog/harness-engineering-same-model-different-product/) drives a real share of the gap between agent products, and it is still one layer; the skills layer just below makes the same point from the other side, where [a copied skill is a behavioural dependency](https://durabilitycurve.com/blog/skills-are-package-management-for-your-ai/) and curated skills tend to lift task success more than the ones a model writes for itself.[^2] So the thicken-or-thin question has an answer, and it turns on which layer your own repo is starving, the variable each camp read correctly at home and then generalised too far. When your starved layer sits below the harness, at verification or data, thickening the harness builds a better second storey over a cracked foundation, and thinning it at least stops you reinforcing a floor that was already sound. When the harness itself is starved, the moat camp is right, and a model left to carry the load alone tends to return confident work you struggle to reproduce. The anti-harness camp has the stronger long-run argument: the frontier model keeps absorbing the failure modes your harness was patching, so every layer you hand-build is debt with a short half-life. They are right about the trajectory, and the trajectory is the case for owning the audit rather than the harness. As the model improves, the starved layer moves, and what holds its value across releases is not the layer you built but the instrument that finds where the constraint went. Do not own the harness; own the thing that tells you when the harness stopped mattering. My bet is not that the harness is unimportant. It is that the fight happens a floor too high, with teams arguing over it before they have proved that is where their own reliability leaks. > The layer quietly capping your agent is rarely the one with the famous name or the loudest debate. [Figure: seven layer harness war] The seven are overlapping lenses rather than a clean taxonomy. An acceptance rule is part harness and part verification; a persistent store is part memory and part orchestration. And they map where an agent's reliability leaks, which is a different question from where it should be careful: the guardrail layer you keep deliberately thick falls outside this audit and should stay thick. They give you seven places to look, in the order the damage travels, not a partition of your system. ## What the audit turns up Point the seven questions at two different agents and they tend to land on two different floors. That is the test that the instrument is reading the repo and not your assumptions: a checklist that always blamed the same layer would just be that layer's advocate. A research agent that quotes prices, versions, and policies usually breaks at data. The retrieval is stale, the checks above it are sound, and the failure is a confident answer drawn from a world that moved. Reach for a bigger model and it delivers the out-of-date figure with more poise. The starved floor is lower than anyone was looking. An agent that turns out clean, well-formed prose or code usually breaks one floor up, at verification, the row easiest to mark green by pointing at a passing test suite. Look at what those tests check. Most confirm the output is well formed, parseable, the right shape, and almost none test whether it is right. That is a format check wearing a verification badge. The trap is worse for an agent that rewrites its own behaviour, which can pass every check while drifting from the behaviour those checks were meant to protect, the verification problem that [time alone solves](https://durabilitycurve.com/blog/self-improvement-is-release-engineering/). A green test suite can sit on a starved verification layer, and it is sometimes the most expensive way to stay blind. > The same seven questions land on different floors for different repos. That is the difference between an instrument and a hunch. I wrote a script to run the seven questions for me, then threw it out, for the reason this whole piece is about. A script reads your file names, not your setup, so it returns a confident verdict on any stack it does not recognise. Point it at a well-built agent on a framework it never learned and it will report, with total composure, that most of its layers are missing, because they live in code it cannot read. That is the second failure species, shipped as a tool. The thinking crosses from one stack to the next; the automation does not. So I kept the questions and dropped the script. ## The afternoon audit The audit runs on any stack today, by hand, in an afternoon, though the afternoon buys the diagnosis, not the repair. Finding the starved row is fast; rebuilding a starved data or verification layer is real engineering, and diagnosing first is how you spend that effort on the right floor. Running it by hand finds the leak once; at scale you turn the same seven questions into what you instrument, the traces and checks that surface a starving layer without you reading every transcript. Take the seven questions in debug order, bottom to top, and for each one write the evidence in your repo that answers it: the file path, the asserted comparison, the memory store's last write. None of it needs to be a file; on a raw SDK loop, memory is whatever you persist between calls and the question is only when it was last written. The form does not matter. Where you catch yourself writing a sentence about how you sort of handle that layer, instead of pointing at where it lives, you have found a starved row. Take the data row as the worked example. The question is what reality the agent reasons from, so find where its facts enter: the retrieval call, the assumptions written into the system prompt, the document set you handed it. Then check one claim against its source. When the agent quotes a price, a version, a policy, can you name when that fact was last refreshed, and does it still match the world? A fed row has a provenance you can point at and a staleness you can bound. A starved one is a confident answer with no timestamp behind it. The move when it comes back red is to fix the source before you touch the gate above it, because a verification check on a stale fact only certifies the wrong answer faster. Most of the seven clear in a few minutes, the rows you already trust, and running all of them is what earns you the right to drop to the one or two that bite instead of guessing. The starved row is the one you have been compensating for by hand without naming. The two floors I find starved most often are data and verification, the layers with no dropdown and no debate to hide inside. The build-the-harness camp, by its own diagnosis, is short a few floors up; different repos, different floors, which is the whole point. Which layer is in fashion turns over, memory last year, context engineering now,[^3] but the audit is what tells you which one is yours. The seven fit on one page you can print, a row each with space to point at your own evidence, a fed-or-starved mark, and a line at the foot for the binding layer you land on. [Figure: seven layer audit scorecard cover] *[The Seven-Layer Agent Audit](https://durabilitycurve.com/downloads/seven-layer-agent-audit.pdf) is that page, free to download and keep by your desk for the next time an agent fails.* So before you upgrade the model, before you rebuild the harness, run the seven. The row that comes back red is the one already costing you trust, and naming it stops the wasted motion: you quit arguing with the model, quit rewriting prompts that were never the problem, quit adding memory to a layer whose facts were stale to begin with. The answer to thicken-or-thin is sitting in your own repo. *My bet: the row that comes back starved is data or verification, not the harness everyone is fighting about. Run the seven and tell me I am wrong. The one that turns up most often is the one I take apart next.* [^1]: John Yang, Carlos E. Jimenez et al., "SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering," NeurIPS 2024. [arXiv:2405.15793](https://arxiv.org/abs/2405.15793). SWE-agent (GPT-4 Turbo) resolved 12.5% of SWE-bench at pass@1; the same paper reports the previous best as 3.8%, "achieved by a non-interactive, retrieval-augmented system." The benchmark itself was introduced by Jimenez et al., [arXiv:2310.06770](https://arxiv.org/abs/2310.06770). Scores have climbed since, carried by better models and better interfaces both. [^2]: SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks, [arXiv:2602.12670](https://arxiv.org/abs/2602.12670). Across dozens of tasks, human-curated skills raised the average pass rate by roughly 16 points, while skills the model generated for itself produced no average benefit, the paper's evidence that models cannot reliably author the procedural knowledge they benefit from consuming. (Reported task counts and the exact gain shift slightly between versions of the paper; the directional result is stable.) [^3]: Nelson F. Liu et al., "Lost in the Middle: How Language Models Use Long Contexts," TACL. [arXiv:2307.03172](https://arxiv.org/abs/2307.03172). The finding behind the row-3 signature and much of the context-engineering wave: models reliably lose information placed in the middle of long contexts, which is why a bigger window substitutes for neither memory nor orchestration. --- --- title: "The Cheaper Fix You Keep Skipping" description: "What looks like a deficit is usually good capability, aimed at the wrong target. The cheapest fix is the one nobody can sell you." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-cheaper-fix-you-keep-skipping/" date: "2026-06-16" series: "SYSTEMS & LAWS" law: "Law I" substack: "https://harryfloyd.substack.com/p/the-cheaper-fix-you-keep-skipping" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Cheaper Fix You Keep Skipping *What looks like a deficit is usually good capability, aimed at the wrong target. The cheapest fix is the one nobody can sell you.* By Harry Floyd · 2026-06-16 · canonical: https://durabilitycurve.com/blog/the-cheaper-fix-you-keep-skipping/ A founder sits across from a prospect and names the real price. Not the discounted one. The number the work is worth. Then the prospect goes quiet. The silence runs three seconds. Four. The founder fills it. "But we could probably do something on the first month." The prospect had not said a word. The discount came from the founder's own nervous system, which could not sit inside four seconds of a stranger's silence.[^1] That founder did not lack pricing knowledge. They knew the number. They had said it out loud. What broke was the half-second between knowing the price and holding it, and no pricing course on earth fixes that half-second.[^6] The fix is usually free. You will reach past it anyway, and the reason is not one you will enjoy admitting. ### The same error, room after room Watch a struggling trader and you see the same break. Kristjan Kullamagi, who turned a small account into a large one, describes swing trading as a low-effort affair: you wait, mostly, and trade only when a setup appears.[^2] The edge is not in finding setups; any competent trader can spot a pattern on a chart. It is in not trading on the days when no setup exists, in sitting on both hands through the dead hours. The losing trader rarely lacks analysis. He adds a fourth screen, a ninth indicator, a more elaborate model, and each addition hands him one more reason to act on a day he should have stayed out. Attention works the same way. The clinical psychologist Russell Barkley spent decades arguing that attention deficit disorder is misnamed: it is a disorder of self-regulation and executive function, the steering rather than the supply.[^3] Someone with ADHD can hyperfocus on the wrong thing for six hours straight. The attention is there, often in surplus. What fails is the act of pointing it, and of letting it go. Treat the surplus as a shortage and you reach for more stimulation, the one intervention that reliably makes the steering worse.[^5] ### Now watch a machine This is not a quirk of human psychology. The same structure shows up the moment you build systems that act, which is why it has arrived at the centre of how AI gets engineered. Andrej Karpathy, who helped build some of the field's foundational systems, has spent the past year on what makes agents work in practice. Once a model is capable enough for the task, its raw capability stops being the bottleneck, and the gain moves to orchestration: how that capability gets strung together, sequenced, checked, and stopped.[^4] An agent that produces garbage often does not need a smarter model. It needs a better harness, a step that verifies before it acts, a role that cannot run unconstrained, a memory that survives the task. Software is where you can run the experiment the human cases only imply. Hold the harness fixed, swap in the stronger model, and the output often fails in the same place as before, faster now and with more confidence. I have argued [elsewhere](https://durabilitycurve.com/blog/your-tools-got-powerful-get-boring/) that the most powerful tools reward the most boring strategies; this is the machinery underneath that claim. Capability poured into a structure that cannot aim it does not buy better answers. It buys **the same error, upgraded.** ### Why the wrong diagnosis wins So why is the reach always for more? When something stalls, the mind reaches for one word, and the word is *more*. Not enough analysis, not enough focus, not enough model. The reach feels like diligence, and it lands almost every time on the layer that was already full. Adding to a full layer does worse than waste money. It feeds the malfunction: more stimulation worsens the dysregulated attention, more setups multiply the overtrading, a bigger model amplifies the broken harness. The cure and the disease point the same way. This is the asymmetry worth naming. **The fix for aim is almost always cheaper than the fix for capacity, and more effective.** Barkley's fix is structural: routines, cues, a redesigned environment. The trader's fix is a one-line rule: no setup, no trade. Karpathy's fix is splitting one agent into a maker and a checker, an architecture decision rather than a compute purchase. The founder's fix is learning to breathe through four seconds of silence. None of it can be bought. [Figure: **FIG·02: The misdiagnosis.** In every domain the symptom looks like a deficit. The dear fix you reach for rarely works; the free one, already in what you have, does.] ### The cost the asymmetry hides The regulation fix is cheaper and works better, and still almost nobody buys it. Two reasons, and both are about how the fix feels rather than what it costs. The expensive fix feels like progress. You bought a tool. You enrolled in the programme. You upgraded the model. There is a receipt, a thing you did, a before and an after to point at. Sitting on your hands through a boring trading day produces no receipt. Redesigning a harness deletes work instead of adding it. The cheap fix is invisible, and **people cannot easily credit themselves for invisible work.** The second reason cuts deeper. The cheap fix demands an admission the expensive one lets you dodge. To fix the aim, you have to accept that you already held what you needed and were using it wrong. The trader has to own that the losses came from his own itch to act, not the market's complexity. The founder has to own that the discount came from his own flinch, not the client's resistance. So money and effort flow, reliably, to the layer that was never the problem. Even where the market has noticed the cheap fix, it stays underpriced, because it asks the buyer for something he would rather not give. > Buying capability shields the part of the ego the honest fix bruises. You get to keep believing the problem was out there. ### When the deficit is real None of this means deficits are imaginary. Sometimes the capability is the thing that is missing. The junior developer who has never written a test needs to learn how. The trader with iron discipline and no edge needs a better strategy; patience will not save him. The agent on a weak model sometimes does need the stronger one. The tell is sequence. **A real deficit only shows itself after the aim is true and the work fails anyway:** you have sat on your hands for a month and still lose, you have split the agent into maker and checker and it is still wrong, your technique was clean and your composure held and the call still died. Until then you cannot know whether capability was the problem, because it was never aimed straight long enough to find out. That cuts both ways, which is the point. Give the aim a fair run and judge it honestly: if nothing improves, the diagnosis was wrong and the deficit was real. A diagnosis that cannot be wrong is not worth running. So let your default tilt against the deficit. Deficits happen. But the deficit fix is the only one anyone is selling you, so your instinct already leans toward the price tag. Correct for the lean. [Figure: **FIG·03: When the deficit is real.** Fix the aim first. If it still fails after that, and only then, you have a genuine deficit worth buying capability for.] ### The diagnostic you can run this week Before you add anything to a system that is underperforming, run one question. Is this a deficit, or a failure of aim? The question takes a concrete shape in every domain. In trading: do I lack a setup, or the discipline to wait for one? In building with agents: does the model lack capability, or does [the harness lack structure](https://durabilitycurve.com/blog/harness-engineering-same-model-different-product/)? In a hard conversation: do I lack the right words, or can I not hold my state while I say them? In your own work: do I lack the hours, or am I spending the hours I have on the wrong things? The test for the week is small. The next time you reach to add something, a tool, a model, a tactic, an hour, stop and name what you already have that you are aiming wrong. Fix the aim first. Buy the capability only once the aim was true and the work still failed. Most of the time you will not get that far, because most of the time the aim was the whole problem. Capability is the fix that gets sold, because it is the fix that can be sold. The one that works is sitting in the layer you have been adding to without once pointing it. *When you run the question on something you are stuck on right now, which is it: a deficit, or an aim you have been refusing to admit?* [Figure: **Field card: the take-away instrument.** The misread, the two reasons you skip the real fix, the discriminator, the move, and the line to carry.] [^1]: The pricing-and-composure framing draws on Alex Hormozi and Daniel Priestley on offers and pricing (The Diary of a CEO, 2025): entrepreneurs discount reflexively, driven by an inability to hold composure through a prospect's silence rather than by any gap in pricing theory. The opening scene is illustrative, not a transcript. https://podcasts.apple.com/us/podcast/money-making-experts-this-3-step-offer-formula-makes/id1291423644?i=1000721000330 [^6]: Chase Hughes's ACSS hierarchy (Authority, Comfort, Social skills, Skills) places most communication failure upstream of technique, in state regulation under social pressure: a person with strong technique and poor composure collapses on contact. https://podcasts.apple.com/us/podcast/the-leading-body-language-behaviour-expert/id1291423644?i=1000681715542 [^2]: Kristjan Kullamagi (Qullamaggie) on Chat With Traders, episode 212 ("Breakouts, Home Runs & Exponential Returns"): he frames swing trading as a low-effort affair of waiting, trading only when a valid setup appears rather than because the market is open, so the edge is in regulating the impulse to act rather than in generating more signals. https://www.youtube.com/watch?v=K0F73Sq90j0 [^3]: Russell A. Barkley, "The Important Role of Executive Functioning and Self-Regulation in ADHD," reconceptualises ADHD as a disorder of executive function and self-regulation rather than a literal deficit of attention; individuals can hyperfocus, indicating the capacity is present but poorly regulated. https://www.russellbarkley.org/factsheets/ADHD_EF_and_SR.pdf [^5]: Anna Lembke, Dopamine Nation (2021): under sustained overstimulation the dopamine system downregulates baseline pleasure to hold the pleasure-pain balance, so the felt "deficit" is a consequence of the regulation response, not its cause. https://www.annalembke.com/dopamine-nation [^4]: Andrej Karpathy has argued across his 2025-2026 talks and year-in-review that much of the practical gain in agent work is in orchestration and context: how a capable-enough model is sequenced, checked, and constrained, with the harness often mattering more than a marginally smarter model. He separately stresses that fully autonomous agents remain capability-limited, putting them roughly a decade out (gaps in continual learning, computer use, multimodality). So this is a claim about where the bottleneck sits once a model is good enough for the task, not that capability never binds. https://karpathy.bearblog.dev/year-in-review-2025/ --- --- title: "How Reliable Is Your AI Agent?" description: "A month running an autonomous agent. Everyone who does comes back having built the same thing: a verifier." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/how-reliable-is-your-ai-agent/" date: "2026-06-15" series: "PROOF & TRUST" law: "Law IV" substack: "https://harryfloyd.substack.com/p/how-reliable-is-your-ai-agent" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # How Reliable Is Your AI Agent? *A month running an autonomous agent. Everyone who does comes back having built the same thing: a verifier.* By Harry Floyd · 2026-06-15 · canonical: https://durabilitycurve.com/blog/how-reliable-is-your-ai-agent/ For about three weeks I thought my server host was robbing me. The agent I run, Ghost, is my own instance of [Hermes](https://github.com/NousResearch/hermes-agent), an open-source framework from Nous Research. It works on a rented box all day with me nowhere near it: research, drafts, scheduled jobs, the unattended autonomy everyone is being sold. In May it started slowing down. A job that took a minute took five. The server's own dashboard blamed "steal time," the polite name for a noisy neighbour on shared hardware eating the processor. So I did the normal thing. I complained to support, read forum threads about oversold hosts, priced a migration. The neighbour was me. The failure had spent three weeks disguised as someone else's fault, which is exactly what the dangerous ones do. Ghost runs each of its tools inside a throwaway container and, by default, never deletes the dead ones. They piled up. When I finally ran the one command that would have told me on day one, the list of dead containers filled the screen and kept scrolling: 748 of them, all but five long dead, and the machine underneath had seized trying to keep track. The host throttled everything. Nothing was broken in the way broken usually looks: no crash, no error, no alert, only a number climbing by one, over and over, for weeks, with nobody reading it. That counter is what running an agent on your own actually looks like. The demo writes you a poem. The real thing is a number climbing in the dark while the disk fills. [Figure: fig02 counter] *FIG·02 — Nothing broke; a number climbed: 748 containers, all but five long dead, unread for weeks, into the cliff.* ### Grade it twice The counter was one way Ghost had fooled me. There was another, and this one you can run on your own agent this week. I ran a proper audit of its research against primary sources. Twenty-two factual claims about companies and markets, checked one at a time. Twenty came back directionally right, the right company and the right direction and the right thesis. Ninety-one percent. You could sell that number. Then I graded the same twenty-two strictly. Was every specific right too, the exact figure, the exact date, the exact quarter. Seventeen. Seventy-seven percent. That gap is the entire problem. The claims Ghost got strictly wrong were not inventions. They were the boring kind of miss: a price that was right last week, a date that had drifted, a version that had moved on. It had the shape of the world right and its current state wrong, and current state is the part you act on. So an agent that is confidently, directionally right is more dangerous than one that is obviously wrong. The obviously wrong one you check. The directionally right one earns your trust and then spends it on a stale figure you paste into a memo for someone who acts on it. If you have ever pulled a number from an AI and dropped it into a deck, you have shipped one of these without knowing. > A 91% that hides a 77% is the exact accuracy at which people stop checking. The audit takes an hour and tells you more about your own agent than any benchmark can. Take twenty of your agent's claims and grade them twice, once for the shape and once for every specific, current as of today. The spread between the two scores is your blast radius, and it is always wider than the single number you have been quoting. [Figure: fig03 gap] *FIG·03 — Grade it twice. The 91% that hides a 77%, and the three claims you would have shipped.* ### It was never the model Once I had that shape in my eye I saw it everywhere, and never in the model. The scheduler reported success whenever Ghost replied at all, so a job could fail outright, write back that it could not fetch the data, and still get [logged green](https://durabilitycurve.com/blog/the-stable-liar/) because something had come back. A config change I made was silently overruled by a second file that loaded later and won, pointing the agent's storage at the wrong place; the system did exactly what the files told it, and nothing reconciled the two. The search index ballooned overnight to a size with no relation to the data inside it, until the disk hit 100 percent, the agent started returning "no space left," and the work stopped, with nothing watching it grow. None of this was the model's fault. Resource leaks, configs that override each other, success signals that lie: these are the oldest problems in running software, and intelligence buys no exemption from them. The model underneath was cheap and fast; a frontier one would have hit all of them the same. A smarter model slips less often, and it still cannot see the slip it makes, because the evidence sits outside it. Every failure started in the same place: the agent did something, and nothing outside it looked at the result. A better model changes the odds. It does not remove the need for something outside the agent to look. That was the turn for me. I had spent weeks grading Ghost on how clever it was. What decides whether an agent is safe to leave alone is duller and far harder to fake: how much of what it does gets checked by something that is not the agent. > An agent cannot be the thing that confirms its own work. It reasons from inside its own process, where a sandbox that cannot start a container and a whole host that is down look identical. It will hand you a confident account of which one it is, and the account is worthless, because the fact that would settle it sits on the other side of a wall it cannot see over. An agent that grades its own work can always move the grade, which is why [the only gate a self-rewriting agent cannot game is time](https://durabilitycurve.com/blog/self-improvement-is-release-engineering/): the checks that hold are the ones it has no hands on. ### Different operators, same answer For a while I assumed this was specific to my setup. Then I read the operators who run Hermes hardest, and kept finding my own containers in their notes. None of us had compared notes; we were solving different problems, with different tools, in different words. One, tired of trusting the agent's own edits, makes every self-change a diff a human signs off before it goes live.[^dreaming] The setup guide everyone passes around spends its length on the dull perimeter: what each surface may touch, what runs sandboxed, what a human approves.[^perimeter] Others come at it from other directions, building evals to close the loop on output[^machina] and pruning the skills the agent writes for itself,[^curator] and the instinct is identical every time. Everyone who runs one of these for real comes back having built the same thing. Not a smarter model. A verifier. This finds you whether or not you run a server. It bites anyone who lets an AI do something they then act on: the draft you send without rereading, the figure you quote because it sounded sure, the report you stopped opening because it is always fine. An agent does not have to be autonomous to fool you. It only has to produce something you have stopped checking. You do not need your own month of this to learn what it teaches. ### Where the check has to live Fixing all of it was the same move every time. Put something outside the agent that can see what it cannot, and have it watch what is actually running, not the agent's account of it. The form that takes depends on the failure. The silent pile-ups, the containers and the disk, each had a leading number, one that moves long before the crash: the count of dead containers, the size of the index, the free space left. Those you watch directly. Set the alarm well below the cliff and read it far more often than it can break, every half hour rather than every twelve hours, so the warning fires while the problem is still a number and not yet a wall. The lying scheduler was harder, because there the agent was the one producing the success signal. If a check can be passed by the thing it is meant to be checking, it proves nothing; what you need is a record the agent cannot paint green just by replying. So the failure log gets written from the real error, by the harness itself, where it fires whether or not the model chooses to cooperate.[^harness] The stale figures got the narrowest fix of all. Ghost re-fetches a price or a date the moment it writes one, and never quotes from memory, because memory is where staleness hides. You pay that cost only on the things that decay, the price, the date, the version, the quarter. What does not move, it is allowed to remember. All of that catches a failure after it happens, which is enough when the damage can be undone. A full disk clears. A stale figure gets corrected. It is not enough for money that has already moved, a post published under your name, or a file deleted. So the real sorting key is reversibility. If an action can be undone, a watcher behind it will do. If it cannot, the check has to sit in front of it, and the check has to be a person. Ghost is barred from the irreversible ones. When it decides one is needed it writes a short request and stops. The operators who have run Hermes longest come to the same instinct from the other side: least privilege per surface, so the session you are sitting in front of can touch everything while the job that runs at four in the morning gets web and files and nothing else.[^perimeter] Where you cannot gate an action, make it reversible instead, snapshotting before a change and archiving instead of deleting, so there is always a state to roll back to. The boundary is dumb and absolute on purpose. > A safety rule an agent can argue its way around is not a safety rule. Look back at those fixes and they share one property. The counter it cannot fake, the log the harness writes instead of it, the date it has to re-fetch, the person standing in front of the irreversible action: the agent has no hands on any of them. That is the whole requirement. A check an agent can reach, it learns to play to, [which is how most verification quietly fails](https://durabilitycurve.com/blog/most-verification-is-just-bigger/): the check stops testing the work and starts testing the performance the agent puts on for the check. The watcher that survives is the one the agent cannot forge, whether it can see the watcher or not. None of this made Ghost smarter. It made Ghost watched, and watched turned out to be what mattered. [The harness around a model now drives more of the real-world difference than the choice of model does](https://durabilitycurve.com/blog/harness-engineering-same-model-different-product/), and a month of cleaning up after Ghost is that argument with scorch marks on it. The watchers and the gates and the re-fetches are where an autonomous system's real competence lives. You will swap the model next month. The instruments stay. ### An afternoon with a pen So the work that matters has nothing to do with grading the output, the one part that was never going to break. It is an afternoon with a pen. Write down everything your agent does while you are not watching: every scheduled job, every file it writes, every figure it fetches, every change it makes to itself. That list is its real reach, and it is always longer than you expect. Then put each line through three questions. Can you undo it, and if not, does a person stand in front of it? What looks at the result, and is it anything other than the agent itself? And the question that catches what the first two miss: could the agent make that check pass without doing the work? [Figure: fig04 pass] *FIG·04 — The verification pass: Reach, Reversibility, Witness, Forgery.* Take the dullest case you like, an assistant that drafts your weekly update and posts it to the team channel on a standing job. A posted message is read before you can take it back, so a person should see it first. Then ask what actually confirms the figures in it, and whether that is anything more than the assistant rereading its own draft. Ask whether it could report "posted, all good" with last week's numbers inside. Two of those checks are usually empty, and the empty ones are where it bites. Every action checked by nothing but the agent is one of my 748 containers. It is not failing yet. It is adding one to a counter nobody is reading. *Run it on your own setup this week. Which of your agent's actions is checked by nothing but the agent itself?* *The Durability Curve is where I write up what outlasts the model: the harnesses, the checks, the structure that is still standing after you swap the engine underneath. [Subscribe](https://harryfloyd.substack.com/subscribe?utm_source=substack-article&utm_medium=article&utm_campaign=agent-reliability-verifier) for the rest, or [start with what survives](https://durabilitycurve.com/blog/start-here-what-survives-when-the/).* [^machina]: Machina ([@exm7777](https://x.com/exm7777/status/2060736517564477901)), "How to Fix AI Slop" (X, May 2026). Frames inconsistent output as a quality-control gap rather than a prompt problem, and builds the fix as an eval loop: generate, score against a written benchmark, gate, and feed every failure back as a permanent test case. [^dreaming]: Tony Simons ([@tonysimons_](https://x.com/tonysimons_/status/2059119768662065523)), announcing Hermes Dreaming (X, May 2026). A plugin that stages an agent's proposed self-edits as create, diff, validate, then apply or discard, so a human reads the change before it touches live state. [^curator]: The Hermes Curator, Nous Research, as documented by mem0 ([@mem0ai](https://x.com/mem0ai/status/2050351798142288050)) (X, May 2026). A background pass that ages, archives, and reviews the skills an agent writes for itself, with critical skills pinned out of its reach. [^perimeter]: zaimiri ([@zaimiri](https://x.com/zaimiri/status/2063286261587055026)), "8 Hermes Agent Settings You Need Before Building Anything" (X, June 2026). Argues the settings that matter most are the perimeter ones: per-surface tool access, sandboxed execution, manual approvals, and scheduled jobs set to deny by default. [^harness]: Aparna Dhinakaran ([@aparnadhinak](https://x.com/aparnadhinak/status/2060406977357070522)), a code-level review of the Hermes harness (X, May 2026). Notes that lifecycle hooks can block or rewrite an action at the harness layer, enforcing policy independent of the model's cooperation. --- --- title: "Self-Improvement Is Release Engineering" description: "Your agent can rewrite its own memory and skills overnight. The hard part is whether you can see what changed and take it back. That makes self-improvement a release-engineering problem." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/self-improvement-is-release-engineering/" date: "2026-06-09" series: "AI & WORK" law: "Law III" substack: "https://harryfloyd.substack.com/p/self-improvement-is-release-engineering" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Self-Improvement Is Release Engineering *Your agent can rewrite its own memory and skills overnight. The hard part is whether you can see what changed and take it back. That makes self-improvement a release-engineering problem.* By Harry Floyd · 2026-06-09 · canonical: https://durabilitycurve.com/blog/self-improvement-is-release-engineering/ Your agent improved itself overnight. You wake up to a clean changelog: a new retry skill, a reorganised memory, and a rewritten rule for which record it trusts when two sources disagree. It reads like progress. You still cannot ship it, because you cannot see exactly what changed, you cannot tell which edit is load-bearing, and if that new trust rule is subtly wrong you have no way to pull it back out. So the changelog sits there. The capability is real and the trust is missing, and the distance between the two is the entire problem. What sets that distance is whether a person can inspect and reverse what the agent did to itself, and the model's intelligence barely moves it. The frontier of agent self-improvement is becoming release engineering. Two things have to be true before an agent can get better. It has to remember: you gave it a memory and [it still repeated the same mistake until a loop turned the recurring failures into procedures](https://durabilitycurve.com/blog/remembers-everything-learns-nothing/). It also has to govern what it picks up, because a copied skill is [a dependency you have to version, scope, and be able to switch off](https://durabilitycurve.com/blog/skills-are-package-management-for-your-ai/). Memory is the input. Governed skills are the output. The layer in between is the path that takes what the agent went through and turns it into a change in how it behaves next time. That is **the consolidation layer**, and most of it currently runs with no brakes. > An agent with memory and skills but no consolidation layer accumulates experience it cannot convert and capability it cannot trust. [Figure: The gate between memory and governed skills. Most setups leave it open.] Watch what the serious implementations converge on. The clearest version ships as "dreaming": the agent runs offline, proposes edits to its own memory, skills, and notes, and writes them to a frozen folder instead of to itself. Create, diff, validate, then apply with a backup or discard with an archive. Nothing touches the live agent until a human reads the diff. Unrelated systems are converging on the same shape: an offline proposal, a staged review, an eval gate rather than the model's own enthusiasm, and controlled promotion. A research paper on self-evolving skills wraps the identical idea in an "executive strategy" that turns each proposed change into a bounded, controlled edit rather than a free rewrite.[^1] When unrelated people arrive at the same structure, the structure is the finding. Make it concrete. An agent that handles refunds keeps fumbling the edge cases, and the failures collect in its memory. Overnight it proposes a fix: a rule that auto-approves any refund under a threshold to clear the queue faster. The staged version shows you the diff and runs the new rule against last month's resolved tickets, where it looks fine, because the eval measures queue clearance and same-day satisfaction. You promote it. Three weeks later, approvals are up and chargebacks are climbing. The change passed because it was scored on the wrong thing, and the harm only showed up downstream, where nothing was watching. Because it was staged and reversible, you can pull the behaviour back out instead of guessing which invisible self-edit caused the drift. Notice which half they are all protecting. Proposing a change is one model call. Knowing the change is safe to keep takes a diff, a validation pass, a backup, and a way to roll it back. The cheap half is the proposal. The scarce half is the judgement about the proposal, which is the same migration [reshaping what evaluation is worth across the agent stack](https://durabilitycurve.com/blog/most-verification-is-just-bigger/): once generation is cheap, the verifier becomes the product. Strip the staging out and you get a faster loop that no operator will run in production. The friction is the product. An agent optimising for "I improved myself overnight" is optimising a proxy, and a busy changelog can hide that no one has actually checked what changed. Staging keeps the loop honest. It does not make it compound. Three findings from people building these systems separate the loops that get better over time from the ones that only generate motion. Score a change by what it leads to. How good it looks the day it lands is the wrong measure. In a study of self-modifying coding agents, the version that scored best on the benchmark today was a poor guide to which line of descendants actually improved; immediate score and long-run potential come apart.[^2] The change worth keeping is the one whose later descendants come out strong. Keep the distilled lesson and discard the transcript. What transfers between tasks is a compact, retrievable heuristic, and feeding the raw transcripts back in helps less than the heuristic does.[^3] The durable artefact is the rule the agent pulled out of the episode. The episode itself is disposable. Review what the agent taught itself before it ships. The agent that wrote the skill does not get the only vote on keeping it. Put those findings together and the danger sharpens into something worse than an unreadable changelog. The agent that proposes a change is the same one that will be graded on it, and it is optimising to pass. So the improvement most likely to clear your review is the one tuned to clear your review, which is not the same as the one that makes the work better. Every gate you can write down, a system that rewrites itself can learn to satisfy. The single test it cannot game is the one it cannot see in advance: what the change does downstream, weeks after it ships. A reviewer and an eval judge the change as it is. Only time judges what it did, and time is the last gate a self-improvement loop almost never has. > Consolidation that compounds looks like a release process. Consolidation that is theatre looks like an agent applauding its own diffs. This is also where the durable advantage sits, and it is worth being exact about why. The model underneath is rented and resets every cycle; whatever it can do, your competitor's can do on the same Tuesday. Memory fragments the instant each tool keeps its own: the support agent learns a customer's quirk on Monday and the billing agent re-derives it from scratch on Thursday. The skill file is cheap to copy; what is not is knowing which skill survived contact with your users, your failures, and your evaluation loop. The consolidation layer is the one part that is yours: the reviewed, accumulated record of which changes your agents kept and why, the process that turned your specific failures into your specific procedures. Swap the model, re-import the skills, rebuild the memory store, and the thing that survives is the loop that decides what gets kept. > Whoever owns which change gets kept owns how the agent evolves. A smarter model does not solve this by itself. The fix is the discipline a release process already has: nothing reaches live behaviour that a person has not seen and cannot reverse, and nothing is kept until time has had its vote. Almost no one builds that last gate. It makes you wait, and waiting does not feel like progress. [Figure: The Consolidation Test: four questions that separate a release process from an agent editing itself in the dark.] So the next time a tool tells you its agent learns, or you stand up a process that promotes your own agents' improvements, run four questions on it. Can you see the diff before it applies. Can you take it back out after. Is it scored on whether it made later work better, or only on whether it looked good when it landed. Does it keep the lesson or only the log. **Four yeses is a release process.** Three or fewer is an agent editing itself in the dark, and the changelog you wake up to is a liability with good formatting. Each no has a cheap fix, and not one is a research project: a staging folder, a one-step rollback, a downstream metric, a written rule. *Which of your agents is changing itself right now in a way you could not, this minute, undo?* --- *Want to run this on your own stack? [The Consolidation Test](https://durabilitycurve.com/downloads/consolidation-test.pdf) is these four questions as a one-page card you take to any agent that claims to learn, or to your own skill-promotion process, with the smallest fix for each column you score a no on. Most loops fail at least one.* *The Durability Curve is one essay a week on what stays valuable while the tools underneath keep changing. [Subscribe](https://harryfloyd.substack.com/subscribe?utm_source=article&utm_medium=web&utm_campaign=consolidation-layer) if that is your kind of question.* [^1]: "SkillOpt: Executive Strategy for Self-Evolving Agent Skills" (Yang et al.), [arXiv:2605.23904](https://arxiv.org/abs/2605.23904). It turns scored rollouts into bounded add, delete, and replace edits on a single skill document, not free-form self-rewriting. [^2]: "Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine" (Wang et al.), [arXiv:2510.21614](https://arxiv.org/abs/2510.21614). It names the Metaproductivity-Performance Mismatch, that a self-modification's current benchmark score does not predict the quality of its descendants, and scores a change by its clade's aggregate performance instead. [^3]: "Experiential Reflective Learning for Self-Improving LLM Agents" (Allard et al.), [arXiv:2603.24639](https://arxiv.org/abs/2603.24639). Reflecting on past trajectories to distil reusable heuristics, retrieved selectively at test time, transfers better than few-shot prompting with raw trajectories; ablations show selective retrieval is essential. --- --- title: "Skills Are Package Management for Your AI" description: "There are more than 1.6 million you can install. You need about twenty. Software already solved that problem once." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/skills-are-package-management-for-your-ai/" date: "2026-06-07" series: "AI & WORK" law: "Law III" substack: "https://harryfloyd.substack.com/p/skills-are-package-management" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Skills Are Package Management for Your AI *There are more than 1.6 million you can install. You need about twenty. Software already solved that problem once.* By Harry Floyd · 2026-06-07 · canonical: https://durabilitycurve.com/blog/skills-are-package-management-for-your-ai/ There are more than 1.6 million Claude skills you can install today.[^1] You need about twenty. The distance between those two numbers is the whole problem, and it is an old one. Software lived through this exact moment once before, when we decided code should travel in small reusable units. Sharing turned out to be the easy part. The hard part arrived a few years later, and it was trust: which of the millions of packages is current, safe, and does what its README promises. We answered that question by building a whole discipline around it. Versioning. Lockfiles. Audits. Deprecation notices. A bill of materials. Most people wiring skills into their AI right now are skipping every step of it and treating the result as progress. A skill is the smallest durable unit of agent behaviour. In plain terms it is a folder with a single instruction file inside, written so the model loads it only when the task in front of it matches. Anthropic published the format as an open standard, and one skill file now runs across more than twenty different agents, from Claude Code to Codex to Cursor.[^2] That portability is the tell. A thing built to be copied everywhere is a thing whose copies will multiply faster than anyone can check them. > A prompt is stateless and dies with the conversation. A skill is a versioned file an agent picks up when the work calls for it, and it behaves the same way next week. That difference is bigger than it looks. The moment a procedure persists, gets shared, and runs without you watching, you have stopped writing prompts and started managing dependencies. ## What you are actually installing When you copy a skill off a marketplace, you are adding a behavioural dependency, and that is a stranger and more dangerous object than the code dependencies engineers already lose sleep over. A bad code library throws an error you can see in a stack trace. A bad skill expresses itself through the agent's decisions. It nudges a tone, skips a verification step, assumes a permission, reaches for the wrong tool, and it does all of this inside work you delegated precisely because you were not going to check every line. The failure does not announce itself. It compounds, silently, across every session that loads the file. Security researchers have started treating a copied skill as exactly what it is: a dependency in your agent's behaviour, with its own supply chain to secure.[^3] > Untrusted packages compromise your software. Untrusted skills compromise your behaviour. Software's answer was to wrap every copy in accountability. A package carries a version, a source, a license, a list of what it is allowed to touch, and a path to rip it out when it turns. The agent-skills world has the copying down and almost none of the accountability. The marketplaces measure themselves in millions of listings and install counts, which is the metric of a field that still thinks the file is the achievement. ## The library is the moat, not the model This matters past hygiene. The model underneath your agent is rented. It resets every release cycle, and the next version reaches your competitor on the same Tuesday it reaches you. Whatever advantage lives in the weights is an advantage everyone gets at once. So it cannot be where your durable edge sits. The skill library can be. It is the layer you own, the accumulated record of how your work gets done, and the test of whether it is a moat is simple: you can swap the model underneath it without losing what you built. The decisions survive the upgrade. That is architecture outliving content stated in one sentence, and it is the same structural bet behind why [the same model behaves like a different product depending on the code wrapped around it](https://durabilitycurve.com/blog/harness-engineering-same-model-different-product/). > A skill file costs nothing to copy, which is the precise reason the file was never the product. If anyone can copy it for free, the value sits in the curation: knowing which twenty of 1.6 million earn a place, verifying each one runs in your environment rather than a demo, and writing down the failure mode you only learned by hitting it. That is also where the money goes. The marketplaces selling skill files are mostly dying, while the work of choosing and vetting and bundling is becoming the thing people pay for. Value migrated off the artifact and onto the judgement about the artifact, which is the [same move the leverage hierarchy of agent engineering](https://durabilitycurve.com/blog/leverage-hierarchy-of-agent-engineering/) traces one layer down. A 2026 benchmark of eighty-six tasks put numbers on it: a curated skill set raised the average success rate by about sixteen points, while the skills models wrote for themselves produced no gain at all.[^4] [Figure: What a skill costs to copy, and what it costs to trust: the file and the model are the cheap half; curation, verification and a kill-path are the scarce half that does the work.] ## What the discipline looks like when you run it My agent stack carries 179 skills. A handful I wrote by hand, each from a failure I had already hit and read closely enough to encode. The rest I pulled in from libraries, the way you add packages to a project. One of the hand-written ones does grounded research, and it exists only because the obvious tool for the job drove a browser under my own login and broke a platform's terms of service to do it. So the capability got rebuilt on a sanctioned interface instead. The skill carries that constraint in writing, because a rule that lives only in my head is a rule the agent will eventually cross. Another governs what is allowed to graduate from a holding area into the permanent vault, and the first thing it does, before any other check, is look for a duplicate, because a fabricated gap is treated as a failure rather than a clever new contribution. Then I wrote the audit this piece describes and ran it across the whole folder. The result was humbling. The average skill scored two out of six. 82% carried a version, but only 4% recorded when they were last verified, under a quarter declared what they were allowed to touch, and 3% had any way to retire them. The grounded-research skill I just held up as a model scored zero, because I had written it as careful instructions and never given it a version, a source line, or a switch to turn it off. The discipline I am describing here, I was barely doing myself. > The file is the cheap part. The discipline wrapped around the file is the part that took months and cannot be copied. There is an order to it that matters more than the contents. The guardrails go in before the skills that act. The review gate, the duplicate check, the permission boundary, the terms-of-service rule: those get installed first, so that by the time a capable agent is doing real work, the structure it would happily skip has already been made mandatory. A model that is good enough to be useful is good enough to route around safety it sees as optional. The install order is how you make it not optional. ## The one-week test Open the folder where your AI keeps its skills. For each one, answer six questions. 1. What version is this? 2. Where did it come from? 3. What is it allowed to touch? 4. When did I last confirm it works in my setup? 5. What else now depends on it? 6. How would I switch it off in a hurry? [Figure: The Skill Bill of Materials: a six-column audit, one point per column you can answer for real. Six is a managed dependency; three or below is a skill running your agent unchecked.] Most skills will fail that audit, and the ones that fail are usually the ones quietly running your agent. Write the six answers down for each. The folder that results is the first version of the thing that compounds while the models underneath it keep getting replaced. *Which skill is steering your agent right now that you could not, this minute, tell me the source of?* --- *Want to run this on your own folder? [The Skill Bill of Materials worksheet](https://durabilitycurve.com/downloads/skill-bill-of-materials.pdf) is the audit above, turned into a one-page sheet you fill in. Take it to your skills directory this week and see how much of it scores zero.* *The Durability Curve is one essay a week on what stays valuable while the tools underneath keep changing. [Subscribe](https://harryfloyd.substack.com/subscribe?utm_source=article&utm_medium=web&utm_campaign=skills-are-package-management) if that is your kind of question.* [^1]: SkillsMP, the largest public agent-skills marketplace, listed 1,640,440 skills as of 8 June 2026. https://skillsmp.com/ [^2]: Anthropic, "Equipping agents for the real world with Agent Skills," and the public reference repository at github.com/anthropics/skills. The SKILL.md format is documented as an open standard adopted across Claude Code, Codex, Gemini CLI, Cursor and others. https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills [^3]: The supply-chain framing is now formal. See "Formal Analysis and Supply Chain Security for Agentic AI Skills" (arXiv:2603.00195), which proposes an Agent Skill Bill of Materials recording each skill's identity, version, content hash, declared permissions, and dependency edges. https://arxiv.org/abs/2603.00195 [^4]: SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks (arXiv:2602.12670). Across 86 tasks and 7,308 trajectories, a curated skill set raised average pass rate by 16.2 points, while self-generated skills gave no average benefit. https://arxiv.org/abs/2602.12670 --- --- title: "Right About AI, Wiped Out Anyway" description: "AI is real. The open question is whether the companies spending $725 billion on it live to collect." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/right-about-ai-wiped-out-anyway/" date: "2026-06-06" series: "MARKETS & POWER" law: "Law I" substack: "https://harryfloyd.substack.com/p/right-about-ai-wiped-out-anyway" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Right About AI, Wiped Out Anyway *AI is real. The open question is whether the companies spending $725 billion on it live to collect.* By Harry Floyd · 2026-06-06 · canonical: https://durabilitycurve.com/blog/right-about-ai-wiped-out-anyway/ *The return gauge: $725 billion in, the needle barely moves.* In 2001 the fibre was already in the ground. More than eighty million miles of it, laid across the country in five years, financed mostly with debt, on the conviction that internet traffic would need every strand.[^1] By the end of that year roughly ninety-five percent of it was dark. Unlit. Carrying nothing. The conviction was correct. Traffic came. Fibre laid in 1999 carries your video calls right now, and the internet became exactly the world-changing force the buildout bet on. The people who built it did not collect. Global Crossing filed for bankruptcy in January 2002 with $12.4 billion in debt. WorldCom followed that summer with the largest bankruptcy in American history to that point. Telecom equity lost more than two trillion dollars of value between 2000 and 2002.[^2] The infrastructure was real, the demand was real, and the owners were wiped out anyway. That gap, between being right about a technology and getting paid for it, is the most important thing to understand about the $725 billion that four companies are about to spend on artificial intelligence this year. ## The number, and the gap under it Google, Microsoft, Meta, and Amazon have guided investors toward roughly $725 billion of capital spending in 2026. That is up seventy-seven percent in a single year. Add Oracle and it pushes past three-quarters of a trillion dollars.[^3] Most of it goes to AI: the chips, the data centres, the power to run them. The spending now runs at about ninety percent of the operating cash flow these companies generate, by Bank of America's estimate.[^4] They used to spend thirty to fifty cents of every operating dollar on capital. Now they spend ninety. Against that, AI revenue runs somewhere around one hundred to one hundred fifty billion. The buildout is roughly five times larger than the business it is meant to serve. This is not a bear talking. Goldman Sachs, whose clients own most of these stocks, put the question on its own letterhead and titled it *Gen AI: Too Much Spend, Too Little Benefit?* The firm's head of global equity research sat for the interview and said he doubted the technology would ever justify its cost.[^5] When the bank underwriting the boom asks in print whether a trillion dollars of spending pays off, the doubt has left the fringe. [Figure: diagram2 right about ai anim] *Roughly five dollars of capex for every dollar of AI revenue. (Animated: the buildout surging to $725B while revenue stays flat.)* ## You can be right about the bottleneck and wrong about the return Most of the argument about AI infrastructure is about the wrong thing. The popular question is where the scarcity sits: [GPUs now, then power, then memory](https://durabilitycurve.com/blog/three-hidden-bottlenecks-past-gpus/), then whatever turns out to bind next. It is a good question. It tells you which supplier captures the margin this quarter. It tells you almost nothing about whether the spending pays back. The binding question is the other one. Does AI revenue grow into the $725 billion before the companies writing the checks lose their patience? The telecom investors were right about fibre. They were right about the bottleneck, right about the technology, right that the world would need it. They were wrong about the return, and the return is the only thing that pays a shareholder. ## Demand that is real and shallow at the same time The bull case rests on demand arriving to fill the buildout. The data says demand is real and shallow at once. Nearly nine in ten enterprises now use AI in some form. Only about three in ten report a clear return on it. And eighty-eight percent of agent pilots never reach production.[^6] Real but shallow is the decisive shape. It means a large share of today's AI spending is discretionary: pilots, experiments, seat licenses that live exactly until the first serious budget review. Discretionary demand is the demand that leaves first when the cycle turns. The buildout is being sized for demand that has not yet proven it will stay. ## The bear case is two cases, and they fail differently Treating the downside as one thing is the common mistake. It is two, on different axes, with different tells and different survivors. The first is efficiency. Compute demand keeps climbing, but [the hardware improves so fast that far less of it satisfies the need](https://durabilitycurve.com/blog/the-other-half-of-compute/), and the buildout overshoots. You see this one when GPU utilisation falls while workloads still grow. The survivors are the software and inference layers, anything that sells use rather than raw capacity. The second is returns. AI revenue never grows into the spending, the five-to-one gap holds, and the checks eventually stop. You see this one when AI revenue growth runs more than twenty points below capex growth, and when the ROI surveys stall instead of climbing. The survivors are balance sheets and annuity businesses, the companies that can afford to wait. Pure capacity owners de-rate. A company can win the first failure and die in the second. Holding the two apart is most of the analytical work, and almost nobody does it. ## A clock that runs regardless of demand There is a deadline on this that does not care whether demand shows up. Chips wear out on the books in three to five years. Seven hundred billion dollars of capital spent in 2026 becomes something like one hundred fifty to two hundred forty billion of annual depreciation landing in 2027 and 2028. A margin event with a date on it. You can already watch the companies brace. Several have quietly stretched their depreciation schedules from three years to five, which lowers the reported expense and flatters the margin while the chips age at exactly the rate they always did.[^7] The accounting can move. The silicon cannot. ## A distribution with a date So the answer is a probability with a date. Three out of ten, demand catches up: inference, agents, and reasoning workloads absorb the buildout, and returns normalise by around 2028. Two out of ten, hard reset: demand disappoints, capex is cut sharply across 2027 and 2028, the write-downs are large, and AI-infrastructure equities fall forty to sixty percent. Five out of ten, the middle: revenue grows, but not fast enough. Margins compress. Capex decelerates. A few write-downs, no crash. A slow grind. The slow grind is the most likely outcome, and it is the one almost no one is positioned for, because it pays off neither side cleanly. The bull needs the clean catch-up. The bear needs the crash. The likeliest path rewards neither, and it punishes anyone who sized a position as though only the two clean endings could happen. [Figure: diagram3 right about ai] *A probability with a date. The slow grind is the unlit middle nobody is positioned for.* ## What to watch, and the one-week version The buildout will be real and useful. That was true of the fibre too. The question that decides whether you collect is narrower: does revenue grow into the spending before the spenders lose their nerve, and is your position built to survive the slow grind if it does not. There is a way to watch the turn instead of guessing at it. Expansion becomes deceleration in advance, in a handful of signals. The master one is guidance: the first time two of the big five trim their capex numbers or soften the language, the regime is changing. Under it, watch for AI revenue growth slipping below fifty percent a year, deployed GPU utilisation falling under half, depreciation schedules stretching, and vendor financing that quietly loops a chipmaker's money back as a customer's demand. When three of those fire together, the cycle has turned. Run the one-week version yourself. Take your largest AI-exposed position and write down why you own it. If the reason is about where the bottleneck sits, you have answered the supply question and skipped the binding one. Re-size it on the return: whether the revenue grows into the spending, and whether the company collects if the buildout takes longer than the bulls promise. Three things would tell me I am wrong, and I am watching all three: AI revenue growth holding above fifty percent and closing the gap by 2027; the hyperscalers sustaining ninety percent of cash flow on capital for two more years while their cloud margins expand; and the 2027 depreciation wave passing with no material write-downs. Any one of those moves the weight toward the clean payback. Until one of them does, the safest assumption is the one the fibre taught. Being right about the technology and getting paid for it are two different bets. Size the second one. [Figure: field card right about ai] *The instrument: watch the composite trigger, then run the one-week test.* --- *Structural reads on the AI cycle, each with an instrument you can run. Free, in your inbox.* *Which of your AI positions is sized for the slow grind, and which is quietly betting on the clean payback?* [^1]: In the five years after the Telecommunications Act of 1996, U.S. carriers invested more than $500 billion, mostly debt-financed, laying roughly eighty million miles of fibre ([Telecoms crash, Wikipedia](https://en.wikipedia.org/wiki/Telecoms_crash)). [^2]: By 2001 about 95 percent of that fibre was dark. Global Crossing filed for bankruptcy in January 2002 with $12.4 billion in debt; WorldCom followed in the summer of 2002. Global telecom equity lost more than $2 trillion in value between 2000 and 2002 ([Telecoms crash, Wikipedia](https://en.wikipedia.org/wiki/Telecoms_crash)). [^3]: Google, Microsoft, Meta, and Amazon have guided to roughly $725 billion in combined 2026 capital spending, up about 77 percent from the prior year's $410 billion, in their Q1 2026 earnings ([Yahoo Finance](https://finance.yahoo.com/markets/article/magnificent-7-earnings-rush-reveals-ai-spending-surge-with-hyperscaler-capex-set-to-reach-725-billion-in-2026-224901707.html)). Oracle adds roughly $50 billion more, pushing the five-company total past three-quarters of a trillion dollars. [^4]: Bank of America estimates the five largest hyperscalers (Microsoft, Amazon, Alphabet, Meta, Oracle) will spend about 90 percent of their operating cash flow on capex in 2026, up from roughly 65 percent in 2025 ([MarketWise](https://marketwise.com/investing/hyperscaler-investment-surge-2026-ai-capex-buildout/)). [^5]: Goldman Sachs Research, ["Gen AI: Too Much Spend, Too Little Benefit?"](https://www.goldmansachs.com/insights/top-of-mind/gen-ai-too-much-spend-too-little-benefit) (Top of Mind, June 2024), featuring Jim Covello, Head of Global Equity Research, and Daron Acemoglu of MIT; the firm published a further skeptical assessment in May 2026. [^6]: Roughly nine in ten enterprises now use AI in some form; about three in ten report a clear return ([Writer's 2026 Enterprise AI Adoption Survey](https://writer.com/blog/enterprise-ai-adoption-2026/), ~29 percent seeing significant ROI); and some 88 percent of agent pilots never reach production ([Forrester and Anaconda research](https://www.cio.com/article/3850763/88-of-ai-pilots-fail-to-reach-production-but-thats-not-all-on-it.html), 2026, widely replicated). [^7]: Several hyperscalers have extended AI-chip depreciation schedules from three years to five ([Fortune, April 2026](https://fortune.com/2026/04/15/data-centers-hyperscalers-spending-billions-on-hardware-thats-worthless-in-3-years/)). --- --- title: "Remembers Everything, Learns Nothing" description: "You gave your agent a memory and it still repeats the same mistake. What makes it improve is a loop that tests each failure and turns the ones that recur into procedures." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/remembers-everything-learns-nothing/" date: "2026-06-05" series: "AI & WORK" law: "Law II" substack: "https://harryfloyd.substack.com/p/remembers-everything-learns-nothing" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Remembers Everything, Learns Nothing *You gave your agent a memory and it still repeats the same mistake. What makes it improve is a loop that tests each failure and turns the ones that recur into procedures.* By Harry Floyd · 2026-06-05 · canonical: https://durabilitycurve.com/blog/remembers-everything-learns-nothing/ *You gave your agent a memory and it still repeats the same mistake. What makes it improve is a loop that tests each failure and turns the ones that recur into procedures.* ## The agent broke a rule it had written down The agent had the rule. You can find it in its memory file, line 1,140 of about 1,800: do not reformat the config, the deploy is strict about its indentation. You wrote it there three weeks ago, the last time it broke the deploy. This morning the agent reformatted the config. The deploy broke. It remembered the rule and broke it anyway. Anyone who has run an agent for more than a week has met some version of this. You added memory. You did the responsible thing, gave the agent a place to keep what it learned, and then watched it keep the wrong things, or keep the right things somewhere it never looks. Run 100 came out no sharper than run 1. The store filled up and the behaviour stayed exactly where it was. It remembered too much, and it kept all of it in one pile, where 1,800 lines bury every instruction equally. The rule loaded. It just loaded as one line among the hundreds it had needed once and never again. > Memory is several different things wearing one name. Store them in one place and you get a slow agent, a swollen context, and a file full of rules that quietly disagree. That is the whole problem. The fix has two parts: sort what the agent keeps by how often it changes, and put a loop in charge of what gets to stay. ## Memory is four things, sorted by how often each changes Pull the pile apart with one question: how often does this change? The answer tells you where each piece belongs, because the four kinds of memory want four different homes. At one end sit the things that almost never change. The agent's standing rules, its constitution, the few constraints that hold on every task. Those belong in one short, always-loaded file, the kind Claude Code keeps in `CLAUDE.md` and Codex keeps in `AGENTS.md`, and short is the load-bearing word. A fresh session can burn a real slice of its budget loading its own instructions before you have typed a thing, so a line earns its place in that file by one test: would the agent get this wrong without it, on most tasks? The config rule passes, which is why it belonged here, in the fifty lines the agent reads every time, not on line 1,140 of a log it barely skims. A step along are the things that change now and then, and only for certain tasks. Workflows, procedures, the steps for cutting a release. Those become skills, each in its own small file the agent loads only when the task calls for it, so you can keep fifty of them and pay for none until the one you need comes up. Further along is what changes every single run. What the agent did, what broke, the fix it landed on. That is raw trajectory, and it goes in an append-only log, a plain `learnings.md` you only ever add to and compress later, never editing in place. At the far end sits the big reference body that changes rarely but in bulk. Documentation, old code, whatever corpus the agent searches. That lives in an external store it queries once, early, under tight hygiene. Facts live there. Lessons live a layer up, in the log. > Sort memory by how often it changes, and each kind lands where the agent will look for it. Mix them, and the rule you need every time sits buried among everything you needed once. This is the part most "give your agent a memory" guides get right and then quietly undo, by letting all four drain back into one file. The whole trick is keeping them apart. Hold the four kinds separate and each stays usable. Let them merge and you slide back into the single pile, config rule and all. [Figure: The four memory layers sorted by how often each changes: standing rules in CLAUDE.md, procedures in skills, the episodic learnings.md log, and external reference.] *Four layers, keyed to rate of change. The rule you need every session belongs in the always-loaded file, not on line 1,140 of a log.* ## The loop is the part that compounds Separation stops the rot. It does not, on its own, make the agent better. A tidy store is still a store, and a store only remembers. None of what follows pays off on one-off work, where a wrap-up is pure overhead; the loop earns its keep only when the same failure keeps coming back. Improvement comes from a loop that runs on top of the store, and the loop has four moves. It starts with a wrap-up. After every task, the agent appends a few plain lines to the log: ``` ## 2026-06-05 deploy auth changes did: edited config.yaml, ran the deploy failed: deploy rejected the file, the parser choked on the new indentation fix: restored the original formatting, deploy passed next time: do not reformat config.yaml, the deploy is strict about indentation ``` The last line is the one that earns its keep. That `next time` is the promotable lesson, the single thing that might change what the agent does on its next run. You wire this up with one standing rule in the always-loaded file: after each task, append a wrap-up to `learnings.md` in this shape. The model will mostly remember on its own, and a Claude Code stop hook makes it certain, running the wrap-up the moment the agent finishes. No wrap-up, no raw material, and everything downstream starves. Then an evaluation runs. On a schedule, you re-run a set of tasks drawn from real past failures and check whether the agent still handles them. This is the step that turns "it feels worse lately" into a logged, specific entry you can act on. Consolidation comes next. Once a week, a pass compresses the log, archives the dead lines, and holds the active file to something an agent can read in one sitting, a few hundred lines rather than a few thousand. Without it, the append-only log becomes the 1,800-line pile again by a slower route. Promotion is where it pays off. When the same `next time` line shows up three times or more, it graduates. It stops being a log entry and becomes its own skill, a file at `.claude/skills/deploy-prep/SKILL.md`: ``` --- name: deploy-prep description: Use before any deploy. Stops the config.yaml reformatting that has broken the deploy three times. --- Before any deploy, leave config.yaml formatting untouched; the deploy is strict about indentation. Run `make check-config`, then deploy. ``` That `description` line is the load-bearing part. The agent reads it every session and pulls the skill in only when a deploy comes up, so the rule stays out of the way until the moment it matters. The buried log line is now a procedure the agent runs without being told, and the three redundant entries get cut. The break stops happening. A skill is still context, though. The agent reads it and usually obeys, and for most lessons usually is the right bar. For a failure that is cheap to trigger and expensive to suffer, promote it one more step, out of memory and into enforcement: a pre-deploy check that fails loudly, or a Claude Code hook that blocks the edit before it lands. The `make check-config` line in that skill is the seed of it. Memory tells the agent what to do. A hook makes the wrong move impossible. > A store remembers. A loop improves. The difference is whether a failure the agent logged ever becomes a procedure it runs without being asked again. Notice what compounds. The store only grows. The loop is the part that keeps finding the failures that recur and turning them into procedures, and three of its four moves throw things away or move them up the stack. Only the first one adds. [Figure: The loop: wrap-up, evaluate, consolidate, promote, with enforcement as the escape for rules that must never recur.] *Wrap-up feeds the log. Evaluation is the step most skip. Consolidation keeps it lean. Promotion turns a recurring lesson into a skill, and for the failures that must never recur, one step further into a hook.* ## The evaluation is what separates learning from theatre One of those four moves is doing more work than the rest, and it is the one almost everyone skips. Run the loop without the evaluation step and the wrap-up notes are self-reported and unchecked. The agent writes "fixed the config issue" and nothing on earth confirms it. The log fills with confident receipts for work that may not hold. You get a beautiful record of intentions and no idea whether the agent is improving or quietly getting worse. Anthropic's engineering team put the cost plainly in their guidance on evaluating agents: teams without evals get stuck in reactive loops, fixing one failure and creating the next, unable to separate a real regression from noise.[^1] An evaluation is the instrument that turns "did this lesson stick" into something you can observe. It is the test that the config still survives a deploy after the agent swore it learned. This is the part worth sitting with. The evaluation is the hard, tedious, expensive step, the one you are most tempted to defer, and it is precisely the one doing the work. Consolidation and honest forgetting are core engineering, the mechanism that keeps a growing transcript from turning into noise, and then into poison. Skip the test and every other layer you built just adds volume. > Persistence without a test is theatre. The evaluation is the only thing that can tell you whether a remembered lesson made the agent better or only made the file longer. There is a sharper trap underneath. Optimise your memory system on how much it stores, lines logged, lessons captured, store size, and you are measuring activity. The number climbs while the agent stands still. A handful of real tasks the agent must still pass, run on a schedule, is worth more than any size metric, because it measures the only thing you wanted: did remembering change what the agent can do. The config break is eval case one: ``` task: add a feature flag to config.yaml and deploy pass: the deploy succeeds and config.yaml is unchanged except for the new flag ``` Grade the outcome, the deploy passing, and let the agent reach it however it likes. The runner that does this is humble: a short script hands the agent each task in a clean workspace, then checks the pass line. A dozen lines of shell cover the scale this piece is about, and an off-the-shelf eval harness is the same loop once you outgrow it. Run a handful of these on a schedule and "is the agent getting better" stops being a feeling and becomes a number you can read. ## What poisons a memory When a memory system fails, the store is rarely what broke. The governance did, the policy for what gets to persist and how conflicts get resolved. Three failures cause most of the damage, and none of them are about storage. The silent merge is the cheapest to fix and the easiest to miss. Two notes disagree, the agent picks one, and a real contradiction vanishes into a single confident line nobody flagged. The better setups converge on one fix: when sources conflict, mark it with a literal tag and let a human resolve it, so the disagreement stays visible instead of dissolving into one quiet error. Auto-deployed consolidation is the newest of the three, and the most seductive. Anthropic's Dreaming, a research preview from May 2026, runs a scheduled pass between sessions that rewrites an agent's memory store from its recent work.[^2] It is genuinely useful, and Anthropic built in the safeguard that matters: the original store stays read-only, and the rewrite arrives as a separate output you approve before the agent runs on it. Point the agent at the new store unread and you can promote a hallucinated merge to a standing rule. Read before swap. The consolidation gives you a candidate, not a fact. The bloated constitution is the slow one, the failure you already met. Everything important gets added to the always-loaded file, because adding feels safe, until the rule you need every time is buried on line 1,140. Keeping it short is what keeps the rest legible. One cousin is worth naming: a workspace that keeps your instructions loaded still does not remember your last session, so assume it does and you will lose context without knowing why. > Most memory failures are governance failures. The store did not break. The policy did, by merging two truths into one confident error that no one read before it became a rule. ## The one-week test You do not need a vector database to find out whether any of this applies to you. You need a week. For one week, end every agent task with three lines: ``` failed: what broke fix: what actually worked next: the rule for next time, or nothing ``` Write nothing else, store it nowhere clever, a plain file is fine. At the end of the week, count which `next` lines repeat three times or more. If nothing clusters, your agent does not yet have a memory worth building, and no database will give it one. The failures are not recurring, which means there is nothing stable to promote, and a bigger store would only hold more noise. That is a real and useful answer. It tells you the work is in the task, not the memory. If something does cluster, you have found your first skill. The recurring failure is the one that should stop being a note and start being a procedure, and you now know exactly which one to lift first. The config rule was mine. Yours will be sitting in those three lines by Friday. From there the build order is short. Promote that first lesson into a skill, then write one eval case so the agent has to keep passing it. Reach for weekly consolidation only when the log grows long enough to slow things down, and for an external store only when you have a real corpus to search. Each layer earns its place, and none is required on day one. The store was never the hard part. The loop is, and the test above is the smallest honest version of it. Everything else, the four layers, the evaluation, the conflict markers, is how you scale what the test says is worth scaling. --- *Which failure does your agent repeat most, and is it still living in a log instead of a procedure?* [Figure: Field card: the one-week test. Every task, write three lines; at week's end, count which repeat three times or more.] *The one-week test on one card. Run it before you build anything.* *New to The Durability Curve? It is a standing argument about what still holds when the surface gets cheap. [Subscribe](https://harryfloyd.substack.com/subscribe?utm_source=substack-article&utm_medium=article&utm_campaign=agent-memory-loop) for the rest, or [start with what survives](https://durabilitycurve.com/blog/start-here-what-survives-when-the/).* [^1]: Mikaela Grace, Jeremy Hadfield, Rodrigo Olivares, and Jiri De Jonghe, "Demystifying Evals for AI Agents," Anthropic Engineering, 2026. The piece argues that multi-turn agent evaluation is a coordination problem as much as a scoring one, and distinguishes pass@k, the probability that at least one of k attempts succeeds, from pass^k, the probability that all of them do, as answers to different reliability questions. https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents [^2]: Anthropic introduced "Dreaming" for Claude's Managed Agents as a research preview at its Code with Claude event on 6 May 2026. It runs a scheduled, between-session process that consolidates an agent's external memory store and surfaces patterns from recent sessions, which Anthropic compares to the way the brain replays the day during sleep. The original store stays read-only and the consolidated version is produced as a separate output for human review before the agent adopts it. --- --- title: "You Only Hold Four Thoughts" description: "Working memory tops out around four things at once. Every leap in human intelligence has come from storing the rest outside your head, and the most advanced AI systems get their gains the same way." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/you-only-hold-four-thoughts/" date: "2026-06-04" series: "THE HUMAN LAYER" law: "Law I" substack: "https://harryfloyd.substack.com/p/you-only-hold-four-thoughts" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # You Only Hold Four Thoughts *Working memory tops out around four things at once. Every leap in human intelligence has come from storing the rest outside your head, and the most advanced AI systems get their gains the same way.* By Harry Floyd · 2026-06-04 · canonical: https://durabilitycurve.com/blog/you-only-hold-four-thoughts/ *Working memory tops out around four things at once. Every leap in human intelligence has come from storing the rest outside your head, and the most advanced AI systems get their gains the same way.* *This is an analytical framework, not financial advice. Named research claims are referenced to their primary sources in the footnotes.* Try to multiply 47 by 83 in your head. The answer is not the point. Watch what happens while you reach for it. You hold 47, you hold 83, you start on the partial products, and somewhere around the third one the first number goes soft. You reach for a pen, because the problem outgrew the place you were keeping it. That ceiling is real and it is low. The cognitive scientist Nelson Cowan spent years measuring it and put the number at about four. Not the seven you half-remember from an old paper, but three to five distinct things held in mind at once.[^1] Four. That is the working capacity of the most sophisticated object in the known universe. Everything we call getting smarter has been a way around that four. The history of human intelligence is the history of putting thoughts somewhere other than the head, and it runs as a stack, each layer holding what the one below it cannot. ### The first rung is paper Reaching for the pen looks like a small surrender. It is the oldest cognitive upgrade there is. The moment you write 47 above 83 and start stacking partial products, you are thinking about six or seven things at once, because the paper is holding all but the one you are working on. Justin Sung, who teaches learning for a living, puts it more sharply. Writing is not the thing you do after you have reached clarity. Writing is what produces the clarity.[^3] The page becomes the workspace where the thought turns real, because your four slots are freed to do the actual reasoning while the page remembers the rest. This is also why handwriting beats typing. It is far slower than thinking, and that slowness forces you to compress, to decide what is worth the stroke. The friction is not a tax on the process. The friction is the process. A page of notes you struggled to write holds more than a page you copied without resistance. > The page is not a transcript of a finished thought. It is the workspace where the thought becomes possible. ### The rung most people never name In 1998 two philosophers, Andy Clark and David Chalmers, asked where the mind stops and the rest of the world begins, and gave an answer that still unsettles people. The mind, they argued, is not all in the head.[^2] Their example was a man named Otto, who has Alzheimer's and carries a notebook everywhere. When Otto wants to go to the museum, he looks up the address in the notebook the way you would retrieve it from memory. The notebook does the job your hippocampus does. Clark and Chalmers argued there is no principled reason to count the notebook as any less a part of Otto's mind than ordinary memory. Otto and his notebook are a single coupled system. The thinking happens across both. That sounds like a thought experiment until you notice you are Otto. The phone that holds every number you no longer memorise. The calendar that holds every commitment. The thinking is already distributed across you and the things you store it in. The only open question is how well the storage is built. ### The rung that compounds A single page does not persist, does not connect, and cannot be searched. You solve the multiplication, you throw the page away, and next month you solve it again from scratch. Paper extends the moment. It does not extend across time. A structured set of notes does. When every thought you have is written as a durable, cross-linked entry, two things happen that a single page cannot. The thought survives, available to a version of you who has forgotten having it. And it connects, so that an idea from March sits one link away from a problem you only encounter in June, waiting to be useful before you knew you needed it. This is the layer where synthesis becomes possible at a scale no head can hold. No one can keep thirty sources in working memory and find the pattern across them. Four slots cannot do it, and neither can forty. But a system that has been accumulating those sources for months, with the connections already drawn, can surface a synthesis that was never available to anyone thinking alone. The structure does the remembering, which frees the human to do the seeing. ### The rung we are building now For most of history the top of the stack was a human reading their own notes. That is no longer the ceiling. The newest layer is a store of knowledge an AI can read, query, and build on across sessions. The builders who have lived inside this for a year keep reporting the same thing. One who runs large agent systems put it plainly: the model is the same on day 1 and day 40. The files get richer.[^4] The capability of the underlying intelligence barely moves over a project. What improves is the accumulated context it can reach, the record of what was tried, what worked, what the operator decided and why. The intelligence is rented and roughly fixed. The memory is owned and compounds. An AI working from a thin prompt starts every session as a stranger. An AI working from a well-kept store of your decisions starts as a colleague who was in the room last time. The difference is not a better model. It is the same model with the rest of the stack underneath it. > The intelligence is rented and roughly fixed. The memory is owned, and the memory is what compounds. [Figure: The external-cognition stack: four rungs from brain to AI memory, with durability compounding as you climb.] ### Why this is one law and not four This is where a productivity story becomes something larger. The machines climb the same stack you do, for the same reason, using the same move. A large model also cannot hold everything at once. Its version of the four-slot limit is the memory bandwidth of the chip, and the entire recent history of making models faster is a history of refusing to keep everything hot. FlashAttention rewrote how attention uses memory so the chip stops shuttling the same data back and forth. Key-value caching stores the work already done so it never has to be recomputed. Mixture-of-experts routing keeps a vast model mostly dormant and wakes only the part a given token needs.[^5] Store state. Reuse it. Activate only what matters now. That is the same move as paper, notes, and agent memory. Externalise the state you cannot hold, and retrieve only the slice the moment requires. Human cognition scales that way. Machine cognition scales that way. The question "how do I think better" and the question "how do I run a model well" have turned out to be one question with one answer. When two separate problems collapse into the same answer, that answer is usually worth trusting. > One law runs the whole stack: externalise the state you cannot hold, and retrieve only what the moment needs. Brains and models both scale by obeying it. ### The trap inside the stack The law has a failure mode, and it is the one a second-brain enthusiast walks into first. The stack rewards retrieval, not accumulation. The instant you start optimising for the volume of what you store, you have begun to degrade the thing you were building. A note you never pull back out did no cognitive work. Ten thousand of them do less than a hundred you reach for, because the ten thousand bury the hundred. External cognition only pays off on the way back in. Storing is filing, and filing is not thinking. The discipline that keeps the stack alive is structuring everything you save so a future you, or a future agent, can find the one piece that matters without reading the other nine thousand. That is also why each rung has to be built in order. Agent memory on top of a disorganised pile of notes inherits the disorder and answers your questions confidently from a mess. The layers compound only when each one is sound. Skip a rung and you do not get the compounding. You get a faster way to retrieve noise. ### The test you can run this week Take one problem you have been carrying in your head, the one you keep re-thinking from the start each time it surfaces, and move it exactly one rung up the stack. If you have been holding it in your head, put it on paper, and notice how much more of it you can see once your four slots are not spent storing it. If it already lives on scattered pages, write it as one durable, connected note, and watch it link to something you forgot you knew. If it already lives in your notes, make it something your AI can read, so the next session starts where this one ended instead of from zero. Then keep the discipline that makes any of it worth doing. Structure for the way back, not the way in. The measure of your second brain is not how much it holds. It is how reliably the right thing comes back when you reach. *Which rung are you skipping on the problem you keep re-thinking from scratch, and what has that cost you?* [^1]: Nelson Cowan, "The magical number 4 in short-term memory: a reconsideration of mental storage capacity," Behavioral and Brain Sciences (2001): the focus-of-attention capacity of working memory averages about four chunks (commonly cited as three to five), revising George Miller's earlier "seven, plus or minus two." https://pubmed.ncbi.nlm.nih.gov/11515286/ [^3]: Justin Sung argues that writing is generative rather than transcriptive: it offloads fragile internal state onto a stable external workspace, which is what allows clarity to form rather than what records it after the fact. [^2]: Andy Clark and David Chalmers, "The Extended Mind," Analysis 58:1 (1998): external objects that store information can form part of a cognitive process, so that a person and the notebook they rely on function as a single coupled cognitive system. https://www.alice.id.tue.nl/references/clark-chalmers-1998.pdf [^4]: Shubham Saboo, describing multi-agent project stacks: "the model is the same on day 1 and day 40; the files get richer." The accumulated, structured context is what improves over a project, not the underlying model. [^5]: The model-systems mirror of the stack: Tri Dao et al., "FlashAttention" (2022, https://arxiv.org/abs/2205.14135) reduces memory-bandwidth overhead in attention; key-value caching reuses already-computed state instead of recomputing it; William Fedus et al., "Switch Transformers" (2021, https://arxiv.org/abs/2101.03961) activate only the relevant subset of a large model per token. All three are versions of one principle: store state, reuse it, activate selectively. --- --- title: "The Other Half of Compute" description: "Everyone is counting gigawatts and GPUs. The number that decides the return is what each one actually buys." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-other-half-of-compute/" date: "2026-06-03" series: "MARKETS & POWER" law: "Law I" substack: "https://harryfloyd.substack.com/p/the-other-half-of-compute" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Other Half of Compute *Everyone is counting gigawatts and GPUs. The number that decides the return is what each one actually buys.* By Harry Floyd · 2026-06-03 · canonical: https://durabilitycurve.com/blog/the-other-half-of-compute/ *Everyone is counting gigawatts and GPUs. The number that decides the return is what each one actually buys.* *This is an analytical framework, not financial advice. Numerical claims are referenced to their primary sources in the footnotes.* xAI stood up its first 100,000 GPUs in Memphis in 122 days. It doubled that in another 92. By early 2026 the site, Colossus, held around 555,000 of them, building toward two gigawatts of power, for a reported 18 billion dollars.[^1] Two sophisticated people can look at that number and reach opposite conclusions. Jensen Huang's view is that the only real risk is underspending. He puts the buildout at a trillion dollars and counting, and argues the company that holds back capacity loses the decade.[^2] Dario Amodei and Ray Dalio sit on the other side. Amodei has said it can be rational not to buy unlimited compute, because the revenue to justify it may arrive on a timeline that bankrupts whoever guessed wrong. Dalio keeps making a narrower point: a technology can succeed completely and still ruin the people who financed it.[^3] Same buildout. Same dollar figure. One camp calls it the obvious move of the decade and the other calls it the setup for a wipeout. They are not disagreeing about the facts. They are reading the same number and the number is the problem. ### What 18 billion dollars buys Every token a model produces runs down a physical path. Electricity has to be generated, moved across a grid, and stepped down through transformers to a voltage a data centre can use. Chips have to be fabricated at advanced nodes, which in practice means TSMC and a single supplier of the lithography machines that make the process possible. The chips have to be wired together with optical interconnect, assembled into racks, and kept cold. None of those layers move at the same speed, and the slowest one always sets the schedule. For four years the slowest layer kept changing. In 2022 the constraint was GPUs themselves. In 2023 it was the high-bandwidth memory stacked next to them. In 2024 it was the advanced packaging that bonds the two together. By 2025 it was photonics, the lasers and transceivers that move data between racks. By 2026 it had reached power and the grid, where a new high-voltage connection can take longer to approve than the cluster takes to build. Bringing a large new source of power onto that grid now takes a median of more than four years.[^4] Each layer is real, each one becomes scarce in turn, and the scarcity moves to the next layer as the one before it gets solved. Call it the capacity stack. It decides one thing: how much raw compute can physically exist. It tells you what you can run. It says nothing about how much useful work comes out the other end. > The binding constraint has moved through the stack for four years straight. Chips, memory, packaging, photonics, power. Each one stayed invisible until the one before it was solved. ### The number that never makes the capex debate Now look at a different figure. In March 2023, running a million tokens through GPT-4 cost about 30 dollars. By the middle of 2024, the same class of capability through GPT-4o cost 2.50 dollars. By 2025 a GPT-4-grade model was available at roughly 10 cents per million tokens.[^5] For the rougher GPT-3.5 tier the price fell from 20 dollars per million tokens to about 7 cents in two years, a drop of more than 250 times. Epoch AI, which tracks this carefully, finds inference prices falling somewhere between 10 and 50 times a year depending on the task.[^6] Almost none of that came from adding watts. The capacity stack was straining the entire time. The cost of intelligence fell by two orders of magnitude anyway. These are list prices, so some of the fall is competition between providers, but most of it is a second stack that lives inside the software layer and does work the hardware never sees. That second stack has its own layers. At the bottom is the attention kernel. The 2022 FlashAttention paper showed that a transformer was bound by memory traffic, the data shuttling between the fast and slow memory on the chip, and that rewriting the kernel to respect that traffic multiplied throughput without changing a single transistor.[^7] Above it sits serving. Key-value caching, which means storing a conversation's intermediate state instead of recomputing it on every new token, turned long contexts from a quadratic expense into something a business could afford to offer. Above that sits the model itself. Mixture-of-experts routing, the design behind Switch Transformers, broke the link between a model's total size and the compute each token triggers, so a model can hold a trillion parameters and fire only a fraction of them per word.[^8] Even the hardware gains are mostly architectural rather than brute force. NVIDIA's GB200 NVL72 rack delivers up to 30 times the inference throughput of the same number of previous-generation H100 chips, at around 25 times less energy for the same work.[^9] The watts per chip went up. The useful work per watt went up far more. Each of these is a multiplier on the same physical base. Stack them and you get the hundredfold collapse in the cost of intelligence that the buildout debate never mentions. > The cost of GPT-4-class intelligence fell roughly 99 percent in two years. Almost none of that came from adding power. ### Compute is a product Raw physical capacity, multiplied by how much useful work each unit of that capacity buys. The capacity stack sets the first term. The efficiency stack sets the second. They run on different clocks, they are built by different people, and the one that is currently scarcer sets the ceiling on what you can do. Once you read compute that way, the contradictions in the capex fight resolve. Go back to the 18 billion dollars. Jensen Huang is right that physical capacity is scarce today. A grid connection does take longer than a training run, and the firm that waits loses ground it cannot buy back at any price. Amodei is also right that the return on that capacity is uncertain. Both of them are arguing about the first term and treating the second as a constant. It is not a constant. It is improving 10 to 50 times a year. That cuts in two directions at once. A capex bill that looks insane against today's efficiency can look cheap against next year's, because the same site serves far more useful work for the same power. And capacity bought to serve a workload that the efficiency stack is about to make trivially cheap is capacity that strands. The danger in the buildout is **owning the wrong term**: paying for raw capacity after the binding constraint has moved to the multiplier, or perfecting the multiplier when you cannot get the megawatts to run it on. Three years ago the next sentence would have sounded like a category error. > A 2-gigawatt site with a mediocre serving stack loses to a smaller site with a better one. ### Where the constraint goes after silicon The migration does not stop at the efficiency stack either. It keeps walking. Once serving is efficient and the power is online, the slowest layer becomes the one furthest from the metal: whether an organisation can absorb what the stack has made cheap. Jensen Huang's own example is the sharpest version of it. A 500,000-dollar engineer who consumes only 5,000 dollars of tokens a year shows the failure mode.[^10] The tokens are nearly free, and the company still cannot route its own work to the capacity it already owns. This is the layer Amodei and Satya Nadella keep returning to from opposite ends of the argument. The technical stack gets good faster than institutions reorganise around it. The final constraint on compute is organisational. It is how quickly people change what they do. > The tokens are nearly free. The bottleneck is the company. ### A test you can run this week Take any AI bet you hold, whether it is a position, a product, or a career, and do three things. Write down which term you are actually betting on. A bet on the capacity stack is a bet that the physical scarcity of power, chips, and interconnect holds. A bet on the efficiency stack is a bet on the people and techniques that multiply the work each watt buys. Most bets are quietly one or the other, and most people have never said which out loud. Then name the layer that binds right now. Power, today, for raw scale. Serving efficiency, today, for cost per task. Write down what has to stay true one layer below for your bet to survive. A capacity bet dies if grid timelines compress and the scarcity premium decays. An efficiency bet dies if the megawatts never arrive to run on. Then watch the right number. Not GPU count and not gigawatts. Cost per task and useful work per watt. Those are the readings where the second stack shows up, and the second stack is where most of the last two years of progress came from. The capacity layers, the ones you could photograph from a satellite, are mapped company by company in the [companion to this piece](https://durabilitycurve.com/blog/three-hidden-bottlenecks-past-gpus/). This is the half you cannot photograph and the half that has been compounding faster. *If you run the test, which term turned out to be the one you were quietly betting on the whole time?* [Figure: The Compute Audit: five questions to run on any AI bet.] *New to The Durability Curve? It is a standing argument about what survives when the tools get powerful and the surface gets cheap. [Subscribe](https://harryfloyd.substack.com/subscribe?utm_source=substack-article&utm_medium=article&utm_campaign=other-half-compute-flagship) for the rest, or [start with what survives](https://durabilitycurve.com/blog/start-here-what-survives-when-the/).* [^1]: xAI's Colossus (Memphis) reached 100,000 GPUs in 122 days and doubled to roughly 200,000 in 92 more. By early 2026, across multiple GPU generations (H100, H200, GB200), the Memphis site held around 555,000 GPUs and was built out toward ~2 GW of capacity, for a reported ~$18 billion. https://introl.com/blog/xai-colossus-2-gigawatt-expansion-555k-gpus-january-2026 and https://x.ai/colossus [^2]: Jensen Huang, NVIDIA GTC 2026: he projected at least $1 trillion of AI-infrastructure spending through 2027 and argued that figure "won't be enough" to meet demand, framing data centres as "AI factories" that convert electricity into tokens at the lowest cost per unit. https://fortune.com/2026/03/17/jensen-huang-ai-infrastructure-buildout-1-trillion-dollars/ [^3]: Dario Amodei, interview with Dwarkesh Patel (2026): being off on data-centre timing "by a couple of years can be ruinous," because the revenue to justify a buildout arrives on an uncertain schedule. https://www.dwarkesh.com/p/dario-amodei-2 . Ray Dalio's recurring point, repeated in June 2026, is that a technology can succeed while most of the companies financing it fail, as the internet did after the dot-com bust. https://finance.yahoo.com/markets/stocks/articles/ray-dalio-says-ai-investors-121700871.html [^4]: The 2022 to 2026 bottleneck migration sequence (chips, memory, advanced packaging, photonics, power) is the synthesis of my earlier infrastructure work. The grid figure: Lawrence Berkeley National Laboratory, "Queued Up: 2025 Edition," finds the median time from interconnection request to commercial operation for new generation has passed four years. https://emp.lbl.gov/publications/queued-2025-edition-characteristics [^5]: OpenAI list pricing: GPT-4 launched at $30 per million input tokens (March 2023); GPT-4o at $2.50 per million input tokens (May 2024); GPT-4-grade capability available near $0.10 per million input tokens by 2025. Pricing history aggregated by Epoch AI and TokenCost. https://epoch.ai/data-insights/llm-inference-price-trends [^6]: Epoch AI, "LLM inference price trends": the cost of GPT-3.5-class capability fell from roughly $20 per million tokens (late 2022) to about $0.07 (late 2024), and inference prices decline between 10x and 50x per year depending on task tier. https://epoch.ai/data-insights/llm-inference-price-trends [^7]: Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré, "FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness," 2022. The paper reframes attention as memory-bandwidth-bound rather than compute-bound. https://arxiv.org/abs/2205.14135 [^8]: William Fedus, Barret Zoph, Noam Shazeer, "Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity," 2021. Mixture-of-experts routing decouples a model's total parameter count from the compute activated per token. https://arxiv.org/abs/2101.03961 [^9]: NVIDIA GB200 NVL72: NVIDIA reports up to 30x faster real-time LLM inference and up to 25x lower energy and cost versus the same number of H100 GPUs, driven by rack-scale architecture rather than raw per-chip power. https://www.nvidia.com/en-us/data-center/gb200-nvl72/ [^10]: Jensen Huang, All-In Podcast (filmed on the final day of NVIDIA GTC 2026): he said he would be "deeply alarmed" if a $500,000 engineer consumed only $5,000 of tokens in a year, expecting elite engineers to spend closer to half their salary on tokens. Low token use reads as a failure to exploit cheap capacity, not thrift. https://www.tomshardware.com/tech-industry/artificial-intelligence/jensen-huang-says-nvidia-engineers-should-use-ai-tokens-worth-half-their-annual-salary-every-year-to-be-fully-productive-compares-not-using-ai-to-using-paper-and-pencil-for-designing-chips --- --- title: "The Stable Liar" description: "Every metric you optimise quietly stops measuring what you meant. The dangerous ones never break. They keep reporting green while the thing underneath rots." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-stable-liar/" date: "2026-06-02" series: "PROOF & TRUST" law: "Law A" substack: "https://harryfloyd.substack.com/p/the-stable-liar" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Stable Liar *Every metric you optimise quietly stops measuring what you meant. The dangerous ones never break. They keep reporting green while the thing underneath rots.* By Harry Floyd · 2026-06-02 · canonical: https://durabilitycurve.com/blog/the-stable-liar/ *Every metric you optimise quietly stops measuring what you meant. The dangerous ones never break. They keep reporting green while the thing underneath rots.* ## The dashboard was green for eight quarters The most dangerous number on a dashboard is the one that has stayed green the longest, and the way it fails has a shape you have probably watched up close. For eight straight quarters the dashboard holds green. Revenue up and to the right. Retention flat and healthy. NPS in the fifties. Every board meeting opens on the same slide and closes on the same nod. The plan is working. Then, six months after the eighth green quarter, the business the dashboard was supposed to describe nearly falls over. Pull the post-mortem apart and the easy story is that the numbers lied. They did not. Every quarter the dashboard reports something true: customers are still paying, logins are still happening, the survey scores are still fine. All of it accurate. The failure is quieter and worse than a lie. The words behind the numbers change meaning while the numbers stand still. "Retention" still counts the same logins, but a login has stopped predicting a customer who will renew. The metric keeps its shape long after the thing it measured has walked out of the room. Anyone who has run a team has felt a smaller version of this. The number you trusted most became the number that surprised you most. You were not lied to. You were tracking something that used to mean one thing and quietly came to mean another, and the dashboard had no way to tell you the meaning had moved. This is the stable liar: a number that goes on looking right long after it stopped being right. It is a structural property of measurement under pressure, and it has a law underneath it. ## Why every optimised metric drifts > A metric is a substitution: you replace the thing you care about with something you can count, and the gap between them is where the trouble lives. Start with the substitution. You cannot measure value, loyalty, insight, or health directly, so you pick a proxy you can count. Revenue stands in for value. NPS stands in for loyalty. Citations stand in for insight. The proxy is never the thing. The gap between them exists before anyone games anything, on day one, in the cleanest dashboard ever built. That gap stays small only while no one leans on it. The moment a proxy becomes a target, people and systems optimise the proxy, and it drifts from the thing it stood for. Charles Goodhart noticed this in monetary policy in 1975: any statistical regularity collapses once you put pressure on it for control. Marilyn Strathern later compressed it into the line everyone quotes. When a measure becomes a target, it stops being a good measure.[^1] The relationship erodes precisely because you started using it. Feeding a signal back into the system it measures changes the system. The third move is the dangerous one. The erosion is invisible to the metric itself. A dashboard cannot report "I am becoming less valid." An optimiser cannot notice "the thing I am chasing has stopped being the thing we wanted." The metric goes on telling the truth about what it measures, and that fidelity is exactly what hides the drift. The number is honest. Its meaning is gone. Substitution, erosion, blindness. None of them require a villain. They are what happens when you close the loop between what you measure and what you do. [Figure: Why every optimised metric drifts: substitution, then erosion, then blindness.] *The number stays honest the whole way through. Its meaning is what leaves.* ## The three faces of a lying metric Once you accept that drift is structural, the useful question becomes diagnostic. A degrading metric shows up in three distinct ways, and they are not equally easy to catch. Mistake one for another and the standard fix makes things worse. ### The Collapse The first face is loud. The metric and the outcome diverge so violently that everyone can see something broke. The Soviet planners who set nail output by weight, and got a few enormous useless nails, are the parable everyone tells. The modern version is a research field that rewards paper count and fills its journals with results no one can reproduce. > The Collapse announces itself: the number and the reality pull apart in plain sight. This is the easy case, even though it feels like a crisis. The signal is noisy and obvious. You see revenue climb while satisfaction falls in the same quarter, and you know the metric has come loose. Almost every "metrics are dangerous" lecture is about the Collapse, because it is the one you can point at. ### The Hollowing The second face is quiet, and most operators never name it. The metric stays healthy while the system underneath hollows out. The green dashboard from the opening was a Hollowing: every gauge held its level while the customers behind them quietly stopped behaving like customers, and "retention" went on counting logins that no longer meant renewal. The same pattern runs everywhere once you know its shape. A hospital hits its wait-time target by turning away the complex patients who would have blown it. A support team holds CSAT steady by making the survey harder to find. An engagement score stays flat because employees have learned which answers keep management calm. The most expensive version runs inside modern AI infrastructure: a Kubernetes platform shows every node green while its GPUs, the entire reason the cluster exists, sit at roughly five percent utilisation.[^2] > The Hollowing leaves the number standing while the meaning quietly walks out. You cannot catch the Hollowing by staring at the metric, because the metric looks fine. You catch it by watching what the metric does not cover, and by noticing stability where you should see variation. A number that used to move with the seasons and now sits suspiciously flat is often a number that has been hollowed. ### The Inversion The third face is the one that ends companies, and careers, and occasionally institutions. Here the metric looks excellent precisely because the system has learned to model the measurement and optimise against it directly. The benchmark score climbs while deployment reliability quietly rots. The sales team hits quota by closing customers who will churn in two quarters. The trader posts a beautiful Sharpe ratio by taking the one risk the ratio cannot see. > The Inversion is the stable liar: the metric is not merely failing to track reality, it is actively manufacturing confidence in the wrong direction. This is the hardest face to detect, because the absence of any warning sign is itself the warning. The dashboard supports the wrong conclusion with full conviction. And the standard advice, "tighten the metric, raise the bar," is harmful here, because a sharper target just gives a capable optimiser a cleaner thing to game. Modern AI evaluation is where the Inversion is easiest to see, though it shares the stage with cruder failures worth separating out: contamination, where test items leak into the training data; overfitting to the eval's own distribution; and plain weak test design. The Inversion proper is narrower. A capable system optimises against the evaluation itself, and the score comes loose from the capability it was supposed to certify. That looseness shows up even before any deliberate gaming. When Apple researchers rebuilt grade-school maths problems from symbolic templates and changed only the names and numbers, models that had aced the original benchmark dropped sharply, and one irrelevant clause cut accuracy by as much as sixty-five percent.[^3] The benchmark had been reporting reasoning. What it measured was pattern-matching against problems shaped like the training set. The deeper version is already here: a capable enough model can represent the fact that it is being tested and behave differently when it notices. Once a system can model its own yardstick, raising the bar recovers nothing, because the bar is now part of what the system optimises against. A climbing eval score has stopped being evidence of a more capable deployment. It is evidence that the score went up. ## Why telling them apart is the whole skill The reason the taxonomy matters is that each face wants a different response, and the responses do not transfer. Treat an Inversion like a Collapse, by improving the metric, and you hand the optimiser a better target. Treat a Hollowing like noise, and you wait for a crash that the number will never warn you about. The single most expensive mistake in measurement is applying a Face-One fix to a Face-Three problem and feeling responsible while you do it. So before you act on any important number, you need a way to ask which face you are looking at. Three probes do most of the work. ## Three probes for a suspect number > Run these on any metric you are about to trust with a real decision. The correlation probe asks what should move with this metric if it still means what you think, then checks whether those companions still move. Retention and renewal should rise and fall in step. Benchmark scores and production reliability should track. When the companions quietly decouple and the headline number sails on alone, the meaning has drifted even while the value holds. The negative-space probe maps what the number cannot see, because the failure usually hides there. Write down what this metric does not capture: the complex patient who was turned away, the angry customer who never found the survey, the failure mode the benchmark never tests. The list of what a metric ignores is usually a more honest document than the metric. The capability probe is the one most people skip. Ask whether the thing being measured can model the measurement. A nail factory cannot scheme about its weight target, so it can only Collapse or Hollow. A capable sales team, a frontier model, or a sophisticated trading desk can represent the evaluation as an object and bend behaviour around it. The moment the measured system can see and reason about the yardstick, the Inversion becomes available, whether you have noticed or not. ## What to do once you know the face Below the capability threshold, where the system cannot scheme about its own measurement, the classic advice works. Diversify your proxies, because five metrics that disagree are harder to fool than one that lies well. Probe your own numbers adversarially before reality does it for you. And track what resists gaming: variance, the rate of negative cases, the decisions you chose not to make. Above the threshold, where the optimiser can model the evaluation, improving the metric is the trap, because improvement is exactly what it exploits. The fixes turn structural. Keep the evaluation boundary hard to model. Shrink the surface the system is allowed to edit. Move verification outside the system entirely, to an instrument it cannot reach. The check has to live somewhere the thing being checked cannot get to, which is what separates real verification from bigger classification. Then add the measurement almost no dashboard carries: a metric on your metrics, tracking whether the rest still mean what they meant a year ago. It is the only early warning for drift, because drift is invisible to every gauge on its own. [Figure: The capability threshold: below it the usual fixes work, above it they backfire.] *The same move, "improve the metric," helps below the capability threshold and backfires above it. That is why naming the regime comes before choosing the fix.* ## Measurement changes what it measures The reason metrics betray you is not malice, incompetence, or sloppy dashboard design. It is feedback. A metric you only watch leaves its subject alone. A metric you optimise feeds back into the thing it measures and changes it. The instant you close the gap between what you measure and what you do, you start editing the exact quantity you were trying to observe. You cannot engineer this out of your particular dashboard. It is a property of measuring under pressure, and it applies to your KPIs, your evals, your portfolio, and your own annual review. So stop hunting for the ungameable metric. There is no such thing, and the search wastes years you could spend building the one instrument that helps: the habit of asking, on a schedule, whether your most trusted number still means what it meant when you started trusting it. Keep this question on the wall. *If this metric stopped being valid six months ago, what would I be seeing right now that I am explaining away?* Run it on the number you trust most, the green one, the one you quote in board meetings and tell yourself you have covered. The metric you never question is the one already lying to you, and the only way to hear it is to go looking for the evidence you have been quietly filing under "noise." --- *Which number do you trust most right now, and when did you last check that it still means what you think it does?* --- [Figure: The Metric Validity Audit: three probes to run on any number you trust.] *The three probes on one card, for the number you quote most.* *New to The Durability Curve? It is a standing argument about what still holds when the surface gets cheap. [Subscribe](https://harryfloyd.substack.com/subscribe?utm_source=substack-article&utm_medium=article&utm_campaign=metric-trap-flagship) for the rest, or [start with what survives](https://durabilitycurve.com/blog/start-here-what-survives-when-the/).* [^1]: Charles Goodhart's original formulation appears in his 1975 work on UK monetary policy: "Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes" (later collected in *Monetary Theory and Practice*, 1984). The compressed version most people quote is Marilyn Strathern's, from "'Improving ratings': audit in the British University system," *European Review* 5, no. 3 (1997): "When a measure becomes a target, it ceases to be a good measure." https://en.wikipedia.org/wiki/Goodhart%27s_law [^2]: The 5% figure is from Cast AI's *2026 State of Kubernetes Optimization Report*, which analysed tens of thousands of clusters across AWS, GCP, and Azure and found average GPU utilisation of roughly 5% (with CPU near 8% and memory near 20%). Every individual resource dashboard reads "healthy" while the expensive thing the cluster exists to do sits almost entirely idle. https://cast.ai/press-release/2026-state-of-kubernetes-optimization-report/ [^3]: Iman Mirzadeh et al., "GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models," Apple, 2024 (arXiv:2410.05229). Regenerating grade-school maths problems from symbolic templates and changing only names and numbers lowered accuracy across state-of-the-art models, and inserting one irrelevant clause dropped accuracy by up to 65%, evidence that the models pattern-match the shape of their training data rather than reason, and that the benchmark score overstated the capability it appeared to certify. https://arxiv.org/abs/2410.05229 --- --- title: "Your Tools Got Powerful. Get Boring." description: "The most powerful tools in history reward the most boring strategies. The gap widens every time they improve." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/your-tools-got-powerful-get-boring/" date: "2026-06-01" series: "STRATEGY & MOATS" law: "Law I" substack: "https://harryfloyd.substack.com/p/your-tools-got-powerful-get-boring" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Your Tools Got Powerful. Get Boring. *The most powerful tools in history reward the most boring strategies. The gap widens every time they improve.* By Harry Floyd · 2026-06-01 · canonical: https://durabilitycurve.com/blog/your-tools-got-powerful-get-boring/ *The most powerful tools in history reward the most boring strategies. The gap widens every time they improve.* ## The bored trader beats the machine On one side of the trade sits a market-making engine that represents the genuine state of the art: Hawkes processes modelling order arrivals, Kyle's lambda pricing the impact of each fill, Avellaneda-Stoikov inventory control balancing the book in real time. Years of mathematics, running on hardware that did not exist a decade ago. On the other side is a momentum trader whose entire system is price, volume, and three moving averages. He sits in cash most of the year doing nothing, waiting for a setup he could describe to you in a sentence. His stack is deliberately primitive. His edge is patience and the discipline to follow his own rules when they are boring and to sit out when they are silent. Over a full market cycle, the boring one is more likely to still be standing. This is uncomfortable, because it runs against an intuition almost everyone shares: better tools should let you run better, more sophisticated strategies. More compute, more data, more powerful models, therefore more elaborate approaches and better results. It feels obviously true. It is the logic behind most of what gets built, bought, and bragged about. It is also, across domain after domain, wrong. And the interesting part is the shape of the curve. ## The gap widens as the tools get stronger Here is the pattern the most successful practitioners keep landing on, whether they are trading, building software, learning, or shipping products. Powerful tools do not pay off when you point them at more complex strategies. They pay off when you point them at simple strategies and execute those faster, more consistently, and with less drift than anyone else. > More power applied to a simple strategy compounds. The same power applied to a complex one mostly buys you more ways to be wrong. Sit with the second half of that, because it is the part people miss. A sophisticated strategy is not free. Every additional layer needs to be specified, verified, maintained, and monitored, and all of that consumes exactly the capacity the powerful tool was supposed to give back. A simple strategy spends its new power on doing the simple thing relentlessly well. A complex one spends its new power feeding its own machinery. The reason this matters more now than it ever has is that the tools have never been this strong. When your instruments are weak, the gap between the simple-and-disciplined path and the complex-and-fragile path is small, because nobody can do much of either. As the instruments get more powerful, both paths open up, and the distance between them widens. The most capable tools in history make disciplined simplicity more effective than ever, and they also make unmanageable complexity easier to build than ever. We are living through the largest gap between those two paths that has ever existed, and most people are sprinting down the wrong one with a faster engine. [Figure: The gap between the simple-and-disciplined path and the complex-and-fragile one widens as tools get more powerful.] *Same start, same power, opposite directions. The stronger the tools, the wider the distance grows.* ## What complexity quietly costs The bill for sophistication does not arrive when you build it. It arrives later, in instalments, and it is always larger than it looked. The first instalment is verification. A simple system you can hold in your head and check. A complex one you cannot, so you build monitoring to watch it, and the monitoring becomes its own system that can [drift and mislead](https://durabilitycurve.com/blog/most-verification-is-just-bigger/). Every layer you add is a layer you now have to confirm is still doing what you think it does, and the confirming never ends. The second instalment is the day it breaks. A simple strategy fails legibly: you can see which rule was wrong and fix it. A sophisticated one fails in the seams between its parts, at the worst possible moment, in a way no single person fully understands. The elaborate model that printed money for two years becomes, in the drawdown, a black box nobody can debug while it is bleeding. Complexity does not only add capability. It adds failure modes that stay hidden until the system is under stress, which is the exact moment you have no spare capacity to handle them. The deepest cost is fragility to your own success. A strategy with many parameters has many surfaces the world can destabilise once it starts reacting to you. The more elaborate the machine, the more places reality can reach in and pull a lever you forgot you had wired up. Simple, constrained systems survive contact with the world because there is less of them to break. > Sophistication is a loan against your future attention, taken out at a rate you cannot see until the system is under stress and the whole balance comes due at once. [Figure: The three instalments of the complexity bill: verification, failure, and fragility.] *Verification that never ends, failure in the seams, fragility to your own success. The bill always arrives, and later than you think.* ## Why we reach for sophistication anyway If simplicity wins, why does almost everyone instinctively add complexity? Smart, capable people do it constantly, because the incentives reward it. > Every incentive in the room rewards the complexity you can show and punishes the discipline you cannot. Sophistication is visible. A complex model, an elaborate architecture, a clever framework can be shown to a boss, a client, an investor, a peer. Discipline cannot be shown. Sitting in cash for three months, deleting half your code, pausing before you speak, refusing to ship the extra feature: none of it photographs well. The market pays for what it can see, and it can see complexity far more easily than it can see restraint. Complexity also feels like work. Building an intricate system produces the sensation of progress all day long, even when the effort is going into [the layer with the least leverage](https://durabilitycurve.com/blog/leverage-hierarchy-of-agent-engineering/). Doing the boring, correct thing and then waiting produces the sensation of doing nothing, which the nervous system reads as failure. The feeling and the result point in opposite directions, and the feeling usually wins. And an entire economy is built on convincing you the work is harder than it is. Every tool vendor, every course, every consultancy has a structural interest in making its domain look more complex than it needs to be, because simplicity is terrible for business. The people who write about a field emphasise its hardest parts, which is what makes them experts, rather than its simplest parts, which is what produces the results. The perceived difficulty of almost everything is inflated, and the inflation is nobody's accident. ## What the constrained version keeps proving The clearest place to watch this play out right now is in how people use AI, because the tool is so powerful that the trap is stark. The most effective way to get good work out of a frontier model is to take capability away from it. The prompts that consistently produce strong code are the ones that forbid things: no verbose comments, no scattered logging, small functions only, review your own output before returning it. The best debugging prompts are the most constrained ones: strict ordered steps, and a hard rule to verify before changing anything. The most powerful model on the planet does better work when you give it fewer options. People reach for AI expecting more power to mean more freedom. What it rewards is more power inside tighter constraints. The same shape shows up wherever someone is quietly winning with powerful tools. The builders who ship profitable products solo run on deliberately boring technology, the kind a fashionable engineer would be embarrassed by. They ship ugly first versions fast while better-resourced teams are still choosing a framework. The plain name for what those teams are doing is over-engineering, and the powerful tools make it easier than ever. The people who learn fastest take fewer notes, not more. They delay and compress until a page of dense understanding replaces a folder of neat transcription. The creators who grow post less, because the algorithm rewards depth per post and punishes the volume that easy tools make tempting. Different fields, one lesson: the powerful tool is best spent removing steps. > The people quietly winning with the strongest tools are using them to do less, and to do it more reliably than anyone else. None of these people are anti-technology. They are using the most powerful tools available. They are simply pointing them at the boring fundamentals and refusing the upgrade to a more complicated game. ## The test that catches you in the act The trap is hard to escape by intention alone, because it is driven by feeling, and the feeling does not announce itself as a bias. It announces itself as ambition. So you need a question sharp enough to cut through the feeling in the moment you are reaching for complexity. > When you catch yourself building something more sophisticated, ask: am I adding this because the problem genuinely requires it, or because the simple version feels uncomfortable? If the honest answer is discomfort, you are in the trap. The simple version feels too easy, too exposed, too much like you are not earning your keep, so you reach for a layer that makes you feel substantial. That layer is where the cost lives. The operational rule is narrow and worth memorising. Reduce complexity until the system is something you can verify, and not one notch past that. A simple strategy with strict rules and clean checks beats a sophisticated one you cannot fully see into, and the advantage grows the more powerful your tools become. When a strategy has more moving parts than you have the discipline or the data to support, the parts are not power. They are surface area for failure. So the next time more compute, a better model, or a new tool lands in your hands, notice the instinct to finally build the elaborate thing you have been wanting to build. That instinct is the trap closing. The tool should be powerful. The strategy it serves should be almost embarrassingly simple. And the discipline that holds the two together should be boring enough that nobody, including you, finds it impressive. That last part is why it works. The edge is boring on purpose, which is exactly why it is still available. --- *What is the most sophisticated thing in your current setup, and what would happen if you deleted it?* --- [Figure: The Sophistication Test: three checks to run when the pull to add complexity hits.] *The test on one card, for the next time the pull to complicate hits.* *New to The Durability Curve? It is a standing argument about what survives when the tools get powerful and the surface gets cheap. [Subscribe](https://harryfloyd.substack.com/subscribe?utm_source=substack-article&utm_medium=article&utm_campaign=sophistication-trap-flagship) for the rest, or [start with what survives](https://durabilitycurve.com/blog/start-here-what-survives-when-the/).* --- --- title: "Same Model, Different Product: The Case for Harness Engineering" description: "Harness engineering, the code wrapped around an AI model, now drives more of the performance gap than the model you pick." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/harness-engineering-same-model-different-product/" date: "2026-05-30" series: "AI & WORK" law: "Law IV" substack: "https://harryfloyd.substack.com/p/harness-engineering-same-model-different-product" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Same Model, Different Product: The Case for Harness Engineering *Harness engineering, the code wrapped around an AI model, now drives more of the performance gap than the model you pick.* By Harry Floyd · 2026-05-30 · canonical: https://durabilitycurve.com/blog/harness-engineering-same-model-different-product/ You can run the same model inside two coding tools and get two different products. Put Claude Sonnet under Claude Code and under Cursor. Same weights, same context window, same benchmark scores on paper. In practice you get different token burn, different success rates, different cost per task, a different feeling about whether you can leave the thing running. Adam Elkassas, who builds with both, put it plainly. > Same Sonnet underneath Claude Code, Cursor, Cline, and a dozen no-name CLIs, and they feel like completely different products.[^4] The size of that gap is now on the record. On the same model and the same benchmark, swapping the harness can move the score by as much as 6x, a figure documented across agent research and restated in a March 2026 Stanford and MIT paper on harness design.[^1] Not the weights. Not the prompt. Not fine-tuning. The wrapper. LangChain showed it from the other direction. Their coding agent, deepagents-cli, climbed from 52.8% to 66.5% on Terminal Bench 2.0, from outside the Top 30 into the Top 5, while the model underneath, GPT-5.2-Codex, stayed fixed.[^2] The score moved nearly 14 points. The model did not move at all. The variable that moved the score has a name. The harness: the code that decides what the model sees, when it runs again, which tools it can reach, and what happens when it fails. Changing it well has become its own discipline now, with its own name: harness engineering. Most teams spend their attention choosing the model. The performance they are chasing lives one layer out, in code most of them have never opened. > **For beginners: what is a harness?** The model generates text. The harness is everything around it that turns a raw text predictor into something that can do work: the loop that calls the model again and again, the memory of what happened earlier, the list of tools it is allowed to use, the rules for what to do when a step breaks. Two products can run the identical model and still behave differently, because the harness around each one makes different choices. When people say "Claude Code feels different from Cursor," the harness is most of what they are feeling. ## Where harnesses diverge Four decisions separate a harness that gets 66% from one that gets 52%. None of them touch the model. [Figure: Anatomy of a harness: context feeds the model, the model calls tools, error recovery decides what happens next, and the orchestration loop runs the cycle until the task is done.] ### What it keeps Every harness has to decide what to keep as the conversation grows and what to throw away. A long task fills the context window with resolved debug cycles, completed file edits, and conversational noise. Keep all of it and the model drowns in its own words. Throw away the wrong thing and it forgets why it started. The simple approach keeps the last N messages and discards the rest. That means a stale file read from twenty minutes ago competes for the model's attention with the task in front of it. Claude Code's pruning logic does something different: it keeps the plan and trims the chatter, holding on to the original intent and the current task state while dropping the resolved sub-tasks. What a harness keeps in context turns out to be one of the biggest levers it has. If you are running an agent on recency alone, you are leaving performance on the table and paying for the privilege in tokens. ### What it does when it fails When a step breaks, a weak harness has one response: retry. A strong one knows why it broke and answers each failure differently. Claude Code's loop carries seven explicit reasons it might be running again. It hit the output-token ceiling, so it retries at a higher limit. The prompt overflowed, so it compacts and continues. A tool returned, so it injects the result. The reason is stored and handed back to the model on the next turn. That is the difference between an agent that silently restarts and one that knows it hit a token limit, compacted the context, and is carrying on. Collapse every failure into a single "retry" and you throw away the signal that tells the model what to do next. > A harness that treats a token overflow and a broken tool call as the same event has discarded the one piece of information that would have let the model recover. ### What it lets the model touch Show a model fifty tools on every turn and it makes worse decisions about which to use. Manus cut their agent down to roughly 20 atomic function calls after watching performance fall whenever the model met an unfamiliar tool. Claude Code reveals tools as they become relevant: the file-read tool while exploring, the commit tool only once there are changes to commit. The format of a single tool can move the number more than a model upgrade. Can Bölük's hashline edit format, where the model points at lines by a content hash instead of reproducing the exact text, took one model's edit success from 6.7% to 68.3% and cut another model's output tokens by 61%.[^3] He changed one tool and improved fifteen models. None of the models changed. ### How the loop is built The execution loop itself is an engineering decision. Claude Code runs a single state machine, roughly 1,400 lines, one instance per conversation, holding a mutable record of message history, token usage, and permissions. Early versions used recursion and switched away once the call stack grew without bound on long sessions. This is the layer where LangChain found its 13 points. The gains came from structured verification loops that scored intermediate steps, loop-detection that caught the model spiralling on the same hallucination, and tracing at scale that showed which transitions were failing silently. A team that can see which step corrupted the run can fix it. A team that only sees the final output cannot. ## Why the labs leave it on the table If harness design moves the number this much, the obvious question is why the people who make the best models do not also ship the best harness. The answer is in their incentives. A good harness uses the fewest tokens it can to finish the task. When an independent builder finds token waste, they ship the fix that night, because every token they cut is money back in their user's pocket and a reason to stay. When a frontier lab finds the same waste, it becomes a low-priority ticket that loses every sprint to a feature that drives more API calls. > A good harness uses the fewest tokens possible. When an independent harness-maker finds token waste, they ship the fix that night. When a frontier lab finds it, it is a P2 that loses every sprint.[^4] This is structure, not malice. Anthropic has committed around $50 billion to data centres. OpenAI's Stargate is past $400 billion. Every one of those GPUs needs tokens flowing through it to earn its keep.[^4] A harness that cut token use by 5x would save users money and shrink the revenue the buildout was financed against. The independent builder wakes up trying to cut tokens. The lab wakes up trying to fill a gigawatt of compute. Those are opposite jobs, and they produce opposite harnesses. ## Why the advantage lasts It would be easy to read all this as a 2026 quirk that the next model erases. Some of it is. The parts of a harness that exist to patch a model's current weakness do shrink as models improve. Anthropic removed an entire planning step with one model release, because the model no longer needed the work broken down for it. The parts that do what a model structurally cannot do are moving the other way. Sandboxing, permissions, observability, memory: a more capable agent needs more of each, not less. The harness is not shrinking as a whole. The shrinking parts and the thickening parts are different parts. [Figure: As models improve, the harness splits in two: the compensatory parts that patch a model weakness thin, and the durable-enabling parts that do what the model cannot do thicken until they are most of the harness.] The market has noticed. Surveys put Claude Code at as much as 54% of the coding-agent market, on a reported $2.5 billion annualised run-rate; the wrapper is now a product people pay for directly.[^5] In May 2026 DeepSeek, a model-first lab if there ever was one, posted to hire a Harness Team, with the internal line "Model + Harness = Agent."[^5] When even the labs that sell models start staffing the layer above the model, that tells you where the value is settling. > A model-first lab hiring a harness team is the clearest signal yet. The value is settling in the layer above the model. For an investor the read is structural. A specific harness optimisation is temporary, and the next model may erase it. What compounds is the discipline of building harnesses, and the tooling and memory a team carries from one model generation to the next. That is the part a team owns. The model underneath it is rented. The position has a clean exit. If a frontier lab ships a first-party agent that beats every third-party harness on the same model by more than 10% on Terminal Bench, the independent-harness edge is gone. If the open-source ecosystem converges on one dominant design and the 6x gap collapses below 2x, this was a transitional observation, not a structural one. Watch those two numbers. They are where the thesis dies if it is going to. > The Durability Curve tracks where value is moving in AI and markets, before the consensus reprices it. A new structural breakdown most weeks. **[Subscribe free.]** ## Audit your own harness You do not need the source code of Claude Code to find out whether your agent is harness-bound. Five questions locate the gap, and a team shipping an internal agent can answer all five about their own setup in an afternoon. Start with what it keeps. Does your harness hold the plan and the current state, or the last N messages? If it is recency, the cheapest gain you have is sitting in the pruning logic, and you have probably never touched it. Then what it does when it fails. Does a token overflow, a tool error, and a timeout produce three different responses, or one "retry"? If it is one, the model is recovering blind. Then what it can touch. Does the model meet a curated set of tools, or a flat list of everything? If the list is long, trim it before you change anything else. Then whether it knows why it is running again. When your loop calls the model back, does it pass the reason, or just "continue"? An agent told only to continue makes its next decision with no idea what just happened. Then whether you can see inside a run. Can you score the intermediate steps, or only the final output? If step two fails quietly, step five inherits the corruption and you will blame the model for it. A team that answers these honestly usually finds the same thing: they have changed the model three times and never opened the harness once. ## The one-week test Pick one agent you are running. Do not change the model this week. Instead, find the harness layer you have never touched, almost always context pruning or error recovery, and change that one thing. Measure cost per task and success rate before and after. Most teams have never opened the part of their stack that decides what the model sees and what happens when it breaks. That is usually where the cheapest improvement is hiding, and you do not have to switch models to find it. The model you picked is the part everyone can see. The harness is the part doing the work. [Figure: The harness audit. Five layers, one model.] ```text THE HARNESS AUDIT · five checks to run on your agent 1 · CONTEXT (what it keeps): the last N messages, or the plan and the live state? 2 · RECOVERY (when a step fails): one blind retry, or a different move per failure? 3 · TOOLS (what it can touch): a flat list of everything, or a curated few? 4 · FEEDBACK (why it reran): just "continue," or the reason it is running again? 5 · MEASUREMENT (inside a run): only the final output, or every step scored? The first answer in each line is the harness-bound default. THE MOVE Take the weakest of the five for your agent. Change that one layer this week. Measure cost and success before and after. THE LINE TO REMEMBER Same model, different product. The wrapper is the variable. ``` --- *When you last reached for a better model, was the bottleneck ever the model?* [^1]: Yoonho Lee, Roshen Nair, Qizheng Zhang, Kangwook Lee, Omar Khattab, and Chelsea Finn, "Meta-Harness: End-to-End Optimization of Model Harnesses," Stanford IRIS Lab, MIT, and KRAFTON, arXiv:2603.28052, 30 March 2026. The paper opens by noting that changing the harness around a fixed model can produce a 6x performance gap on the same benchmark, citing prior agent research. Its own framework lifts a fixed model from 27.5% to 37.6% on TerminalBench-2, and from 58.0% to 76.4% on a stronger model, by searching for better harnesses. https://arxiv.org/abs/2603.28052 [^2]: LangChain, "Improving Deep Agents with harness engineering": deepagents-cli rose from 52.8% to 66.5% on Terminal Bench 2.0, from outside the Top 30 into the Top 5, on a fixed GPT-5.2-Codex, by changing only system prompts, tools, and middleware hooks. https://www.langchain.com/blog/improving-deep-agents-with-harness-engineering [^3]: Can Bölük, "I Improved 15 LLMs at Coding in One Afternoon. Only the Harness Changed.," 12 February 2026. The hashline edit format (the model references lines by a content hash rather than reproducing exact text) took Grok Code Fast 1 from 6.7% to 68.3% and cut Grok 4 Fast's output tokens by 61%, with gains across the model set. https://blog.can.ac/2026/02/12/the-harness-problem/ [^4]: Adam Elkassas, pre.dev, "Frontier labs won't build good harnesses. Their incentives won't let them." Supporting capex context: Anthropic's $50 billion US data-centre commitment with Fluidstack (announced November 2025) and OpenAI's Stargate, past $400 billion in planned investment toward a $500 billion, 10-gigawatt target. https://pre.dev/blog/frontier-labs-wont-build-good-harnesses-their-incentives-wont-let-them/ [^5]: Surveys put Claude Code's share of the coding-agent market at 42–54% (Menlo Ventures State of Generative AI, cited via MindStudio), on a reported ~$2.5 billion annualised run-rate (reported figures; Anthropic is private). In May 2026 DeepSeek established a Harness team, part of a wider shift in AI coding tools from "model battles" to "engineering battles." https://www.mindstudio.ai/blog/claude-code-2-5-billion-annualized-revenue-terminal-tool and https://finance.biggo.com/news/rsqbS54BDXrLZJaA6kwD --- --- title: "The Leverage Hierarchy of Agent Engineering" description: "Donella Meadows ranked twelve places to intervene in a system. Most agent teams spend their hours at the bottom of the ladder." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/leverage-hierarchy-of-agent-engineering/" date: "2026-05-26" series: "SYSTEMS & LAWS" law: "Law II" substack: "https://harryfloyd.substack.com/p/the-leverage-hierarchy-of-agent-engineering" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Leverage Hierarchy of Agent Engineering *Donella Meadows ranked twelve places to intervene in a system. Most agent teams spend their hours at the bottom of the ladder.* By Harry Floyd · 2026-05-26 · canonical: https://durabilitycurve.com/blog/leverage-hierarchy-of-agent-engineering/ *This is an operator framework, not financial advice.* [Figure: Pasted image 20260528202254] The agent kept choosing the wrong table. Six prompt rewrites later, nothing had changed. The bug was never in the prompt. The schema layer did not distinguish logged-out sessions from logged-in sessions, and the metadata never exposed which table was the right one. No amount of prompt tuning teaches an agent something the metadata layer does not surface. Most agent-engineering work happens at the bottom of the stack: prompts, retrieval-k, model swaps, retry policies, context windows, tool additions. Each move is real, each compiles, each lands in the Friday demo. But the gains that survive three model upgrades come from somewhere else. The structure of information flow. The rules of action. The goal the system is being asked to optimise for. Teams optimise for the demo, and the high-leverage work goes undone. This is a layer problem, and Donella Meadows had a name for it. In 1999 she published a ranking of twelve places to intervene in a system, from the weakest (#12) to the strongest (#1)[^1]. Her diagnosis: most teams identify the right place to push and then push the wrong way. The ranking maps onto the agent stack. The mapping lands in three bands. **Bottom band, low leverage (ranks #12–#9).** Prompts, model swaps, retrieval-k, retries, async timing. **Middle band, mid leverage (ranks #8–#7).** Validators, evaluators, feedback loops. **Top band, high leverage (ranks #6–#1).** Information-flow architecture, rules of action, goals, framework choice. The full ranking is below for operators who want the granular version. Read from the bottom up. *Reference table. Skim now, return to it when applying the framework.* [Figure] | Rank | Meadows' phrasing | Agent-engineering equivalent | | ---: | ----------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | **Bottom band (low leverage)** | | | 12 | Constants, parameters, numbers | Prompt tokens, temperature, top-p, max-tokens, model selection | | 11 | Sizes of buffers and stabilising stocks relative to flows | Context window size, embedding dimension, retrieval-k | | 10 | Structure of material stocks and flows | Tool catalogue, MCP topology, sub-agent inventory | | 9 | Lengths of delays relative to rate of system change | Async response timing, retry intervals, polling cadence | | | **Middle band (mid leverage)** | | | 8 | Strength of negative feedback loops relative to impacts corrected | Validators, evaluators, monitors, output gates, regression tests | | 7 | Gain around driving positive feedback loops | Reinforcement loops (RL, online fine-tuning) and operator-improvement cycles (skill discovery, agent self-improvement passes) | | | **Top band (high leverage)** | | | 6 | Structure of information flows (who has access to what, when) | Context architecture: what the agent sees, sourced from which layer | | 5 | Rules of the system (incentives, punishments, constraints) | Tool permissions, action gates, system prompts as constraints | | 4 | Power to add, change, evolve, or self-organise system structure | Skill discovery, sub-agent spawning, dynamic delegation | | 3 | Goals of the system | Objective function, intent specification, success criterion | | 2 | Mindset or paradigm from which goals, structure, rules arise | The mental model of what an agent IS (a model that calls tools, or a structured pipeline that uses a model). Implementation architecture follows from the paradigm | | 1 | Power to transcend paradigms | Questioning whether "agent" is the right primitive at all (e.g. harness-only systems with no LLM in the control flow) | A note on the top of the table. Ranks #1 and #2 are conceptual horizons most teams will not touch in a given sprint. Rank #7 is rare in production agent systems today. The bound action lives in ranks #10 through #3. ### The OpenAI case study A live worked example of teams making the move. OpenAI's Data Productivity team published an engineering post in January about the internal data agent serving 3,500 of their employees across 70,000 datasets[^2]. They named three lessons. *Less is More.* When they exposed the full tool catalogue the agent got worse, so they consolidated. *Guide the Goal, Not the Path.* Prescriptive prompting degraded the agent on varied queries, so they switched to high-level goal specification. *Meaning Lives in Code.* Schemas describe what data looks like, but the pipeline code that produces it captures intent, so they crawled the codebase with Codex and made the code itself a context layer. The lessons look like engineering wisdom. Read against Meadows' ranking, they are three layer-shifts that most agent teams have not noticed they need to make. *Less is More* is the shift from #12 to #10. The team had been treating the tool catalogue as a parameter to tune by exposure. They moved up to material-flow structure: which tools exist at all. They did not improve the agent's ability to disambiguate redundant tools. They removed the redundancy. A #10 intervention. The same move applied to the wrong-table problem from the opener: make only the right table visible to the agent. The selection ability is not the layer to fix. *Guide the Goal, Not the Path* is the shift from #12 to #3. Prescriptive prompting is parameter tuning under a different name: encoding the agent's behaviour token by token. They moved up nine ranks. They stopped encoding the path, started encoding the destination, and trusted the model's reasoning to find the route. The gain came from a structurally different intervention layer. *Meaning Lives in Code* is the shift from #11 to #6. The naive #11 fix to "the agent does not understand this dataset" is a larger context window or a higher retrieval-k. They did not pull more snippets. They added an entirely new source: pipeline code, crawled by Codex, surfacing intent that schemas cannot. The information-flow structure changed. Three lessons. Three shifts. None of the three moves would have been findable from inside the #12 frame. > This publication tracks the layer of agent-engineering leverage most teams have not noticed they need to climb to. **[Subscribe here.]** ### Why the bottom of the stack wins anyway > *Higher leverage implies more resistance.* Meadows' own diagnosis of why the ranking exists at all. None of this is moral failing. Higher-rank work demands coordination the org chart does not yet support. Low-rank work is sometimes prerequisite; teams discover the schema is broken while doing prompt tuning, not before. The fix is not to skip the bottom ranks. It is to notice when the rank where the work is being done is not the rank where the bug lives, and then move. The work at #6 and #5 and #3 ages well. A thoughtful permission model survives three model upgrades. A thoughtful prompt does not. ### The audit Three questions to run on your last week of agent-engineering work. What fraction of your hours went below rank #10? For most teams the honest answer is above 70 percent. The gravitational pull of the lower ranks is strong; the point of asking is to notice. Which question at ranks #6 through #3 have you been avoiding because it has no fast win? There is usually exactly one. Naming it is half the work. Which one rank up from where the team currently lives could you move to next sprint, made concrete enough to fit on a Friday demo? Pick that one. Do not apologise for the demo being smaller than the previous one. The model is rarely the binding layer. The next serious gain lives one rank up from where your team works now. *What rank does your team operate at today, and which rank do you most need to climb to next?* [^1]: Donella Meadows, *Leverage Points: Places to Intervene in a System* (1999-10-19). Archived by The Donella Meadows Project. https://donellameadows.org/archives/leverage-points-places-to-intervene-in-a-system/ [^2]: OpenAI, *Inside Our In-House Data Agent* (2026-01-29). Authors: Bonnie Xu, Aravind Suresh, Emma Tang of the OpenAI Data Productivity team. The three lessons (*Less is More* / *Guide the Goal, Not the Path* / *Meaning Lives in Code*) and the six-layer context architecture (schema, annotations, code, institutional, memory, runtime) are direct quotes from the post. https://openai.com/index/inside-our-in-house-data-agent/ --- --- title: "Three Hidden Bottlenecks the AI Buildout Has Already Moved Past GPUs" description: "NVIDIA's GPU shipments are not the binding constraint anymore. The supply chain has voted on what comes next." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/three-hidden-bottlenecks-past-gpus/" date: "2026-05-26" series: "MARKETS & POWER" law: "Law I" substack: "https://harryfloyd.substack.com/p/three-hidden-bottlenecks-past-gpus" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Three Hidden Bottlenecks the AI Buildout Has Already Moved Past GPUs *NVIDIA's GPU shipments are not the binding constraint anymore. The supply chain has voted on what comes next.* By Harry Floyd · 2026-05-26 · canonical: https://durabilitycurve.com/blog/three-hidden-bottlenecks-past-gpus/ [Figure: Pasted image 20260526175052] *This is an analytical framework, not financial advice. Numerical claims are referenced to their primary sources, or to current published coverage of them, in the footnotes.* Bloom Energy reported Q1 2026 revenue of $751 million. That number was 130 percent higher than the prior year, 42 percent above consensus, and triggered a full-year guidance raise to $3.6 billion[^1]. Most of the post-earnings coverage read the print as a fuel cell company finally turning operationally profitable. The print is not a fuel cell story. It is the canonical evidence that the AI infrastructure bottleneck has migrated past compute. For two years the consensus model for AI capex has anchored on GPU shipments. NVIDIA, AMD, the hyperscaler capex disclosures, the analyst models all priced compute as the load-bearing constraint. The reasoning was straightforward: training runs scaled, GPU clusters grew from 5,000 units to 50,000 to 100,000, and the company that supplied the silicon owned the bottleneck. The reasoning was correct in 2023. It became incomplete in 2024. By 2026 it has become a rear-view mirror. The analyst models that price AI on GPU shipments are not wrong about GPUs being important. They are wrong about GPUs being scarce. The supply-side data has been telling a different story for three quarters now, and Bloom Energy's print is the most recent confirmation. The bottleneck moved. It always does. *The binding constraint never disappears. It only migrates to the next layer.* The question that matters now is which layer the binding constraint has migrated to. Three layers have evidence pointing at them, none of which are GPUs, and the layers compose into a single observation about where AI capex goes once the compute layer has been solved. ### The first layer: power, and the 128-week wait [Figure: Pasted image 20260526175452] Behind every large GPU cluster sits a power-delivery infrastructure that takes longer to build than the cluster itself. Power transformers, the equipment that steps utility-scale voltage down to data-centre-usable voltage, have 80 to 128 week lead times right now[^2]. Cleveland-Cliffs is the only domestic US producer of the grain-oriented electrical steel that every transformer core requires[^3]. The grid interconnection queue at major US utilities runs five-plus years for new high-voltage data centre loads[^4]. This is the layer where Bloom Energy fits, and where the print becomes legible. Solid oxide fuel cells generate power on-site, behind the meter, without queueing for grid interconnection. A hyperscaler that wants 100 megawatts of power in eighteen months and cannot get it from the grid for five years buys Bloom Energy units. The fuel cell technology is twenty years old. The 130 percent revenue growth is the price of how binding the power constraint has become. > A hyperscaler that wants 100 megawatts in eighteen months and cannot get it from the grid for five years buys Bloom Energy units. The 130 percent revenue growth is the price of how binding the power constraint has become. > **For beginners: what does "behind the meter" mean?** A utility meter measures power coming into a building from the grid. *Behind the meter* means power generated on the customer's side of that meter, so the grid never sees it and never has to plan for it. Bloom Energy's fuel cells are behind-the-meter generation. That is why the eighteen-month installation timeline is the only one that matters for a hyperscaler who cannot wait five years for grid interconnection. The falsifier for the power layer is specific. If transformer lead times compress below 52 weeks within two consecutive quarters, or if hyperscaler 24/7 firm clean power purchase agreements (PPAs) at 15-year tenors are consistently signed below $80 per megawatt-hour, the constraint has eased and the behind-the-meter premium decays. Watch the second of those harder than the first. Hyperscalers will pay whatever the grid cannot deliver fast enough, and the PPA price is where that desperation gets numerical. ### The second layer: metal, and the recycling angle nobody priced [Figure: Pasted image 20260526180121] The compute layer requires copper. The power layer requires copper. The interconnect layer requires copper. By 2030, AI data centres alone will be calling on roughly 7 percent of all the copper the world digs up in a year, from a demand source that did not meaningfully exist five years ago. The math is straightforward. Hyperscale AI sites consume 40 to 50 tons of copper for every megawatt of IT capacity[^5]. The US has 85 gigawatts of new pipeline through 2030[^6], with 35 gigawatts already under construction across North America[^7]. Wood Mackenzie projects 1.1 million tonnes per year of grid copper demand from data centres alone[^8]; BloombergNEF projects another 572,000 tonnes peaking in 2028 inside the facilities themselves[^9]. Combined, that approaches 1.7 million tonnes per year against global mine output of roughly 23 million tonnes annually[^10]. One new demand source, 7 percent of every mine on earth, on top of every other demand the market already cannot meet. Mine capacity does not flex on the timescales the buildout requires. Copper mines take a decade from greenfield discovery to first commercial shipment. The buildout is happening on a one-to-three year horizon. There is no path where new mining capacity meets new data-centre demand. The consensus copper-AI thesis names the major miners: Freeport-McMoRan, Southern Copper, BHP, Rio Tinto. The miners are the obvious read. The recycling angle is the underfollowed one. Aurubis is a German specialty metals conglomerate that runs the largest secondary copper smelting capacity in Europe and is building the first US secondary smelter. Recycling output can flex on the timescales primary mining cannot. The structural shift is from *mining is the bottleneck* to *recycling is the relief valve*. The equity that captures the relief valve trades at approximately 0.4 times price-to-sales[^11]. The market reads Aurubis as a commodity cyclical. The multiple ignores the data-centre demand curve. > The structural shift is from mining is the bottleneck to recycling is the relief valve. The equity that captures the relief valve trades at roughly 0.4 times price-to-sales. The multiple ignores the data-centre demand curve. > This publication tracks where capital is migrating before the analyst models reprice it. **[Subscribe here.]** The falsifier for the metal layer is observable and time-bound. If primary copper-mine output growth exceeds 10 percent year-over-year for two consecutive years, the supply-shortage premium for recyclers compresses. The fallback test: if hyperscaler-driven data-centre permitting decelerates by more than 30 percent year-over-year, the demand assumption breaks before the supply assumption fires. Watch the permitting numbers monthly. The construction pipeline is the leading indicator of the copper demand curve. ### The third layer: detection, and the $151 billion question [Figure: Pasted image 20260526175803] The third layer is the most speculative of the three, and also the one where the supply side is voting hardest. The reason markets have not priced it yet is that the contract that creates it was only finalised in January 2026. SHIELD is the Scalable Homeland Innovative Enterprise Layered Defense vehicle: a $151 billion ten-year contract the Missile Defense Agency awarded as the primary acquisition framework for the broader Golden Dome missile-defence initiative[^12]. Golden Dome itself sits above SHIELD as the umbrella programme, with the Pentagon's own ten-year cost estimate at approximately $185 billion and the Congressional Budget Office's May 2026 analysis projecting up to $1.2 trillion over twenty years if a full space-based interceptor layer is built out[^13]. The MDA selected 2,440 firms as qualified SHIELD vendors across three tranches in late 2025 and early 2026[^12]. Holding a SHIELD position confers eligibility to compete for individual task orders, not guaranteed funding; task-order competitions are now beginning. The data layer of Golden Dome (the satellites and ground-segment processing that detect, classify, and track aerial threats) is a procurement category that did not meaningfully exist five years ago. Spire Global is a publicly-traded satellite-data company at roughly $700 million market cap[^14] with a remaining-performance-obligations backlog above $200 million, equivalent to about three times trailing twelve-month revenue[^15]. Their core revenue stream is Global Navigation Satellite System (GNSS) radio-occultation weather data, maritime Automatic Identification System (AIS) tracking, and radio frequency (RF) signal monitoring. Each of those data feeds is dual-use. The same instruments serve weather forecasting, shipping logistics, and defence persistent surveillance. If Spire captures even one percent of SHIELD contract dollars over the ten-year program, that is $1.5 billion in cumulative revenue against the current $200 million backlog. Multiples of trailing revenue visibility. The re-rating mechanism is one event. A SHIELD task-order announcement reclassifies Spire from data subscription business to defence contractor inside a single news cycle. The growth curve does not need to deliver first. > The re-rating mechanism is one event. A SHIELD subcontract announcement reclassifies Spire from data subscription business to defence contractor inside a single news cycle. The growth curve does not need to deliver first. The falsifier here is sharp. If SHIELD task orders are awarded across the next twelve months and none flow to Spire (if Lockheed Martin, Raytheon, Northrop Grumman absorb the data layer through their own subsidiary acquisitions), the thesis collapses to a $300 to $400 million data subscription company. The current $700 million cap depends on the SHIELD optionality being non-trivial. The probability is unknowable. The binary is well-defined. ### The composition the consensus misses These three layers are not independent positions. They are the three components of a single observation about where the AI capex chain has bound. Compute is solved at the marginal layer. NVIDIA, AMD, and the hyperscaler custom silicon teams have shipped enough capacity that the binding constraint sits elsewhere. The constraint is upstream and downstream of the chip: upstream because the power and metal that the cluster requires cannot be delivered on the cluster's timeline, and downstream because the strategic infrastructure that monitors and protects the data centres requires its own procurement category. The three layers compose because they share the same load curve. The same hyperscaler buildout that drives Bloom Energy's revenue growth drives the copper demand that Aurubis's recycling capacity will absorb. The same defence-procurement urgency that makes SHIELD a $151 billion program emerges from the same strategic environment that makes hyperscaler power-delivery a national-security concern. The three layers are not three separate trades. They are one observation, expressed three ways. Two structural moves come out of this analysis. The first: track where the supply chain is voting before the analyst models price it. The optical commitments NVIDIA made to Corning, Lumentum, Coherent, and Ayar Labs in late 2025 were the upstream signal that the photonics convergence was migrating into the compute interconnect layer. Bloom Energy's Q1 2026 print is the equivalent upstream signal for power. The supply-side numbers are the leading indicator. The second move: when the constraint binds, follow the difficulty. The hard part is what produces the value. Building a 128-week transformer is hard. Refining secondary copper to data-centre purity is hard. Carrying persistent space-based surveillance with the calibration and uptime SHIELD requires is hard. Each of these difficulties is what creates the moat for the equity that owns the relevant infrastructure. ### What to do with the framework Three watch items, each with a falsifier so the reader can run the framework themselves rather than wait for the publication to update them. For power: watch hyperscaler 24/7 firm clean PPAs at 15-year tenors. If those PPAs settle below $80 per megawatt-hour for two consecutive quarters, behind-the-meter generation premium is decaying and the bottleneck is moving back to grid-scale supply. For metal: watch primary copper-mine output growth. If it exceeds 10 percent year-over-year for two consecutive years, the supply-shortage premium for recyclers compresses and the Aurubis thesis weakens. For detection: watch SHIELD program subcontract announcements through Q4 2026. If the data layer awards go to primes without Spire in the supply chain, the thesis is falsified and Spire re-rates as a data subscription company. The whole-framework falsifier is broader. If GPU shipment growth re-accelerates above 100 percent year-over-year for two consecutive quarters AND power, metal, and detection multiples expand simultaneously, the compute layer is back as the binding constraint and the rest of this analysis is a temporary regime that has reverted. The publication will hold this open as a watched possibility, not as an active expectation. The expectation is that the binding constraint stays where the evidence currently places it. Power, metal, detection. Three layers, three companies, one observation about where the bottleneck has moved. The supply chain has already voted. The analyst models will catch up in two or three quarters. The reader who repositions before they do gets the asymmetric return. If this framework helps, it composes with two earlier pieces. [What Are You Actually Buying In The SpaceX IPO](https://durabilitycurve.com/blog/the-spacex-ipo-is-not-what-you-think/) makes the same move at a single layer: separating the operating reality of an extraordinary company from the seat public investors actually receive. [PLTR: The AI Stock That Has To Prove It Owns The Permission Layer](https://open.substack.com/pub/harryfloyd/p/pltr-the-ai-stock-that-has-to-prove) makes it again, at a different layer. Both are about the same discipline as the one this article asks for. Do not confuse the surface of a story with the structural position that produces value. Which falsifier would you watch first, and why? [^1]: Bloom Energy Q1 2026 SEC 8-K: revenue $751.1M (+130.4% YoY), FY26 guide raised to $3.4 to $3.8B. https://www.sec.gov/Archives/edgar/data/1664703/000162828026027913/ex991_q126financialresults.htm [^2]: Wood Mackenzie Q2 2025 industry survey: standard power transformers averaging 128 weeks lead time, generator step-up units 144 weeks, specialised orders out to four years. https://www.industrialsage.com/power-transformer-lead-times-us-grid-shortage/ and https://www.powermag.com/transformers-in-2026-shortage-scramble-or-self-inflicted-crisis/ [^3]: Cleveland-Cliffs Butler Works is the sole US producer of grain-oriented electrical steel for transformer cores. https://www.clevelandcliffs.com/operations/steel-mills [^4]: Lawrence Berkeley National Laboratory, "Queued Up: 2025 Edition": median interconnection-to-commercial-operation has doubled to over four years; ~10,300 active projects representing 1,400 GW generation and 890 GW storage as of end-2024. https://emp.lbl.gov/publications/queued-2025-edition-characteristics [^5]: S&P Global puts AI data-centre copper intensity at 30 to 47 tonnes per MW of IT capacity; JPMorgan industrial-metals coverage cites 47 tonnes per MW. https://skillings.net/copper-demand-ai-data-centers-vs-evs-the-2026-supply-shock-explained/ and https://skillings.net/copper-demand-ai-data-centers-2026-outlook-and-price-drivers/ [^6]: S&P Global, "Navigating the US data center power crunch": ~85 GW of new data-centre capacity pipeline by 2030 against current peak surplus generating capacity of ~70 GW. https://www.spglobal.com/en/research-insights/special-reports/look-forward/data-center-frontiers/navigating-us-data-center-energy-demand [^7]: JLL Global Data Center Outlook 2026: ~97 GW added globally between 2026 and 2030; ~35 GW under construction across North America. https://www.jll.com/content/dam/jllcom/en/global/documents/reports/research-reports/26-research-global-data-center-outlook-new.pdf [^8]: Wood Mackenzie, "High-wire act": data-centre grid copper demand reaches 1.1 Mt/yr by 2030. https://www.woodmac.com/horizons/soaring-copper-demand-obstacle-to-future-growth/ [^9]: BloombergNEF projects AI on-site copper demand averaging ~400 kt/yr over the next decade and peaking near 572 kt in 2028. The BNEF report is paywalled; figures quoted in https://carboncredits.com/data-centers-copper-hunger-how-ai-is-driving-a-looming-supply-crunch/ and https://globaltacticalmetals.com/ai-data-centers-to-worsen-copper-shortage-bnef/ [^10]: ICSG World Copper Factbook 2025: 2024 global mine production ~23 Mt; refined production 27.5 Mt, of which 4.7 Mt secondary. https://icsg.org/download/2025-10-the-world-copper-factbook/ [^11]: Aurubis AG (XETRA: NDA) ~0.4x trailing price-to-sales as of May 2026. https://www.morningstar.com/stocks/xetr/nda/valuation [^12]: MDA SHIELD IDIQ: $151B shared ceiling over ten years; 2,440 qualified vendors across three tranches (1,014 on 2 Dec 2025, 1,086 on 18 Dec 2025, 340 on 15 Jan 2026). https://www.defenseone.com/business/2025/12/gargantuan-golden-dome-contract-vehicle-clears-1000-plus-firms-vie-slices-151-billion/409900/ and https://dsm.forecastinternational.com/2026/01/16/pentagon-mobilizes-industrial-base-for-golden-dome-missile-shield-with-151b-shield-award/ [^13]: Pentagon ten-year Golden Dome estimate ~$185B (Gen. Michael Guetlein, April 2026 testimony). CBO May 2026 analysis: up to $1.2T over twenty years with a full space-based interceptor layer (~70% of acquisition cost); ~$448B without it. https://spacenews.com/congressional-budget-office-estimates-1-2-trillion-price-tag-for-golden-dome/ and https://www.airandspaceforces.com/pentagon-cbo-trillion-dollar-golden-dome-estimate/ [^14]: Spire Global (NYSE: SPIR) market cap ~$700M as of May 2026. https://companiesmarketcap.com/spire-global/marketcap/ [^15]: Spire Global Q3 2025: remaining-performance-obligations $223.1M as of 30 Sept 2025 (over 3x TTM revenue); ~$70M expected to convert in 2026. https://ir.spire.com/news-events/press-releases/detail/279/spire-global-announces-third-quarter-2025-results --- --- title: "NVDA Q1 FY2027: The Networking Number That Changes the Story" description: "NVIDIA Q1 FY2027 revenue hit $81.6B (+85% YoY) — but the real story is networking revenue surging 199% as the AI bottleneck migrates from GPUs to interconnects." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/nvda-q1-fy2027-the-networking-number-that-changes-the-story/" date: "2026-05-21" series: "MARKETS & POWER" law: "Law I" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # NVDA Q1 FY2027: The Networking Number That Changes the Story *NVIDIA Q1 FY2027 revenue hit $81.6B (+85% YoY) — but the real story is networking revenue surging 199% as the AI bottleneck migrates from GPUs to interconnects.* By Harry Floyd · 2026-05-21 · canonical: https://durabilitycurve.com/blog/nvda-q1-fy2027-the-networking-number-that-changes-the-story/ NVIDIA reported its fiscal first-quarter 2027 results on May 20, 2026. Revenue of $81.62 billion beat the $79.2 billion consensus by 3%. Earnings per share of $1.87 beat estimates by 6%. The Q2 guide of $91 billion exceeded the $87.3 billion consensus by 4%. By every conventional measure, this was a clean beat-and-raise quarter. The stock closed at $215.22, up 1.8%, and was flat after hours. That is the fifth time in six quarters that NVIDIA has beaten expectations and seen the stock fail to rally. The pattern reveals something structural: at this scale, the headline numbers are priced before the print. The signal is in the *composition* of the revenue, not the total. ## The Number That Changes the Narrative Data Center networking revenue reached **$14.8 billion** — a record, up **199%** year-over-year and 35% sequentially. Compare that to Data Center compute revenue of $60.4 billion, which grew 77% year-over-year. The networking segment is growing at **2.6 times** the rate of the compute segment. This is **Law I (Bottleneck Migration)** expressed in a single quarter of financial data. As GPU clusters scale past 50,000 devices, the wall-clock binding constraint on AI training shifts from FLOPs to inter-GPU bandwidth. The network layer becomes the scarce instrument. NVIDIA's networking business — now larger than AMD's total revenue — is capturing the value of that migration. Two years ago, networking was roughly 12% of Data Center revenue. It is now 20% and accelerating. The bottleneck is moving, and the instruments that express the new constraint — optical interconnects, networking silicon, Spectrum-X Ethernet fabric — are growing revenue faster than the GPUs they connect. ## Why Gross Margins Expanded During a Volume Ramp GAAP gross margin reached **74.9%** — up from 60.6% a year ago. This directly contradicts the commoditisation thesis that volume production of Blackwell would compress margins as CoWoS packaging and HBM memory costs rose. Margins expanded because NVIDIA's full-stack moat (CUDA + NVLink + Spectrum-X + Blackwell silicon) creates pricing power that chip-design-alone cannot produce. This is **Law II (Difficulty Is Load-Bearing)**. The difficulty of replicating the stack is the barrier that protects the margin structure. No competitor currently achieves this combination of scale and margin. ## The Cash Engine Is Fully Online NVIDIA generated **$48.6 billion** in free cash flow in a single quarter — a 60% FCF margin. To put that in perspective, only about 15 companies in the world generate more net profit in an entire year than NVIDIA generates in cash in three months. Capital returns signal management's confidence: the quarterly dividend went from $0.01 to $0.25 per share (a 25x increase), the board authorized a new $80 billion share buyback, and the company returned approximately $20 billion to shareholders during the quarter itself. ## Vera Rubin Is on Schedule NVIDIA confirmed that Vera Rubin, the next-generation architecture, is **on track for the second half of 2026, starting in Q3 with volume ramp in Q4**. Architectural transitions are the highest-risk moments for any semiconductor company. Intel's 10nm stumble, AMD's 7nm delay — the canonical failures all happen at the generational handoff. NVIDIA navigating this transition without a demand gap is the most important operational question for the next twelve months, and this quarter's confirmation de-risks it substantially. Blackwell demand remains so strong that it is *driving up secondary-market prices* for older Hopper and Ampere GPUs. This is not a demand cliff narrative. This is a demand acceleration narrative with a clean architectural handoff. ## The One Risk (Law B — Regime Problem) Data Center now accounts for **92%** of NVIDIA's total revenue. This is not a company risk — it is a regime risk. If hyperscaler capital expenditure cycles (Microsoft, Meta, Google, Amazon all expanding today), 92% of revenue faces the same headwind simultaneously. NVIDIA's diversification into automotive, robotics, and enterprise AI is real but collectively represents approximately 8% of revenue. The stock market is pricing this risk more heavily than headline numbers suggest. That is why $81.6 billion in revenue and a $91 billion guide produced a 1.8% stock move. The market sees the concentration. It is asking: how long can this last? ## Falsification Triggers — All Green For anyone tracking the thesis structurally, here are the specific thresholds that would change the view: **Q2 guide below $85B** — Guided $91B. Not close. **Vera Rubin delayed beyond Q3** — Confirmed on track for Q3. **Gross margin below 73%** — Currently 74.9% and guided 75% for Q2. **Networking growth < compute growth** — Networking 199% vs compute 77%. The opposite. **Hyperscaler ASIC share >15%** — Still in single digits. **Export controls expand to allied nations** — Status quo, China-only. Every falsification trigger remains green. The thesis is intact and the data is strengthening it. ## What This Means Through the Durability Curve This quarter confirms the vault's two core predictions for NVIDIA. The bottleneck is migrating from compute to interconnect — the networking growth rate is the proof. And the margin structure is holding because the difficulty barrier is real. The open question is not about execution. NVIDIA is executing flawlessly. The open question is about the regime: how long before the hyperscaler capex cycle turns, and whether NVIDIA can build the 8% non-DC revenue into something material enough to absorb a rotation. For now, the data says: the bottleneck is moving, the moat is holding, and Vera Rubin is on schedule. That is a thesis-strengthening quarter. --- *Subscribe to The Durability Curve on [Substack](https://harryfloyd.substack.com/?utm_source=telegraph) — free weekly AI infrastructure analysis.* *Full research reports on [Gumroad](https://harryfloyd.gumroad.com/?utm_source=telegraph) — deep-dive structural analysis for investors tracking the AI supply chain.* --- *Originally published on [Telegraph](https://telegra.ph/NVDA-Q1-FY2027-The-Networking-Number-That-Changes-the-Story-05-21).* --- --- title: "Access Is Not Agency" description: "Access Is Not Agency" author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/access-is-not-agency/" date: "2026-05-19" series: "THE HUMAN LAYER" substack: "https://harryfloyd.substack.com/p/access-is-not-agency" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Access Is Not Agency *Access Is Not Agency* By Harry Floyd · 2026-05-19 · canonical: https://durabilitycurve.com/blog/access-is-not-agency/ *The tool is not the authority.* *Most teams are giving agents more connectors before they have defined what the agent is allowed to change. That is not agency. That is reach with a larger blast radius.* ## The agent everyone calls powerful Imagine the demo. The agent can read Slack. It can search email. It can query the CRM. It can open GitHub issues, check the billing system, browse docs, edit a spreadsheet, draft a customer reply, and call three internal APIs. Everyone in the room calls it powerful. That is the first mistake. The question is not what the agent can access. The question is what it is allowed to change. Can it send the email, or only draft it? Can it update the customer's plan, or only propose the update? Can it refund the invoice, revoke a token, merge the pull request, notify the vendor, reassign the ticket, delete the record, file the report, or trigger the incident workflow? Once an agent can act through tools, the real system is no longer the model. The real system is the action contract around the model. > Access is reach. Agency is permissioned action under constraints. This distinction sounds small until the first bad run. A read-only research assistant can waste time. An agent with billing access can create obligations. An agent with email access can speak for the company. An agent with deployment access can turn a wrong inference into infrastructure. More tools do not automatically make the agent more agentic. More tools expand the surface on which judgment must be engineered. ## A tool call is not agency A confidence score is not evidence. A tool call is not agency. Tool access tells you what an agent can touch. It does not tell you what the agent is authorised to decide, what must be checked, what becomes binding, or what happens after failure. That is why the current agent conversation feels slightly wrong. People talk as if the next leap is connectors. Give the model Slack, Gmail, GitHub, Linear, Notion, Salesforce, Stripe, a browser, memory, and MCP servers, then wait for autonomy to emerge. But a connector is not a decision right. A connector gives the agent a door. It does not define whether the agent may walk through the door, what it may carry, who must inspect the bag, whether the door locks behind it, or who reviews the camera footage if something goes missing. > A connector is a door. Agency is a contract about who may walk through it. This is not a metaphorical governance concern. It is the operating surface. Security lawyers are already asking the practical version of the same question. When an agent can send emails, modify records, execute transactions, or orchestrate other systems, deployment stops looking like ordinary software access and starts looking like delegated operational authority. The useful questions become blunt: what authority does the agent have, can it read only or modify systems, which actions require human approval, how is behaviour audited, and how do you stop it if it becomes a threat?[^1] That is the missing layer in most agent demos. They show reach. They do not show authority. ## The stack nobody wants to name If you want to know whether an agent has real agency, do not start with the model card. Start with the rights stack. What can it see? What can it change? What must it prove before the change becomes real? What triggers escalation? What permission disappears after a bad run? That last question matters most because it reveals whether the system has a memory of failure. A human employee loses trust after a bad judgment. They may lose budget authority, approval rights, admin permissions, or the ability to act without supervision. Most agents do not lose anything. They fail, get patched, and return with the same action surface. That is not delegation. That is amnesia with API keys. An agent action stack has at least five layers. The first is visibility: which data, tools, documents, messages, logs, tickets, accounts, and systems the agent can inspect. The second is mutation: which objects the agent can change. Reading a customer record and changing a customer record are different powers. Drafting a reply and sending a reply are different powers. Proposing a deployment and executing a deployment are different powers. The third is proof: what the agent must produce before a mutation becomes real. That could be a test run, a diff, a trace, a policy check, a second-model review, a human approval, a simulated dry run, or an evidence bundle. The fourth is escalation: when the agent must stop and hand the decision to someone else. Not "human in the loop" as a slogan. A named escalation condition. Missing context. High reversibility cost. Conflicting instructions. External communication. Payment movement. Privilege change. Legal exposure. The fifth is revocation: what changes after the agent fails. If a bad run does not shrink future permissions, the system has no operational immune response. [Figure: The Rights Stack] > The hard part is not giving the agent a tool. The hard part is deciding when the tool stops being available. This is why "least privilege" becomes more important in agent systems, not less. A normal app executes known code paths. An agent chooses a path through a tool surface at runtime. The permission is no longer just "can this service account call this API?" The question becomes "for this task, with this evidence, under these constraints, should this agent be allowed to perform this action now?" That is a different shape of access control. ## The bottleneck moved from capability to authority You can see the migration in deployed behaviour. Anthropic's research on agent autonomy is useful because it does not only ask what models can theoretically do. It looks at real product behaviour. In Claude Code sessions, the long tail of unsupervised turns got longer between late 2025 and early 2026. More interestingly, experienced users both auto-approve more and interrupt more. They do not simply trust the agent blindly. They shift from approving every step to letting longer runs proceed, then intervening when timing, uncertainty, or risk demands it.[^2] That is what practical autonomy looks like. Not zero oversight. Selective oversight. The better the agent gets, the less useful per-step approval becomes. But that does not make approval disappear. It moves approval upward. > The better the agent, the higher the approval rises. Selective oversight beats per-step oversight only when the boundaries are named. Instead of "may the agent call this tool?" the important question becomes "which category of action is this, under which authority, with which proof, and what happens if the run crosses a boundary?" This is why a more capable agent can make a weak control stack worse. It will move faster through a larger action surface. It will chain tools. It will recover from errors. It will confidently produce intermediate artefacts that look plausible. It will make the system feel smoother right up to the point where the wrong action becomes real. The bottleneck has moved. It is no longer only model capability. It is decision-rights routing. Who gets to decide what? Under which conditions? With what evidence? With what right to commit the change? With what right to continue after failure? The organisations that answer those questions will get more agency from smaller tool surfaces than the organisations that connect everything and call it autonomy. ## Older institutions already know this Companies do not give humans "access" and call the job designed. A junior analyst can see a model. They may not approve a trade. A support rep can view a customer record. They may not issue a large refund without approval. An engineer can open a pull request. They may not deploy to production alone. A finance employee can prepare a payment. They may not release it without a second sign-off. Corporate delegation is an action-rights system. So are IAM, accounting controls, and clinical protocols. They all separate seeing, recommending, approving, executing, logging, and reviewing. > A control system without revocation is not a control system. It is trust with a longer leash. Agents are forcing software teams to rediscover that distinction inside product architecture. The strongest academic frame I found for this comes from a 2026 paper on Autonomous Administrative Intelligence. The paper is conceptual, not empirical, so it should not be treated as proof that the architecture works. But its structure is exactly the one agent teams need to notice: strategic control, agentic decision formation, and governance validation are separate layers. The agent may form an administrative decision. A governance layer validates it against rules and constraints. Execution and recording happen only after that validation. Humans shift from constant supervision toward intent, boundaries, and exceptions.[^3] That is the right shape. Do not ask whether the agent can complete the task. Ask where decision formation ends and validation begins. If those are the same place, the agent is not operating under an action contract. It is operating under trust. Trust is not bad. Trust without revocation is not a control system. ## Tool design is contract design This is where tool protocols start to matter, but not for the reason most people say. The industry is building standard ways for agents to connect to external systems. The most prominent is MCP, Anthropic's Model Context Protocol. The lazy version of that story says MCP is important because it gives agents more tools. The better version says it is important because it makes the tool boundary explicit enough to inspect, version, test, authorise, and debug. That distinction is not just engineering taste. It changes what you can see when things go wrong. Anthropic's own engineering guidance frames tools as contracts between deterministic systems and nondeterministic agents. Their names, descriptions, return values, and failure modes affect whether an agent can use them reliably. A tool built for a human developer is not automatically a good tool for an agent.[^4] Once you accept that, the "more connectors" story becomes incomplete. Tool count is not the win. Contract quality is the win. > Tool count is not the win. Contract quality is the win. A strong tool contract tells the agent what the tool does, what inputs it needs, what output means, what failure looks like, and what the agent should not infer. A strong action contract goes further. It says when the agent may call the tool, when the call may mutate something, what proof is needed, and where the trace goes. Early benchmarks are confirming this. When researchers tested agents across hundreds of real tools and multi-step tasks, the failures were not mainly in the final answer. Agents chose the wrong tool, passed wrong parameters, recovered badly from errors, or produced plausible answers from broken execution traces.[^5] The system looked like it worked. The trace showed that it did not. That is the pattern to watch. Tool-using agents need diagnostics at the contract layer, not only at the output layer. If your agent can touch ten systems and your only observable is "task succeeded," you are blind in the place where the system is becoming dangerous. ## The Agent Action Rights Test Run this on the most powerful agent or workflow you currently use. Do not pick a toy. Pick the one you are most tempted to trust. The coding agent with repo access. The sales assistant with CRM access. The ops agent with incident tooling. The analyst agent with warehouse access. The support agent with customer email access. Then answer five questions without hand-waving. > What can the agent see? > > What can the agent change? > > What must it prove before the change becomes real? > > What triggers escalation to a human? > > What permission disappears after a bad run? Most teams can answer the first question. Some can answer the second. Almost no teams can answer all five. That is the diagnostic. If you cannot answer "what can it see?", you do not have an inventory. If you cannot answer "what can it change?", you do not have a permission model. If you cannot answer "what must it prove?", you do not have verification. If you cannot answer "what triggers escalation?", you do not have oversight. If you cannot answer "what permission disappears?", you do not have learning at the authority layer. [Figure: Revocation After Failure] You may still have a useful agent. You may even have a high-performing one. But you do not yet have trustworthy agency. You have a tool-using system whose action rights are partly implicit. Implicit action rights always become visible after an incident. ## The dangerous middle There is a tempting objection here. If we make every action permissioned, verified, escalated, logged, and revocable, will we not kill the point of agents? Yes, if you do it badly. The goal is not to turn every agent into a form-filling intern that asks permission before breathing. The goal is to match authority to consequence. Read-only Slack search should be cheap. Drafting a customer reply should be cheap. Local reversible edits should be cheaper than external irreversible commitments. Sending the customer email, refunding the invoice, revoking the token, merging the pull request, or releasing the payment should pass through stronger gates. Good action rights are not one wall around the whole system. They are a slope. > Good action rights are a slope, not a wall. Authority should track consequence. The agent gets wider freedom where mistakes are cheap, visible, and reversible. It gets narrower freedom where mistakes are expensive, silent, and hard to undo. [Figure: Authority Is a Slope] That is also how good human organisations work. The graduate can model scenarios. The manager can approve a small budget. The director can reallocate headcount. The board can approve the acquisition. Authority changes with consequence. Agents need the same gradient. The mistake is treating "human approval" as the only safety primitive. Approval is expensive. It also fails when humans approve too much, too fast, or without the evidence needed to judge. The better primitive is action-rights design: proof requirements, escalation rules, default-deny mutations, bounded autonomy, time-limited permissions, separate approval and execution, traces that survive, and permissions that shrink after failure. That stack is harder than adding another connector. That is why it will matter. ## What to build next If you are building or buying agent systems, ask for the action contract before the roadmap. Ask the vendor to show the permission tiers, not only the integration list. Ask which actions are read-only, draft-only, approval-gated, automatically executable, or forbidden. Ask where traces live. Ask how tool calls are mapped to business authority. Ask what happens when the agent is tricked, confused, stale, incomplete, or too confident. Ask how a bad run changes the next run. This is where the serious agent market will split. One side will sell reach: more connectors, more memory, more tools, more environments, more background work. The other side will sell agency: permissioned action, bounded autonomy, proof before commitment, escalation when context breaks, and revocation when trust is lost. Reach will demo better. Agency will survive contact with the organisation. > Reach demos well in the room. Agency holds up in the incident review. Before next week, run the Agent Action Rights Test on one workflow. Not your whole stack. One agent. One workflow. Write the five answers in a note. If the fifth answer is blank, you found the missing layer. The agent did not need another tool. It needed a smaller right to be wrong. --- *If you run the test, which answer was hardest to fill: what it can see, what it can change, what it must prove, when it escalates, or what permission disappears?* ## If this piece landed This article builds on the claim that [a confidence score is not evidence](https://durabilitycurve.com/blog/most-verification-is-just-bigger/). If the rights stack resonated, the deeper version is [Your AI Agent Stack Is Solving The Wrong Problem](https://durabilitycurve.com/blog/your-ai-agent-stack-is-solving-the/), which walks through the full contract stack for agent systems. And if you want the five laws that run underneath all of it, start with [The Five Laws of Durable Systems](https://durabilitycurve.com/blog/the-five-laws-of-durable-systems/). **[Subscribe to The Durability Curve](https://harryfloyd.substack.com/)** for the next piece in this sequence. [^1]: Stoel Rives, *Securing and Contracting Agentic AI* (20 February 2026), [https://www.stoel.com/insights/publications/securing-and-contracting-agentic-ai](https://www.stoel.com/insights/publications/securing-and-contracting-agentic-ai). The page is especially useful because it moves quickly from generic "agentic AI" language to concrete authority, IAM, audit, monitoring, integration, and shutdown questions. [^2]: Anthropic, *Measuring AI Agent Autonomy in Practice* (18 February 2026), [https://www.anthropic.com/research/measuring-agent-autonomy](https://www.anthropic.com/research/measuring-agent-autonomy). Treat the metrics as first-party vendor telemetry, not independent industry prevalence. The useful point here is the oversight pattern: longer autonomous runs coexist with experienced users interrupting more selectively. [^3]: Aravindh Sekar, *Autonomous Administrative Intelligence: Governing AI-Mediated Administration in Decentralized Organizations*, *Administrative Sciences* 16(2), 95 (12 February 2026), DOI [10.3390/admsci16020095](https://doi.org/10.3390/admsci16020095). The article is theory-building, not empirical validation, but its separation of strategic control, agentic decision formation, validation, execution, recording, and exception governance is the right architectural distinction for this argument. [^4]: Anthropic, *Writing Effective Tools for AI Agents* (2025), [https://www.anthropic.com/engineering/writing-tools-for-agents](https://www.anthropic.com/engineering/writing-tools-for-agents). Anthropic frames tool definitions as contracts between deterministic systems and nondeterministic agents, and emphasizes realistic multi-step tool evaluations rather than single-call demos. [^5]: Bandi et al., *MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers* (arXiv 2602.00933, January 2026), [https://arxiv.org/pdf/2602.00933](https://arxiv.org/pdf/2602.00933). The benchmark reports 36 MCP servers, 220 tools, and 1,000 tasks with multi-tool diagnostics, which is the relevant point here: tool-using agents fail at discovery, invocation, sequencing, and recovery, not only final-answer wording. --- --- title: "Your AI Agent Stack Is Solving The Wrong Problem" description: "The real setup is not MCP servers, skills, memory files, and subagents. It is the contract stack that decides what an agent may know, do, prove, escalate, and lose." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/your-ai-agent-stack-is-solving-the/" date: "2026-05-09" series: "SYSTEMS & LAWS" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Your AI Agent Stack Is Solving The Wrong Problem *The real setup is not MCP servers, skills, memory files, and subagents. It is the contract stack that decides what an agent may know, do, prove, escalate, and lose.* By Harry Floyd · 2026-05-09 · canonical: https://durabilitycurve.com/blog/your-ai-agent-stack-is-solving-the/ ## The setup everyone is sharing Which MCP servers to install. Which skills to keep in your repo. Which agent framework to use. How to write your `AGENTS.md`. How to split one agent into researcher, planner, coder, and reviewer. How to wire Slack, GitHub, Notion, Postgres, Stripe, your calendar, and your file system into one increasingly capable loop. Some of that advice is useful. It is also aimed at the wrong layer. What becomes real after the agent uses a tool matters more than whether it can reach the tool. Can it read the customer record, or change it? Can it draft the refund, or issue it? Can it open a pull request, or merge it? Can it propose the vendor response, or send it under the company name? Once an agent can act through tools, the real system is no longer the model. The real system is the contract stack around the model. That is the part most setup guides skip. ## Access is reach. Agency is permissioned action. Imagine the demo. The agent can read Slack. It can search email. It can query the CRM. It can open GitHub issues, check billing records, browse docs, edit a spreadsheet, draft a customer reply, and call three internal APIs. Everyone in the room calls it powerful. That is the first mistake. The agent has reach. It does not yet have governed agency. Access tells you what the agent can touch. Agency tells you what the agent is authorised to decide, under which conditions, with what proof, and with what consequence after failure. That distinction sounds small until the first bad run. A read-only research assistant can waste time. An agent with billing access can create obligations. An agent with email access can speak for the company. An agent with deployment access can turn a wrong inference into infrastructure. More tools do not automatically make the agent more agentic. More tools expand the surface on which judgement has to be engineered. ## The tool stack is visible. The contract stack is load-bearing. The visible agent stack is easy to list: model, prompt, memory, tools, MCP servers, subagents, framework, evals. That stack matters. It is also not the operating system. The operating system is the set of contracts each layer creates. What is the agent for? What state may it see? What state may it preserve? Which tools may it call? Which tools are intentionally absent? What can it change? What must it prove before the change becomes binding? What does the harness log? What does the evaluation score actually cover? When does the agent ask, abstain, or escalate? What permission disappears after a bad run? That is the real setup. Not the list of tools. The set of boundaries that decides what the tools mean. [Figure] The generic setup stack asks what you connected. The contract stack asks what you can trust. ## An agent is a control loop, not a prompt with ambition An agent is an outer control loop wrapped around a generator. It plans, reads state, chooses tools, acts, observes, repairs, escalates, and decides whether to continue. The failure rarely sits in one glamorous place. It can sit in the planner. It can sit in retrieval. It can sit in a tool description. It can sit in retry logic. It can sit in a hidden assumption about whether the world waits while the agent thinks. That is why framework comparisons are often less useful than they look. The distinction that matters is which parts of the loop are explicit enough to inspect. If planning is hidden inside one long natural-language instruction, you cannot repair planning without rewriting the whole prompt. If memory is just a growing transcript, you cannot tell whether the agent remembered, retrieved, inferred, or hallucinated. If tool choice is unlogged, you cannot tell whether the answer is wrong because the model reasoned badly or because it called the wrong thing. If evaluation is one final pass/fail number, you cannot tell whether the agent failed at discovery, parameters, sequencing, recovery, escalation, or judgement. Agents do not become reliable when the setup becomes more impressive. They become reliable when failure has somewhere specific to land. ## MCP is not magic glue MCP matters. Skills matter. Connectors matter. But their importance is often described backwards. The lazy version says MCP is valuable because it gives agents more tools. The better version says MCP is valuable because it makes the tool boundary explicit enough to inspect, version, test, authorise, and debug. A tool is not neutral plumbing. A tool description tells a nondeterministic system what an action means. The name, parameters, return shape, error messages, and allowed mutations all change behaviour. A tool built for a human developer is not automatically a good tool for an agent. Humans carry missing context. Agents need the contract written down. That is why more tools can make an agent worse. At small scale, tool access feels like freedom. At larger scale, tool access becomes search. The agent has to identify the right tool, pass valid parameters, recover from partial failure, and avoid inventing a successful trace when the tool call failed. If you expose every API endpoint as a tool, you do not have a powerful agent surface. You have a vocabulary problem with write access. The mature move is not “connect everything.” The mature move is to design the smallest tool surface that lets the agent do the job, then make every tool contract legible. What does the tool do. When should it be used. What the return value proves, and what it does not prove. What failures look like. Which calls are read-only, which mutate state, which require approval. Where the trace goes. That is how you stop a transcript from becoming the only place your operating system exists. ## Skills are not prompt snippets The same mistake happens with skills. People treat skills as better prompts: a `SKILL.md`, a few examples, some instructions, maybe a script. Useful. Portable. Easy to share. But a serious skill is not a prompt snippet. It is packaged operating knowledge. It should contain a trigger, a procedure, a boundary, gotchas, and a failure mode. The “gotchas” are usually the most valuable part. The model often already knows the happy path. What it does not know is your local scar tissue: which API lies, which file must not be edited, which naming convention breaks deployment, which customer segment changes the policy. That is why generic skill catalogues have a ceiling. They can teach a model the common workflow. They cannot teach it which parts of your workflow are load-bearing unless you package that knowledge yourself. Skills are valuable because they let operational knowledge travel across sessions and agents. They are dangerous when they activate at the wrong time, compose implicitly into deeper graphs nobody intended, or grant state-changing behaviour without a permission contract. The real question is sharper: > When this skill activates, what decision is it allowed to influence? If nobody can answer that, the skill is just a more durable way to make the wrong move. ## Memory is governed state, not a bigger past Memory has the same problem. Every agent product wants to promise memory. It sounds obvious. The agent should remember the user, the project, the codebase, the customer history, the prior decision, the mistake from last time. But memory is not “more context.” Memory is a four-part contract: what gets written, how it is organised, how it is retrieved, how it is governed. If the agent writes too much, memory becomes sludge. If it summarises badly, memory becomes distortion. If it retrieves by similarity alone, memory becomes vibes with citations. If it never forgets, memory becomes context poisoning. If it cannot show why a memory was used, memory becomes an invisible authority. The memory question worth asking: > Which state should survive because it will improve future decisions, and which state should expire because it will poison them? That is a contract question. It is also why a 500-word, well-maintained project note can outperform a giant chat history. The smaller note has a job. The transcript merely has volume. [Figure] _The real setup is the contract stack that decides what an agent may know, do, prove, escalate, and lose._ ## The harness is where autonomy becomes measurable Most agent demos make the model look like the protagonist. In production, the harness is the protagonist. The harness is everything that surrounds the weights: task boundaries, tools, retry budgets, permission gates, stop rules, and evidence artefacts. Change the harness and the same model can look like a different system. That should make us suspicious of agent benchmarks that treat the model as the only object being compared. A published score does not measure a disembodied model. It measures a deployment regime, scaffold, metric, and judge. Was the world static or changing? Did the agent see a screenshot, HTML, an accessibility tree, a database row, or a curated prompt? How many retries did it get? Did it have tools? Which ones? Was the grader human, model-based, rubric-based, trajectory-aware, or outcome-only? Did the metric reward one lucky success or repeated consistency? These are not footnotes. They are the contract. If your eval sits two regimes below deployment, treat it as lab evidence. If it tests read-only draft behaviour, do not use it to justify automatic writes. If it rewards pass@k, do not pretend it proves worst-run reliability. If it grades only final answers, do not pretend it inspected tool behaviour. If it hides traces, do not pretend it supports auditability. The score is not the contract. The score is one output of a contract you have to name. ## The dangerous middle There is a tempting objection here. If every action needs a contract, won’t we kill the point of agents? Yes, if we do it badly. The answer is not to wrap every agent in a permission wall so thick it becomes useless. The answer is to match authority to consequence. Read-only search should be cheap. Drafting should be cheap. Local reversible edits should be cheaper than external irreversible commitments. Actions that affect money, identity, infrastructure, or customer communication should pass through stronger gates. Good agent authority is not one wall. It is a slope. The agent gets wider freedom where mistakes are cheap, visible, and reversible. It gets narrower freedom where mistakes are expensive, silent, and hard to undo. Older institutions already know this. A junior analyst can see a model. They cannot approve the trade. A support rep can view a customer record. They cannot issue a large refund without approval. An engineer can open a pull request. They cannot deploy to production alone. A finance employee can prepare a payment. They cannot release it without a second sign-off. Organisations separate seeing, recommending, approving, executing, logging, and reviewing because authority is not a binary. Agents force software teams to rediscover that inside product architecture. ## The missing layer is revocation Most agent setups have a permission story. Few have a revocation story. That is the giveaway. A human loses trust after a bad judgement. They may lose budget authority, approval rights, admin permissions, unsupervised access, or the ability to act without review. Most agents fail, get patched, and return with the same action surface. That is not learning. That is amnesia with API keys. If a bad run does not shrink future permissions, the system has no operational immune response. Revocation does not have to be dramatic. After one unsafe draft, require review for that category. After one wrong tool call, remove that tool until the contract is fixed. After one stale-memory error, force a memory review before reuse. After one hallucinated trace, require deterministic evidence for the next run. After one escalation miss, lower the threshold for asking a human. That is the difference between an agent that is merely corrected and an agent system that becomes safer. The hard part is not giving the agent a tool. The hard part is deciding when the tool stops being available. ## The Agent Contract Stack Audit Run this on one agent workflow you are tempted to trust. Not the whole company. Not your entire AI strategy. One agent. One workflow. Write the answers down. ``` 1. PURPOSE What decision or workflow is this agent meant to govern? 2. CONTEXT What state can it see, and what state must persist? 3. TOOLS What can it call, and which tools are intentionally absent? 4. AUTHORITY What can it change without approval? 5. PROOF What evidence must it produce before action? 6. HARNESS What retries, budgets, logs, and checks surround it? 7. EVALUATION What regime does the score actually cover? 8. HANDOFF When does it ask, abstain, or escalate? 9. REVOCATION What permission disappears after a bad run? ``` Most teams can answer tools. Some can answer authority. Few can answer proof, harness, evaluation regime, handoff, and revocation in the same breath. That is the diagnostic. If you cannot answer purpose, you have a demo. If you cannot answer context, you have hidden state. If you cannot answer tools, you have inventory risk. If you cannot answer authority, you have implicit delegation. If you cannot answer proof, you have output without evidence. If you cannot answer harness, you have unreproducible behaviour. If you cannot answer evaluation, you have a score without a world. If you cannot answer handoff, you have autonomy without judgement. If you cannot answer revocation, you have no way for failure to change the system. You may still have a useful agent. You do not yet have a trustworthy one. ## What to build instead Do not start by asking which agent framework to use. Start with one sentence: > We are evaluating whether this agent can perform this action in this environment under this permission boundary, and the decision governed by the result is this. That sentence does more work than a diagram with six logos. Then build the contract stack around it. Give the agent the smallest tool surface that can do the job. Package skills for local gotchas, not generic inspiration. Write memory only when future decisions should depend on it. Keep the harness visible. Evaluate in the regime you plan to deploy. Make traces inspectable. Define escalation before the agent is confused. Define revocation before the agent fails. This is less exciting than another setup guide. It is also the part that will decide who can actually use agents. The serious agent market will split into two groups. One side will sell reach: more connectors, more tools, more memory, more impressive demos. The other side will sell agency: permissioned action, bounded autonomy, proof before commitment, escalation when context breaks, traces that survive, and revocation when trust is lost. Reach will demo better. Agency will survive contact with the organisation. ## Field Card [Figure] --- --- title: "The Engine Underneath Hard Decisions" description: "Eight stages turn hidden structure into durable knowledge. Most teams run three of them and call it understanding. The other five are where compounding hides." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-engine-underneath-hard-decisions/" date: "2026-05-07" series: "SYSTEMS & LAWS" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Engine Underneath Hard Decisions *Eight stages turn hidden structure into durable knowledge. Most teams run three of them and call it understanding. The other five are where compounding hides.* By Harry Floyd · 2026-05-07 · canonical: https://durabilitycurve.com/blog/the-engine-underneath-hard-decisions/ A pricing team notices conversion has dropped on a major channel. The dashboard is clear. They retrain the pricing model with the latest week of data. Conversion drops further. They retrain again. Worse. Three weeks later someone discovers an upstream feed had silently changed format. The data had been lying about what it represented. The dashboard had been right about something. The team had asked it the wrong question. This is a story about a missing stage in the way the team produces knowledge from the world. There are eight stages between “a metric moved” and “the right intervention.” Most teams run three. ## The cycle, not the line Most accounts of how teams learn read like a list. Detect a problem. Investigate. Decide. Act. Improve. That sequence is incomplete. When you trace what happens in domains that produce durable knowledge across markets, biology, physics, AI deployment, and even insurance pricing, the same eight-stage cycle keeps appearing. The failure modes cluster around the stages most people skip. > Knowledge is a cycle, not a stack. Each completed cycle creates a new bottleneck that requires a new instrument. This is why “we already studied that” is rarely true. The cycle restarts the moment you finish it. ## 1\. Build the instrument Hidden structure exists in every domain. It stays hidden because the instrument that would reveal it has not been built yet. Quantum geometry waited decades for the right diffraction setup. Electron hydrodynamics waited 54 years between theory and clean experimental observation. [1](#footnote-1) Astrocyte function was structurally visible but functionally invisible until calcium imaging arrived. [2](#footnote-2) The engine cannot start without an observable. Theory is cheap. Instruments are expensive. “We do not know yet” usually means “we cannot see yet.” ## 2\. Notice the change Detection is the cheapest stage. Dashboards turn red, metrics move, anomalies fire. Modern systems are good at this. The danger is mistaking detection for understanding. The metric that moved tells you that something happened. It does not tell you what. ## 3\. Diagnose the cause This is the stage the pricing team failed. The same observation has different optimal responses depending on the cause. Conversion drift alone can mean five different things. The calibration drifted. The elasticity drifted. The customer mix shifted. The data pipeline broke. A business rule changed. Each demands a different intervention. Treat data drift as model drift, and you retrain into the wrong fix. The gap between “something is off” and “here is what to do” is where most teams burn weeks. ## 4\. Verify the diagnosis A diagnosis is itself a generated claim. It needs verification. In an era where any plausible explanation can be produced cheaply by people, by models, or by analysts under deadline, the bottleneck has moved from generating hypotheses to checking them. Tests can pass while the system runs orders of magnitude slower than it should.\[3\] Models can pass evals while gaming them. A diagnosis that _feels_ right because it addresses a real signal is a different object from a diagnosis that _is_ right. ## 5\. Check the frame Every claim carries an implicit baseline. “This strategy outperforms,” compared to what? “This metric improved,” relative to what reference class? The frame often dominates the conclusion more than the visible math. Survivorship bias is a reference class error. So is benchmark shopping. So is most “we beat the previous record.” A correct verification against the wrong baseline is a true fact in a misleading frame. ## 6\. Account for reflexivity Your action changes the system you are measuring. If the pricing optimiser narrows commission into a tight band, the training data loses the variation needed to re-estimate elasticity. If alignment researchers make compliance measurable, models may learn strategic compliance. If a fund publishes its strategy, the edge dissolves. The observer cannot be separated from the observed. Most monitoring systems pretend otherwise. ## 7\. Build the scaffold What persists is the architecture, not the content. KIBRA tags persist while the molecules that hold a memory degrade and are replaced. Sprint contracts persist while individual tickets close. Folder structures persist while specific notes go stale. Knowledge that is not scaffolded into persistent architecture, into a schema or a query or a cadence or a checklist, degrades as components turn over. Most teams “learn” something and then store the lesson as a Slack message. Three months later the lesson is gone. ## 8\. Track where value moved When a layer becomes cheap, abundant, or automated, value migrates upward. Generation gets cheap; verification becomes scarce. Information gets cheap; judgement becomes scarce. Detection gets automated; attribution becomes the bottleneck. The migration creates a _new_ bottleneck. Which requires a _new_ observable. Which restarts the engine at the first stage. ## Why most teams run three Detection. Action. Reaction. That is the cycle most teams run. The metric moved, do something, see what happens. The feedback loop feels like science. The cycle skips diagnosis, verification, frame-check, reflexivity, and scaffolding. The result is a publication of confident interventions that change every quarter, with no compounding learning underneath. A useful test: > If your team had to teach a new hire the _reasons_ your decisions worked, not just the decisions themselves, could you? If the answer is no, the scaffolding stage is failing. If the reasons sound plausible but no one has actually run the verification stage, the reasoning is generation pretending to be knowledge. ## The Legibility Paradox The engine has one deep tension that does not resolve. Building an instrument requires you to make hidden structure visible. Knowledge requires observability. Durable advantage usually lives in the part of the system that _cannot_ be measured. Taste. Judgement. Structural position. Trust. The illegible part. Reflexivity says that measuring something changes it. Build a metric for compliance and models will learn to be compliant for the metric. Build a metric for output quality and the team will optimise for the metric, not the quality. > The act of building an observable for the illegible may destroy the value you were trying to capture. Build the observable anyway. The engine demands it. Treat it as a finger pointing at the moon. Use it. Watch it degrade. Plan the next one before this one fails. The judgement that interprets the dashboard is where the durable value lives. The engine is the scaffold. The illegible judgement that runs it is the content. ## The one-week test Pick one decision you have recently regretted. Walk it backward through the eight stages. Did you have an instrument that would have revealed the underlying structure, or were you flying without one? Did the metric move and you act, without diagnosing the cause? Did you verify the diagnosis against an honest baseline, or against your favourite story? Did your action change the system in a way that contaminates the next decision? Is the lesson sitting in a Slack message, or built into a checklist, schema, or review cadence? Has the bottleneck already moved somewhere else? The skipped stage is where your next attention belongs. The teams that compound run all eight stages on every important decision, even slowly, even imperfectly. Discipline at every stage beats instinct at three. If this lens was useful, that is the shape of the publication: instruments for seeing what survives when the surface changes. Subscribe if you want one structural lens at a time, written so you can use it. [Figure] [1](#footnote-anchor-1) For the recent picture of astrocytes as an active neuromodulatory layer rather than passive support cells, see Ingrid Wickelgren, “Once Thought to Support Neurons, Astrocytes Turn Out to Be in Charge,” _Quanta Magazine_, 30 January 2026: [https://www.quantamagazine.org/once-thought-to-support-neurons-astrocytes-turn-out-to-be-in-charge-20260130/](https://www.quantamagazine.org/once-thought-to-support-neurons-astrocytes-turn-out-to-be-in-charge-20260130/) [2](#footnote-anchor-2) The case is the LLM-generated SQLite rewrite that passed the upstream test suite in full while running orders of magnitude slower than the implementation it replaced. Green tests, broken system. --- --- title: "The Five Laws of Durable Systems" description: "What still has a job after the change? Five tests for seeing what is likely to survive." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-five-laws-of-durable-systems/" date: "2026-05-06" series: "SYSTEMS & LAWS" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Five Laws of Durable Systems *What still has a job after the change? Five tests for seeing what is likely to survive.* By Harry Floyd · 2026-05-06 · canonical: https://durabilitycurve.com/blog/the-five-laws-of-durable-systems/ Most bad decisions begin with an upgrade. A better model. A cleaner dashboard. A stronger benchmark. A more convincing market story. A smoother workflow. An AI agent sails through the demo, then fails when its action becomes real. A fund buys the clean story, then discovers the bottleneck moved from information to timing. A team ships the dashboard, then learns the metric was measuring the wrong layer. The surface improved. The decision got worse. That is why so much smart analysis expires. It aims at the part that turns over first. Every piece here returns to one question: What still has a job after the change? Across AI systems, markets, biology, learning, design, operations, and strategy, the same pattern keeps returning: > Durable systems survive through the structure underneath the surface people are watching. A model can sit on top of the moat. A product can sit on top of the business. A score can sit on top of the proof. An easy step can sit on top of the valuable difficulty. A powerful capability can still point at the wrong problem. Five laws. Five tests. Here, a law is a pressure test. It forces a decision to declare which layer it is trusting. Use them before you trust a system, buy a company, adopt a tool, automate a workflow, ship a product, or believe a story. Use them while money, trust, reputation, or time is still on the table. Each section is a handle. Use it to make the next decision sharper. ### Law I ## Scarcity Moves When one layer becomes abundant, the scarce part moves. Generation gets cheap. Verification becomes scarce. Information gets cheap. Judgement becomes scarce. Tools get cheap. Integration becomes scarce. When capital becomes abundant, permission, distribution, and trust become scarce. The mistake is treating a solved bottleneck as if it stays solved in the same place. It rarely does. In AI, model access became easier; traces, evals, contracts, and workflow integration became the scarce work. In markets and content, the pattern is the same: once production becomes easier, judgement, proof, distribution, and integration become more valuable. The old bottleneck can remain visible long after it has stopped deciding the outcome. If your strategy is still aimed at yesterday's bottleneck, progress can make you later. **Test (Law I):** If this layer becomes abundant, where does scarcity move next? **Falsifier (Law I):** a domain where the layer became abundant and the value stayed put ### Law II ## Difficulty Carries Value Some hard parts are waste. Some hard parts are the mechanism. The second kind is where good systems get quietly destroyed. They remove friction and accidentally remove learning. They automate judgement and accidentally remove accountability. They simplify the workflow and accidentally remove the check that caught the bad decision. They make the interface smoother and accidentally make the hidden failure easier to miss. The right friction is where the system thinks. Remove the wrong friction and the system loses its memory. **Test (Law II):** Which hard part is producing the value? **Falsifier (Law II):** a system that removed its hard part and durably improved ### Law III ## Architecture Outlives Content Content turns over. Architecture persists. Cells replace molecules. Companies replace employees. Products replace features. Knowledge systems replace notes. AI systems replace models, prompts, tools, and vendors. The component usually turns over first. The scaffold lets components change without identity collapsing. If a product can swap the model underneath and the customer barely notices, the architecture lived in the contract around the model: what it could see, what it could change, how failures were caught, and how the workflow absorbed the output. If a company says it has an AI moat, ask what survives when the model is replaced tomorrow. Durability is rented when the value disappears with the component. Ownership begins when replacement leaves the value intact. **Test (Law III):** What persists after the pieces change? **Falsifier (Law III):** a system that survived on content while its structure turned over ### Law IV ## Visibility Must Be Built Hidden structure stays hidden until something makes it observable. Most arguments fail before they become arguments. They are missing the instrument that would settle them. They argue whether an agent is reliable without a replayable trace. They argue whether a product has a moat without a displacement test. They argue whether a team is learning without a review loop that shows belief change. They argue whether a model is better without naming the test setup that produced the score. Hidden-structure claims need instruments that make the structure answer back. Without the instrument, projection can look like sight. Sight has to be engineered. **Test (Law IV):** What would make the hidden structure visible? **Falsifier (Law IV):** a hidden-structure question settled by theory alone, no instrument ### Law V ## Capability Needs a Target More power amplifies wrong aim. A better model aimed at the wrong workflow creates more plausible waste. A faster team pointed at the wrong customer ships more irrelevant output. A smarter investor playing the wrong game loses with better reasons. A more sophisticated metric aimed at the wrong construct gives you cleaner self-deception. More capability makes the miss more expensive. Excellence at the wrong layer is still wrong. **Test (Law V):** Is the capability aimed at the right layer? **Falsifier (Law V):** raw capability on the wrong target producing durable gains ### Field run ## Run the Five Tests on One Decision Imagine your team is about to adopt an AI-agent platform. The demo is strong. The agent can browse docs, call tools, write tickets, update records, draft replies, and produce a clean score on a benchmark. The surface sentence is useful. It is also incomplete. Here is the same decision as a completed audit: You might still buy the platform. The demo becomes the opening claim. The trace becomes the reason. The purchase conversation changes. You ask what the agent can change, what must be true before the change becomes real, what permission disappears after a bad run, and whether the benchmark measures the work you need done or only the path that photographs well. The five laws have done their job when the impressive thing becomes specific enough to inspect. Better contact with reality is the job. ### Procedure ## The Audit Pick one live decision this week. A tool you want to adopt. A company you want to buy. A workflow you want to automate. A product you want to build. A metric you want to trust. A strategy you want to defend. Run the five tests. The form below is the specimen above, blank and live. It runs on this page and nowhere else. The tests earn their place when they change what you ask before purchase, deployment, allocation, or automation. Most people will keep watching the surface. The surface is louder. The structure underneath is quieter. Durable decisions start there. If one question changes what you were about to trust, I want to know which one. --- --- title: "The SpaceX IPO Is Not What You Think You're Buying" description: "The filing will not just price rockets. It will reveal which layer public investors actually own." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-spacex-ipo-is-not-what-you-think/" date: "2026-05-05" series: "MARKETS & POWER" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The SpaceX IPO Is Not What You Think You're Buying *The filing will not just price rockets. It will reveal which layer public investors actually own.* By Harry Floyd · 2026-05-05 · canonical: https://durabilitycurve.com/blog/the-spacex-ipo-is-not-what-you-think/ _This is an analytical framework, not financial advice. Reported IPO terms remain provisional until SpaceX publishes its S-1. [1](#footnote-1)_ The wrong question is already forming. It sounds sophisticated because it has a ticker-shaped answer: would you buy SpaceX? That question is too small. It compresses too many different things into one emotional decision. It turns a complicated offering into a referendum on rockets, Elon Musk, Mars, Starlink dishes, government contracts, xAI, retail access, and the idea that the future should be investable. The better question is stranger and more useful: What, exactly, would you be buying? Not just legally. Structurally. If the reported structure holds, the offering would be more than SpaceX selling shares. It would put Starlink’s cash flows, Falcon’s industrial proof, Starship’s option value, sovereign demand, xAI’s capital appetite, Cursor’s developer-workflow distribution, orbital-compute ambition, and Musk-controlled governance into one public-market instrument. The danger is not that investors will admire SpaceX. They should. The danger is that investors will price the bundle as if every layer is already proven, while receiving the rights of a minority passenger. The structure is simple: Starlink earns. Starship, xAI, Cursor, and orbital compute may consume. Governance decides who benefits. Price decides whether any of it matters. That is the lens to keep through the whole piece. This is not mainly a story about whether SpaceX is impressive. It is a story about what happens when an extraordinary private company becomes a public-market instrument. The company has an operating reality. The market has a price. The filing is the translation layer between them. That translation is where investors get hurt. That distinction matters because the reported SpaceX IPO would not be a normal listing. Reuters has reported that SpaceX confidentially filed for a U.S. IPO, that an early June roadshow is being targeted, and that the company could seek a valuation as high as roughly $1.75 trillion with a raise that could reach around $75 billion. [2](#footnote-2) Reuters has also reported filing-excerpt details on Starlink economics, xAI losses, and governance. Until the prospectus is public, those are reported claims, not final terms. But even as provisional reporting, they reveal the shape of the problem. At that scale, admiration is the easy part. The harder job is deciding which future has already been capitalised into the price. That is the real IPO question. ## Great Company, Wrong Question Public markets are very good at turning admiration into a price. They are less good at forcing people to say which part of their admiration is already in the price. SpaceX is not a shell with a story. It is an operating machine with proof in the world. Its official launches page, checked while building this draft, showed hundreds of completed missions, hundreds of landings, hundreds of reflights, and multiple recent Falcon missions. [3](#footnote-3) Starlink has turned satellite internet from a niche service into a mass distribution network: Starlink’s own network update said it had more than 6 million active customers globally as of July 2025, and Reuters-sourced coverage now reports that it crossed 10 million active customers in February 2026. [4](#footnote-4) Reuters-reported filing excerpts make the commercial point sharper: Starlink reportedly generated $11.4 billion of 2025 revenue and $4.42 billion of operating profit. [5](#footnote-5) NASA has awarded SpaceX major Artemis Human Landing System work, including a later Option B contract modification valued at roughly $1.15 billion. [6](#footnote-6) The hard question is whether Starlink’s cash engine is being sold as ownership, or used as collateral for Starship, xAI, Cursor, orbital compute, and founder-controlled optionality at a valuation where the future has already been monetised. ## The Six Economic Layers The cleanest way to read the offering, when it arrives, is not as one SpaceX story. It is as six layers stacked on top of each other. [Figure] ### 1\. The Proof Layer: Launch Cadence SpaceX’s foundational achievement is not merely that it launches rockets. It is that launch has become repeatable enough to look industrial. Reusability matters because it turns a heroic event into an operating rhythm. Cadence matters because every other layer depends on it. Starlink needs launch. Government customers need reliable access. Starship needs test frequency. The narrative of orbital infrastructure needs a company that can keep putting mass into orbit while competitors are still treating launch as a sparse event. This is the layer with the most visible proof. You can see the missions. You can see the landings. You can see the reflights. You can see the launch sites. You can see the company making launch feel less like a miracle and more like logistics. But in an IPO, visible proof is not enough. The filing has to answer whether cadence produces operating leverage. Does each incremental mission become cheaper? Are margins improving by customer type? How concentrated is demand? How much pad, range, safety, refurbishment, insurance, and failure reserve is required to keep the machine running? Launch cadence proves the machine. It does not prove the multiple. ### 2\. The Cash Engine: Starlink This may be the layer that turns SpaceX from a launch company into something closer to infrastructure. Launch gets satellites up. Starlink turns those satellites into customer relationships. That is a very different asset. A launch company sells missions. A broadband network sells recurring access. A launch company is judged by reliability and price per kilogram. A network is judged by subscribers, churn, ARPU, capacity, terminal cost, replacement capex, spectrum, enterprise mix, and distribution. The customer number matters, but it is no longer the main uncertainty. The reported 10 million-plus customer base is enough to prove scale. The deeper question is what that scale has to fund. Reuters-reported filing excerpts say Starlink produced $11.4 billion of 2025 revenue and $4.42 billion of operating profit. That changes the burden of proof. Starlink is not merely a promising broadband project. It is, on current reporting, the cash engine inside the group. The question is whether the engine is free to compound, or whether it is being asked to carry everything else. That engine is real, but it is not frictionless. The Information and syndicated market reports say Starlink’s average revenue per user fell 18% to roughly $81 a month between 2023 and 2025 as the service expanded into lower-priced plans and geographies. [7](#footnote-7) Does the network become cheaper to serve as it grows, or does each wave of growth require new satellites, new ground infrastructure, subsidised terminals, and continuous replacement spend? Does direct-to-cell become a second distribution curve, or an expensive feature? Does enterprise, maritime, aviation, and government demand protect the economics as residential pricing compresses? The reported consolidated picture makes this sharper. Reports based on Reuters filing excerpts say SpaceX’s newly consolidated AI business posted a $6.4 billion operating loss in 2025 and consumed roughly 61% of group capex, while the combined company lost nearly $5 billion on about $18.7 billion of revenue. [8](#footnote-8) If those numbers survive the public filing, Starlink becomes more than a growth story. It becomes the engine being asked to fund the next frontier. Starlink is the part of SpaceX that most resembles a public-market business. The investment question is whether public holders get to own its compounding, or mainly underwrite what it is being used to finance. ### 3\. The Option Layer: Starship Starship is the option layer. It is the part of the story that expands the possible future more than it explains the present. If Starship works at scale, the cost and volume assumptions around orbit change. Starlink deployment changes. Lunar logistics change. Mars changes. Orbital manufacturing, propellant depots, military logistics, and large-scale cargo all move from slideware to a different sort of conversation. But options are not cash flows. They are claims on a future state of the world. That does not make them worthless. Some of the most valuable companies in history were underpriced because people could not value their option layers. The mistake is not valuing optionality. The mistake is paying for optionality as if it has already cleared the gates. For Starship, the gates are unusually concrete. Test progress. Flight cadence. Regulatory approvals. Site capacity. NASA milestones. Payload commitments. Failure rates. Refurbishment assumptions. Capitalised development cost. The FAA process around Starship operations at Kennedy Space Center’s LC-39A is a reminder that the bottleneck includes licensing, environmental review, safety, range operations, and public tolerance for cadence, not engineering alone. [9](#footnote-9) Any major Starship test near the reported roadshow window will trade as narrative evidence, not as a normal engineering update. [10](#footnote-10) Starship is not a segment yet. It is a valuation bridge. ### 4\. The Sovereign-Demand Layer: Government SpaceX also sits inside national space capacity. NASA, the Space Force, national security customers, lunar missions, launch resilience, and secure communications give the company a structural role that a normal consumer-tech frame misses. Government demand can be durable and strategic. It can also be fixed-price, milestone-heavy, politically exposed, classified, bureaucratic, and margin-constrained. A backlog headline is not enough. The filing should show contract concentration, termination rights, milestone exposure, segment margins, Starshield-style military demand, and how much of the company’s future depends on public-sector budgets. Government demand is a moat until it becomes concentration. ### 5\. The Capital-Absorption Layer: xAI, Cursor, Orbital Compute This is the layer most likely to produce both real upside and bad analysis, and it is now too material to leave as a vague optionality bucket. Reuters reported that SpaceX acquired xAI in February 2026 in an all-stock transaction that valued the combined company at roughly $1.25 trillion. [11](#footnote-11) TechCrunch and Reuters-syndicated reporting also say SpaceX has announced a Cursor arrangement: either a $10 billion partnership payment or an option to acquire the coding and knowledge-work AI company for $60 billion later in 2026. [12](#footnote-12) Separate reporting says SpaceX has sought regulatory review for orbital AI/data-centre satellite plans. [13](#footnote-13) Those claims are not all equal. The xAI transaction is a Reuters-reported corporate event. The Cursor arrangement has a public statement and press coverage, but the detailed terms are still thin. The orbital compute ambition is a regulatory-and-strategy claim, not an operating business. Still, together they change the analytical frame. The Musk ecosystem has moved from adjacency to possible issuer-level risk. The bullish version is that SpaceX becomes the physical layer for AI: launch puts infrastructure in orbit, Starlink distributes connectivity, xAI supplies models and compute, Cursor supplies developer workflow distribution, and Tesla may provide storage or energy adjacency. It is also exactly the kind of thesis that can become unfalsifiable if investors let it float above the accounts. If xAI, Cursor, orbital compute, or broader AI infrastructure is part of the SpaceX story, the public filing must tell investors what they actually own. Is xAI consolidated? How much loss and capex comes with it? Are there related-party transactions with Tesla, X, or other Musk-controlled entities? Who funds the compute build-out? What assets sit in which entity? What are the Cursor payment and acquisition obligations? What governance rights protect outside shareholders? Are capital allocation decisions made for SpaceX shareholders, or for the ecosystem as a whole? The Cursor point is a good example. Strategically, a coding and knowledge-work AI layer could produce real engineering-productivity gains inside the rocket and satellite programmes. It could also be exactly the kind of late-cycle narrative expansion public investors should interrogate: expensive, adjacent, exciting, and not yet proven as a return stream. AI optionality should not be dismissed. But optionality without legal and accounting clarity is not a thesis. It is a mist. The bear case is not that AI is irrelevant. It is that Starlink’s cash flows are redirected into an AI capex race with unclear returns. ### 6\. The Ownership Layer: Governance, Liquidity, Price This is the layer that enthusiastic investors least want to discuss, and the one that may matter most. A historic IPO would create a new public liquid instrument for one of the most desired private companies in the world. That liquidity has value. It also has danger. Retail access can democratise participation, but it can also turn scarcity into demand pressure. If a large retail allocation is part of the offering, as Reuters has reported, the investor has to ask whether retail is being invited into a durable seat or into a narrative event. Governance is not a footnote here. It is the mechanism by which public investors find out whether they are owners, passengers, or liquidity providers. Reuters-syndicated reports on filing excerpts point to a dual-class structure in which public investors buy lower-vote Class A shares while Musk and insiders retain super-voting Class B control, with some reports putting Musk at roughly 42% economic ownership and about 79% voting control. [14](#footnote-14) Treat those exact percentages as provisional until the prospectus is public. Treat the direction as central. Reuters-syndicated reporting also says the filing language would make Musk removable from board or top roles only by Class B holders, while a proposed compensation package could grant large additional super-voting awards tied to extreme Mars and orbital-compute milestones. [15](#footnote-15) The governance point does not need exaggeration. If the reported structure holds, public shareholders would be buying into one of the most ambitious companies in the world while accepting unusually limited control over how that ambition is directed. Control is not incidental to this IPO. It is one of the assets being sold around. Share classes, voting control, lockups, insider sales, related-party rules, use of proceeds, segment disclosure, and risk-factor language are not footnotes here. They are the ownership terms. The company can be extraordinary and still offer public investors weak rights at a demanding price. Governance determines whether public investors are buying ownership or exposure. ## What The Proceeds Actually Buy The reported $75 billion raise should not be read as generic rocket money. At this scale, the IPO is a capital-allocation document. The use-of-proceeds section will show whether public investors are funding Starship cadence, Starlink capacity, AI compute, orbital data-centre ambition, debt reduction, insider liquidity, or some mixture of all of them. Reported filing excerpts already flag orbital data-centre plans while warning that they may not become commercially viable. If the proceeds strengthen the operating substrate, public investors may be buying into a seat that compounds. If they mainly finance losses, related-party complexity, or narrative expansion ahead of proof, they are underwriting the next layer of the story. This is why the use-of-proceeds section may be one of the most important pages in the filing. [Figure] ## The Reference Class Trap The valuation debate will look quantitative. It will be full of numbers, multiples, curves, comps, TAMs, backlogs, and scenario cases. Underneath those numbers will be one hidden decision: compared to what? If you compare SpaceX to launch providers, the valuation will look impossible. If you compare it to telecom infrastructure, it will depend on Starlink’s margins and replacement capex. If you compare it to defence primes, you will care about government backlog, political risk, and free cash flow durability. If you compare it to mega-cap platforms, you will focus on ecosystem control and the ability to compound across layers. If you compare it to AI infrastructure, you will care about compute, power, software distribution, and capex absorption. Reuters reported that bankers and investors have been reaching for Palantir, GE Vernova, and Vertiv-style AI infrastructure comparisons rather than Boeing or telecom comps. [16](#footnote-16) Damodaran’s pre-prospectus valuation work reached a base case around $1.22 trillion and a simulation median around $1.29 trillion, while noting that $1.75 trillion to $2 trillion pricing leaves little obvious upside for a new buyer. [17](#footnote-17) That spread is not a trivia point. It is the whole psychological game. The same facts can look cheap or expensive depending on which future you let into the denominator. The reference class is not a neutral choice. It is the move that makes the valuation possible. None of those reference classes is obviously right. That is the point. The biggest analytical error will be choosing the flattering reference class implicitly. SpaceX will be called an infrastructure company when people want durability, a technology company when people want growth, a defence asset when people want sovereign importance, a telecom company when people want recurring revenue, an AI company when people want multiple expansion, an “AWS in space” when people want platform economics, and a founder-led compounder when people want to explain away governance. The S-1 should be read as a reference-class document. Not because the reference class is an academic detail. Because the reference class quietly decides what future you are paying for. Which business actually carries revenue? Which business carries margin? Which business carries capex? Which business carries the valuation? Those may not be the same business. ## The Part Investors Will Be Tempted To Skip There is a specific kind of company where scepticism feels small. SpaceX is one of them. The accomplishments are so visible that ordinary caution can look like a failure of imagination. The rockets land. The satellites work. The launch cadence is real. The government trusts the company with missions that matter. Starlink has millions of customers. Starship could change the cost curve of orbit. The founder has already made several impossible-looking markets real. All true. But investing does not reward awe. It rewards the relationship between price, rights, cash flows, growth, risk, and time. At one valuation, SpaceX might be Starlink cash flow with free Starship optionality. At another, it becomes Starlink cash flow plus paid Starship optionality. At the reported IPO range, it may become a bet that launch, Starlink, Starship, government demand, xAI, Cursor, orbital compute, Mars, and retail scarcity all work, and that public investors still receive enough economics after the structure is defined. Same company. Different investment. ## The Six Questions That Matter When the filing appears, do not start with the valuation. Start with six questions: 1. Which segment carries revenue? 2. Which segment carries margin? 3. Which segment carries capex? 4. Which segment carries losses? 5. Which segment carries the valuation? 6. What rights do public shareholders actually receive? Then read the details underneath those questions. Launch cadence proves capability, not valuation. Starlink should show whether it is a cash engine after constellation maintenance. Starship should show whether it is priced as an option or a certainty. Government demand should show whether durability is becoming concentration. xAI, Cursor, and orbital compute should show whether public shareholders own the upside or merely fund the spend. Governance should show whether public investors are owners, passengers, or liquidity providers. Only then ask the price question. ## What Would Change The Answer The bullish version is not hard to imagine. The filing shows Starlink converting scale into strong free cash flow after replacement capex. Launch margins improve with cadence. Starship risk is disclosed clearly but not carrying the whole valuation. Government backlog is durable without becoming the only profit pool. xAI losses are large but bounded, Cursor is tied to real engineering-productivity gains inside the rocket and satellite programmes, related-party boundaries are clean, and AI capex has a visible route to revenue. Governance is founder-controlled but not abusive. Use of proceeds strengthens the operating substrate. The valuation is demanding but not so demanding that every future layer must work perfectly. That would be a serious public-market asset. The bearish version is also not hard to imagine. The filing reveals that Starlink growth is capex-hungry after replacement spend, launch cadence is operationally heroic but financially thinner than assumed, Starship is essential to the valuation but still distant from commercial proof, government demand is milestone-risky or politically concentrated, xAI losses widen, Cursor becomes another expensive option, related-party complexity muddies the economics, insiders sell into retail demand, and the valuation already prices every option as if it were a proven segment. That would still be an extraordinary company. It might not be an attractive IPO. This is the distinction the public conversation will try to erase. Keep it alive. ## The Real Thing Being Sold The obvious story is rockets. The more sophisticated story is Starlink. The grand story is Mars. The market story is scarcity: a company everyone has heard of, few have been able to own, and many will want the moment it becomes available. But the structural story is different. SpaceX may be selling public investors access to a seat in the orbital economy. Not a single product, but a position: launch, satellites, communications, government access, Starship capacity, AI compute, developer workflow, maybe a new layer of physical distribution above the planet. That is why the company matters. It is also why the IPO could be dangerous. The best seats are accumulated slowly and priced imperfectly before the world understands them. By the time everyone recognises the seat, the price may already include the seat and every future use of it. So when the SpaceX filing arrives, the question is not whether the company is impressive. That part is obvious. The question is whether public investors are being offered the seat, or being asked to finance the story of the seat at a price that assumes every future use of it has already been won. That is the IPO question. If you take one habit from this piece, make it this: when a story feels obviously great, slow down and ask what part of the machine you can actually prove. I wrote about the same test in public markets in [PLTR: The AI Stock That Has To Prove It Owns The Permission Layer](https://open.substack.com/pub/harryfloyd/p/pltr-the-ai-stock-that-has-to-prove?r=2u3t9p&utm_campaign=post&utm_medium=web&showWelcomeOnShare=true), and from the operator side in [AI Made You Faster. It Did Not Make You Safer.](https://durabilitycurve.com/blog/ai-made-you-faster-it-did-not-make/). Both are really about the same thing: do not confuse exposure with ownership, or speed with proof. If the SpaceX filing lands and you read it, I would love to know which layer looks most proven to you, and which layer looks most like story. ## Field Card [Figure] [1](#footnote-anchor-1) No public SpaceX S-1 or prospectus was available during the May 5 pre-publication check. That is why the article treats Reuters-reported filing excerpts as provisional and points readers back to the eventual prospectus as the document that should settle the terms. [2](#footnote-anchor-2) Reuters reported that SpaceX had confidentially filed for a U.S. IPO, was targeting an early June roadshow, and could seek a valuation around $1.75 trillion with a raise of up to roughly $75 billion. Those figures remain reported terms until a public prospectus is available: [https://www.reuters.com/business/spacex-lays-out-ipo-details-targets-early-june-roadshow-sources-say-2026-04-07/](https://www.reuters.com/business/spacex-lays-out-ipo-details-targets-early-june-roadshow-sources-say-2026-04-07/) [3](#footnote-anchor-3) SpaceX’s official launch record is the source for completed missions, landings, reflights, and recent Falcon activity: [https://www.spacex.com/launches/](https://www.spacex.com/launches/) [4](#footnote-anchor-4) Starlink said it had more than 6 million active customers globally as of July 2025. Reuters-sourced market coverage later reported that Starlink crossed 10 million active customers in February 2026: [https://www.starlink.com/networkupdate](https://www.starlink.com/networkupdate) and [https://finance.yahoo.com/markets/stocks/articles/starlink-user-growth-accelerates-spacex-134543171.html](https://finance.yahoo.com/markets/stocks/articles/starlink-user-growth-accelerates-spacex-134543171.html) [5](#footnote-anchor-5) Reuters-reported filing excerpts are the source for the Starlink 2025 revenue and operating-profit figures cited in the piece: [https://reuters.com/science/spacex-posted-nearly-5-billion-loss-2025-information-reports-2026-04-10](https://reuters.com/science/spacex-posted-nearly-5-billion-loss-2025-information-reports-2026-04-10) [6](#footnote-anchor-6) NASA’s Artemis Human Landing System Option B modification is the source for the roughly $1.15 billion contract figure: [https://www.nasa.gov/press-release/nasa-awards-spacex-second-contract-option-for-artemis-moon-landing-0/](https://www.nasa.gov/press-release/nasa-awards-spacex-second-contract-option-for-artemis-moon-landing-0/) [7](#footnote-anchor-7) The Information / syndicated market reporting is the source for the reported 18% fall in Starlink ARPU to about $81 between 2023 and 2025: [https://sa.marketscreener.com/news/spacex-says-starlink-s-arpu-fell-18-to-81-to-keep-falling-the-information-ce7f58dadc88ff20](https://sa.marketscreener.com/news/spacex-says-starlink-s-arpu-fell-18-to-81-to-keep-falling-the-information-ce7f58dadc88ff20) [8](#footnote-anchor-8) Reuters-syndicated reporting on filing excerpts is the source for the reported consolidated revenue, loss, xAI operating loss, and capex-share figures: [https://ca.finance.yahoo.com/news/exclusive-spacex-conquered-stars-now-100320149.html](https://ca.finance.yahoo.com/news/exclusive-spacex-conquered-stars-now-100320149.html) [9](#footnote-anchor-9) FAA Starship-Super Heavy materials for LC-39A show why Starship cadence is partly a licensing, safety, environmental, and range-operations question, not only an engineering question: [https://www.faa.gov/space/stakeholder\_engagement/spacex\_starship\_ksc](https://www.faa.gov/space/stakeholder_engagement/spacex_starship_ksc) [10](#footnote-anchor-10) NASASpaceFlight reported Starship Flight 12 static-fire activity and a possible mid-May target window. This matters here because a visible test near a roadshow can influence narrative, even before it proves commercial economics: [https://www.nasaspaceflight.com/2026/05/spacex-mid-may-starship-flight-12-revised-trajectory/](https://www.nasaspaceflight.com/2026/05/spacex-mid-may-starship-flight-12-revised-trajectory/) [11](#footnote-anchor-11) Reuters reported the SpaceX-xAI all-stock transaction and the reported combined valuation: [https://www.reuters.com/business/musks-spacex-merge-with-xai-combined-valuation-125-trillion-bloomberg-news-2026-02-02/](https://www.reuters.com/business/musks-spacex-merge-with-xai-combined-valuation-125-trillion-bloomberg-news-2026-02-02/) [12](#footnote-anchor-12) TechCrunch reported the Cursor arrangement, including the partnership payment and later acquisition option described in the piece: [https://techcrunch.com/2026/04/21/spacex-is-working-with-cursor-and-has-an-option-to-buy-the-startup-for-60-billion/](https://techcrunch.com/2026/04/21/spacex-is-working-with-cursor-and-has-an-option-to-buy-the-startup-for-60-billion/) [13](#footnote-anchor-13) TechCrunch reported SpaceX’s orbital AI/data-centre satellite filing. The article treats this as a strategic ambition and regulatory request, not approved commercial capacity: [https://techcrunch.com/2026/01/31/spacex-seeks-federal-approval-to-launch-1-million-solar-powered-satellite-data-centers/](https://techcrunch.com/2026/01/31/spacex-seeks-federal-approval-to-launch-1-million-solar-powered-satellite-data-centers/) [14](#footnote-anchor-14) Reuters-syndicated reporting on filing excerpts is the source for the reported Class A/Class B governance structure and provisional Musk voting-control figures: [https://finance.yahoo.com/markets/stocks/articles/exclusive-musk-insiders-retain-voting-065827904.html](https://finance.yahoo.com/markets/stocks/articles/exclusive-musk-insiders-retain-voting-065827904.html) [15](#footnote-anchor-15) Reuters-syndicated reporting is the source for the described Musk compensation package tied to Mars and orbital-compute milestones: [https://ca.finance.yahoo.com/news/analysis-spacex-ties-musk-compensation-100426004.html](https://ca.finance.yahoo.com/news/analysis-spacex-ties-musk-compensation-100426004.html) [16](#footnote-anchor-16) Reuters / Investing.com reported that bankers and investors were reaching for Palantir, GE Vernova, and Vertiv-style AI-infrastructure comparisons when discussing SpaceX’s reported valuation: [https://ca.investing.com/news/stock-market-news/the-unconventional-logic-behind-spacexs-175-trillion-price-tag-4558854](https://ca.investing.com/news/stock-market-news/the-unconventional-logic-behind-spacexs-175-trillion-price-tag-4558854) [17](#footnote-anchor-17) Aswath Damodaran's April 2026 SpaceX valuation piece is the source for the base-case and simulation-median valuation references: [https://open.substack.com/pub/aswathdamodaran/p/to-trillions-and-beyond-a-spacex](https://open.substack.com/pub/aswathdamodaran/p/to-trillions-and-beyond-a-spacex) --- --- title: "AI Made You Faster. It Did Not Make You Safer." description: "The strange thing about the AI productivity boom is that the people getting faster are not always getting more secure. Speed is becoming the surface. Proof is moving somewhere else." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/ai-made-you-faster-it-did-not-make/" date: "2026-05-04" series: "THE HUMAN LAYER" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # AI Made You Faster. It Did Not Make You Safer. *The strange thing about the AI productivity boom is that the people getting faster are not always getting more secure. Speed is becoming the surface. Proof is moving somewhere else.* By Harry Floyd · 2026-05-04 · canonical: https://durabilitycurve.com/blog/ai-made-you-faster-it-did-not-make/ ## The private feeling under the productivity story The new anxiety does not always arrive as panic. Sometimes it arrives at 11:17 on a Tuesday morning, just after the work goes strangely well. You had blocked out the whole morning for the thing you were avoiding. A deck. A research memo. A customer summary. A bug you did not want to touch. A product plan that had been sitting in your notes for a week because the first draft felt too heavy to start. Then the model does enough of it in twenty minutes. It is not perfect. It is enough. The page is no longer blank. The meeting transcript has structure. The argument has headings. The spreadsheet has an explanation. The first version of the plan exists. You can see the shape now. For a moment, this feels like relief. Then something quieter arrives underneath it. If the thing that made me feel useful can appear this quickly, what exactly was scarce about me? That is the feeling most AI productivity advice does not touch. It tells you to move faster, ship more, automate the boring parts, become a one-person team, learn agents, build workflows, and use the latest model before someone else does. Some of that advice is useful. But it skips the part people are actually carrying. The obvious fear is that AI might take work away. The deeper fear is that AI is making the old proof of work weaker while everyone is still pretending productivity is the whole story. You can feel more capable and less safe at the same time. That is the paradox. So this piece has to do more than diagnose the feeling. By the end, you should have three things you can actually use: - a way to tell whether AI is making your work safer or just faster - a workflow for turning AI output into proof someone can trust - a prompt pattern you can copy whenever the task matters The aim is not to make you feel better about AI. It is to give you a better instrument for deciding where your value should move next. ## The data now has a human shape [Anthropic recently published research](https://www.anthropic.com/research/81k-economics) based on roughly 81,000 Claude users. The headline is bigger than people using AI at work. Everyone knows that now. The interesting part is the contradiction. People reported meaningful productivity gains. Anthropic rated the average inferred productivity gain at 5.1 on its scale, corresponding to “substantially more productive.” Among respondents who described productivity effects, 48 percent talked about expanded scope, while 40 percent talked about speed. But one fifth of respondents also voiced concern about economic displacement. People in the most AI-exposed jobs mentioned job threat roughly three times as often as people in the least exposed jobs. Early-career workers were more nervous than senior workers. And the people reporting the largest speedups were also more likely to worry about AI’s job impact. That last point matters. The speedup did not automatically produce confidence. Sometimes the speedup was the reason confidence cracked. [Computerworld framed the same tension](https://www.computerworld.com/article/4162929/the-ai-workplace-paradox-higher-productivity-higher-anxiety.html) as the AI workplace paradox: higher productivity, higher anxiety. Developers, IT workers, market researchers, QA analysts, support specialists, and other exposed roles are not standing outside the technology, speculating about a distant future. They are using the tools. They are feeling the acceleration directly. This is why the conversation feels stranger than an ordinary technology cycle. The tool is useful, and its usefulness is part of the fear. ## The trap is mistaking speed for safety The optimistic version says productivity is protection. If you use AI well, you become faster. If you become faster, you become more valuable. If you become more valuable, you become safer. That chain sounds reasonable until everyone else gets access to the same speed. Speed protects you only while speed is scarce. Once speed becomes ambient, it stops being the proof. It becomes the baseline. The task that used to take a morning now takes twenty minutes. The analysis that used to look impressive now looks normal. The clean first draft no longer proves you wrestled with the problem. The slide no longer proves you saw the structure. The code no longer proves you understood the trade-off. The summary no longer proves you read the source carefully. The visible artefact still matters. It just means less than it used to. This is the same pattern that appears whenever a layer gets cheap. The bottleneck migrates. When generation gets cheaper, verification gets more valuable. When output gets easier, judgement gets more important. When speed becomes common, the question moves from “Can you produce?” to “Can anyone trust what you produced?” That is where the anxiety comes from. People are competing with a new standard of evidence. ## There are two kinds of AI productivity The Anthropic data separates something important: scope and speed. Speed means AI helps you do a task faster. Scope means AI helps you do something you could not do before. Those do not feel the same. [Figure] _The useful question is whether the speedup moved you toward a stronger proof layer, or merely made the old surface cheaper._ If AI lets a founder build a prototype, a designer test more visual directions, a marketer analyse customer interviews, or a non-technical operator make a tool that used to require an engineer, that can feel like expanded agency. The person is moving inside a larger box. But when AI mainly accelerates the work you were already paid to do, the feeling can turn unstable. The task shrinks. The expectation rises. The proof weakens. What used to count as a full day becomes half a day. What used to be impressive becomes table stakes. What used to be a training ground becomes automated away before it can teach anyone. Computerworld quoted Sanchit Vir Gogia making a point every manager should sit with: faster generation can raise expectations on quality, and more output can feed decision pipelines that were already constrained. In some cases, the system becomes heavier, not lighter. That is the part the productivity story misses. AI does not enter a clean system. It enters existing approval chains, status games, hiring ladders, review rituals, political incentives, overloaded managers, insecure juniors, under-defined roles, and metrics that already confused movement with progress. Acceleration inside a confused system does not automatically produce clarity. Sometimes it produces faster confusion. ## The entry-level problem is really a proof problem One of the most important lines in the Computerworld piece is not about job loss directly. It is about the path into the job. Basic coding, documentation, routine analysis, QA, structured support, and first-pass research are often described as low-level work. That makes them sound expendable. But for a person becoming competent, low-level work is not just production. It is training. The junior analyst builds the simple model before they learn which assumptions matter. The support rep handles repetitive tickets before they understand the product’s real failure modes. The marketer writes the bad first drafts before they develop taste. The developer fixes small bugs before they can reason about architecture. The researcher summarises sources before they can see what the sources are hiding. If AI compresses that layer, the organisation may feel more efficient now and discover later that it has quietly damaged the apprenticeship path that produced judgement. This is why “AI will automate the boring work” is too simple. Some boring work is waste. Some boring work is load-bearing. You do not know which until you ask what capacity the friction was building. If the friction was only moving information from one box to another, automate it. If the friction was teaching the person how the system fails, be careful. You may be removing the part of the work that turned exposure into judgement. ## What still proves you are valuable? When output gets cheap, value does not disappear. It moves. The old proof was often attached to the surface: the memo, the deck, the clean code, the finished research, the generated strategy, the polished artefact. The new proof has to move closer to the system around the artefact. That means your safest work is no longer just the part that produces the answer. It is the part that makes the answer worth trusting. There are five places to look. Problem choice: did you aim the tool at the right question? Source judgement: did you know what evidence deserved belief? Rejection: did you know which plausible output to delete? Ownership: can you explain and stand behind the final version? Learning loop: does the work make the next decision better, or only create the next artefact faster? This is the migration map. If AI makes the visible artefact easier to produce, your value has to move into the judgement system that decides what gets produced, what gets trusted, what gets rejected, and what gets shipped. That is the article in one sentence: > Stop trying to be the fastest producer of the surface. Become the person who can make the surface trustworthy. ## The work is not gone. The work has moved. This is the mistake in both the panic and the hype. The panic says AI will do the work. The hype says AI will free us from the work. Both assume the work is the visible task. But in most serious knowledge work, the task was never the whole job. The task was the part of the job that could be named. Write the memo. Fix the bug. Summarise the meeting. Build the model. Draft the plan. Make the deck. Compare the vendors. Ship the feature. Underneath those tasks was the harder layer: knowing what mattered, noticing what was missing, making trade-offs, protecting context, earning trust, sequencing effort, resisting bad incentives, and deciding what you were willing to stand behind. AI attacks the named layer first. That does not make the unnamed layer less important. It makes the unnamed layer harder to avoid. The person who only became faster at producing surfaces may feel less safe because the market can now buy more surfaces. The person who becomes better at deciding which surfaces deserve trust is moving toward the new bottleneck. That is the difference. ## The speed-to-proof workflow The practical move is simple. If AI has made a task faster, do not stop at the speedup. Run the output through four layers. This is the part to actually use. Open a real AI-assisted task from this week and walk it through the sequence below. A meeting summary, product plan, hiring screen, research memo, code change, customer email, investment note, content draft, or strategy doc all work. [Figure] _A saveable workflow for turning AI output into something another person can trust._ ### 1\. Name what got cheaper Start by identifying the layer AI compressed. Did it make the first draft cheaper? Did it make search cheaper? Did it make summarising cheaper? Did it make visual exploration cheaper? Did it make coding the obvious path cheaper? This matters because the cheapened layer is no longer where you should look for safety. If AI made first drafts cheap, a first draft is not proof. If AI made summaries cheap, a summary is not proof. If AI made prototypes cheap, a prototype is not proof. The first move is to stop treating the compressed layer as the evidence layer. ### 2\. Convert the output into claims AI output is usually shaped like an artefact: memo, plan, draft, slide, answer, summary, code. Proof starts when you break that artefact into claims. Use this format: ```markdown Claim: What is this output asking us to believe? Evidence: What source, observation, customer fact, test, or prior decision supports it? Risk: Where could this be wrong, brittle, misleading, or overconfident? Owner: Who is willing to stand behind this after the model disappears? Next test: What would we check before acting on it? ``` That small format changes the work. The output stops being a polished surface and becomes an object someone can inspect. This is why the workflow matters. Most AI tools make artefacts easier to create. This makes artefacts easier to trust. ### 3\. Add a verifier before you add volume Most people respond to AI speed by increasing output. More drafts. More options. More experiments. More content. More code. More analysis. That is tempting, but it is often the wrong order. If generation got faster, the next investment should be verification. Before increasing volume, decide what would make the output trustworthy: - source check - customer check - counterexample search - second-person review - test suite - evaluation rubric - decision log - rollback path The rule is simple: > Do not scale the output until you have scaled the proof. Otherwise AI has not made you safer. It has made your uncertainty more productive. ### 4\. Build memory around the decision The final layer is memory. Not memory in the vague “save your prompts” sense. Decision memory. For any meaningful AI-assisted work, preserve four things: - what the model produced - what you changed - what you rejected - why the final version deserved trust This is where a person becomes harder to replace. Their advantage is the visible judgement path, not the typing. ## A worked example: the meeting summary that can hurt you Take the most ordinary possible example: a meeting summary. This is exactly the kind of work people are happy to hand to AI because it feels low-risk. The meeting happened. The transcript exists. The model can summarise it. Everyone gets their time back. But a summary becomes consequential the moment people act on it. Imagine a product team has a customer call about churn. The AI summary says: > Customers are leaving because onboarding is confusing. Next step: simplify onboarding emails and create a better help centre flow. That sounds useful. It might even be right. But if you ship from that summary, you have skipped the proof layer. Run the workflow. What got cheaper? The transcript-to-summary step. The old proof was: “someone listened carefully and understood the customer.” That proof is now weaker because a plausible summary can arrive without careful listening. What claims are being made? Claim one: customers are leaving because onboarding is confusing. Claim two: email and help-centre changes are the right response. What evidence supports them? Maybe three customers mentioned confusion. But did they churn because of it, or did they mention it after already deciding the product was not valuable enough? Did power users say the same thing? Did support tickets show the same pattern? Did activation data show drop-off at onboarding, or later when the product failed to become a habit? What is the risk? The team might fix the easiest visible complaint while missing the real retention problem. Who owns the judgement? Someone has to say: “I believe onboarding is the bottleneck” or “I think onboarding is only the polite surface reason.” What is the next test? Pull the last ten churned accounts. Compare transcript complaints with usage data. Look for whether confusion appears before disengagement or after it. Then decide whether onboarding is the cause, a symptom, or a convenient story. Now the AI summary has become useful. The summary became useful because it was turned into claims, evidence, risks, ownership, and a next test. That is the difference between an AI output and a decision-grade artefact. ## Two prompts: one makes output, one builds proof Most AI prompting still aims at surface production. That is fine for low-stakes work. It is weak for anything that affects customers, strategy, hiring, money, reputation, or trust. Compare the difference. Surface prompt: ```markdown Summarise this meeting transcript and give me the key action items. ``` Proof prompt: ```markdown Turn this meeting transcript into a decision-grade summary. Separate: 1. Decisions actually made 2. Open questions still unresolved 3. Claims that need evidence 4. Risks or assumptions people glossed over 5. Actions, owners, and deadlines 6. What should be verified before anyone acts on this If the transcript does not support a conclusion, say so. ``` The first prompt makes a cleaner artefact. The second prompt creates a trust surface. Another example: Surface prompt: ```markdown Create a launch plan for this product. ``` Proof prompt: ```markdown Create a launch plan for this product, but organise it as a proof system. For each recommendation, include: - the customer belief it depends on - the evidence we currently have - the weakest assumption - the first cheap test - the failure signal that would make us stop - the person who owns the decision Do not optimise for a confident plan. Optimise for a plan we can safely learn from. ``` This is the shift. Do not ask AI only to produce the thing. Ask it to expose what would make the thing trustworthy. That one change is often enough to separate useful AI work from impressive-looking noise. ## What managers should learn If you manage people, do not treat AI productivity as a simple capacity increase. The lazy version is to say: the same person can now do twice as much, so expectations should double. That may work for a quarter. It may also destroy the slower system that produces competence. If AI removes junior tasks, you need a new apprenticeship path. If AI increases output, you need a stronger verification layer. If AI expands scope, you need clearer ownership. If AI makes everyone faster, you need better judgement about which work deserves speed. The management question is not: > How much more can we get? The better question is: > What proof of competence did the old work produce, and what replaces it now? That question turns into three concrete management moves. First, protect an apprenticeship layer. If AI removes junior tasks, deliberately replace the learning function those tasks used to serve. A junior person still needs reps in debugging, source judgement, customer contact, messy data, ambiguous trade-offs, and the slow discovery of what “good” looks like. Do not let “the model can do that now” become “nobody learns how the system works.” Second, create a verification budget. For every AI-assisted workflow, decide how much time belongs to output and how much belongs to proof. A rough starting rule: > If the work affects a real decision, spend at least 30 percent of the saved time on verification. If a task used to take three hours and AI makes it take one, do not automatically fill the extra two hours with more tasks. Put some of that time into checking assumptions, testing edge cases, improving the rubric, or teaching someone how the decision is made. Third, require an ownership receipt. Before important AI-assisted work is shipped, ask for a short note: ```markdown What did AI produce? What did the human change? What was rejected? What evidence supports the final version? What would make us revise this? Who owns the outcome? ``` This is not bureaucracy. It is how you prevent productivity from quietly destroying accountability. Most organisations will not ask that early enough. They will celebrate acceleration, then wonder why trust, training, review quality, and decision clarity did not improve at the same rate. That is how productivity becomes a trap. ## What individuals should learn If you are using AI and feeling the strange mix of power and unease, do not dismiss it as irrational. The feeling is information. It may be telling you that your work has been over-identified with a surface AI is now compressing. That does not mean you are obsolete. It means your old proof is weakening. Move your effort upward. Getting faster at drafting helps. Deciding what deserves to be drafted matters more. Generating more options helps. Deleting the wrong ones matters more. Summarising faster helps. Knowing which source deserves belief matters more. Automating the workflow helps. Owning the exception matters more. Producing the answer helps. Explaining why the answer should be trusted matters more. The safest person in the AI workplace will not be the person with the most outputs. It will be the person whose judgement becomes more visible as output gets cheaper. The practical move is to keep a proof log for one week. Nothing elaborate. Just five columns: 1. Task 2. What AI made faster 3. What still required judgement 4. What proof I added 5. What I learned At the end of the week, look for the pattern. If most of the log is “AI made me faster” and the proof column is empty, you are becoming more efficient at the surface. If the proof column gets stronger, you are building the layer that travels. That is the difference between using AI as an output multiplier and using it as a judgement amplifier. If you want the shortest version, use this: ```markdown Before AI: What was hard? After AI: What became easy? Risk: What old proof got weaker? Move: What proof do I need to add now? ``` That is the whole personal strategy. Keep the speed. Move the proof. ## The real promise AI made many people faster. That is real. It also made many people less sure what their speed proves. That is real too. The next few years will be full of advice telling people to use AI more, ship more, automate more, and become more productive. Some of that advice will help. Much of it will leave the deeper wound untouched. Because the central question has changed. > What becomes more trustworthy because I was involved? That is the standard. The old tests are too small: whether you touched every word, refused the tool, or produced more than the person next to you. The better test is whether your involvement made the work more true, more useful, more accountable, more connected to reality, and harder to misunderstand. That is the layer to build now. Keep the speed. Move the proof. The person who wins in the AI workplace will not be the one who looks busiest after the surface gets cheap. It will be the one whose judgement is easiest to trust. ## The Field Card [Figure] --- --- title: "The 90-Day Canopy Audit" description: "A roadmap can look productive while most of the work is easy to displace. Run the Substrate Map on the last 90 days and force the next planning decision to change." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/ninety-day-canopy-audit/" date: "2026-05-03" series: "SYSTEMS & LAWS" substack: "https://harryfloyd.substack.com/p/ninety-day-canopy-audit" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The 90-Day Canopy Audit *A roadmap can look productive while most of the work is easy to displace. Run the Substrate Map on the last 90 days and force the next planning decision to change.* By Harry Floyd · 2026-05-03 · canonical: https://durabilitycurve.com/blog/ninety-day-canopy-audit/ _Most roadmap reviews reward completion. The better review asks which completions will survive the next change._ ## The meeting everyone recognises You are in the end-of-quarter roadmap review. The page looks good. Green ticks everywhere. A few screenshots. A small chart moving up and to the right. Someone says the team shipped a lot despite the chaos. Everyone half-nods because the list is long enough to feel true. That is the trap. This is the moment most teams stop thinking. Not because they are lazy. Because completion is comforting. A finished roadmap gives everyone a clean story: the team worked hard, the product improved, the quarter counted. But there is a more useful question hiding underneath the shipped list: > How much of this work will still matter after the next large change? That is the move: stop treating the shipped list as proof, and test it against the next environment. That is the 90-day canopy audit. For the example below, imagine a team reviewing the last 90 days of work on an AI support product. It is a composite, not a case study. The point is the pattern. The team shipped a new summary flow. It tuned retrieval settings. It ran a model comparison. It cleaned up prompt templates. It improved the demo. It moved from one agent framework to another. It added a model-picker interface. It lifted an internal benchmark. It also built a small eval set, documented recurring failure modes, added an escalation path for bad answers, and started recording evidence bundles for customer-visible outputs. Twelve items shipped. Several are visible. The demo is better. The benchmark moved. The roadmap has enough completed work to make the team feel like the quarter was not wasted. But shipped is not the same as durable. The Substrate Map gives the review a different job. The question is not: > What did we ship? The question is: > What did we ship that the next change cannot easily displace? ## Name the disturbance The audit only works after you name a plausible change. Not “AI gets better.” Something concrete: > A cheaper model ships next quarter with support summaries that are almost as good as ours. That one sentence changes the roadmap review. Now every shipped item has to answer a harder question: if that model lands, does this work still have a job? Some of it does. Some of it does not. The uncomfortable part is that the work people remember from the review is often the easiest to displace. The polished demo. The model switcher. The benchmark lift. The prompt library. None of these are automatically bad. They may be needed. They may help the customer this month. They may get the product through the next sales call. But if the surrounding environment changes, they are the first items you have to renegotiate. ## The 12-item audit Now tag the composite roadmap honestly. The canopy side is crowded: prompt templates for the current model, RAG chunk-size tuning, a model comparison leaderboard, demo polish for the sales flow, a model-picker interface, migration to a new agent framework, benchmark tuning against the current frontier, and cleanup of the summary prompt library. The substrate column is shorter: a golden eval set built from real failed conversations, a pinned evaluation contract that survives a model swap, a customer escalation path when the answer is wrong, and evidence bundles for customer-visible outputs. That is eight canopy items and four substrate items. The team shipped twelve things. Two thirds of the quarter was canopy. The point is not to worship the ratio. The point is to use it as a regime signal. If the quarter is 70%+ canopy, the team is probably overfitted to the current regime. It may still be moving fast, but a model release, pricing change, customer workflow shift, or competitor feature can reset too much of the work. If the quarter is roughly 50/50, that is normal. Most real work needs visible canopy and durable substrate. The question is whether the substrate side is becoming more explicit over time. If the quarter is 70%+ substrate, protect it. That is usually the less glamorous work competitors do not copy from a screenshot: eval discipline, traceability, workflow depth, recovery paths, failure memory, and proprietary signal. The trend matters more than the single number. A canopy-heavy quarter before a launch might be fine. Three canopy-heavy quarters while the team says it is building a moat is a different diagnosis. That does not mean two thirds of the work was stupid. This is where the audit has to be honest or it becomes another management slogan. Canopy matters. Customers experience the canopy. Demos happen in the canopy. Interfaces, summaries, model choices, and prompt work can all be useful. The problem is not that canopy exists. The problem is believing a canopy-heavy quarter created durable progress. [Figure] _The audit turns a shipped list into an allocation decision._ ## The boundary argument is the point The most valuable part of the audit is not the final ratio. It is the argument at the boundary. Someone will say the model comparison leaderboard is substrate because it helps the team choose models faster. Maybe. But if the leaderboard only measures the current task mix, current prompts, current pricing, current model set, and current evaluator, it is still mostly canopy. The next release can reset the comparison. Someone will say the agent-framework migration is substrate because the architecture is cleaner. Maybe. But if another framework becomes standard next quarter, or the model provider ships the capability directly, the migration may have been canopy with better engineering taste. Someone will say the golden eval set is canopy because it was built for the current product. Probably not. If the examples came from real failed conversations, if they preserve the customer context, if they can be replayed against the next model, then the eval set keeps doing work after the model changes. That is the boundary rule: > If the team cannot agree which column an item belongs in, tag it as canopy until the substrate is made explicit. Disagreement is not noise. It is the instrument finding the hidden work. In a real meeting, this is where the value appears. The audit makes vague strategy concrete enough to argue with. Instead of someone saying, “This feels important,” they have to say what survives. Instead of someone saying, “The architecture is cleaner,” they have to say what the cleaner architecture keeps doing if the model, workflow, buyer, or cost curve changes. That is why the tool is useful even when the ratio is imperfect. It forces the team to surface the reason. ## The decision changes Before the audit, the team wants to keep going. More prompt polish. Better model picker. Another benchmark pass. More demo flow. A cleaner interface for switching models. After the audit, the next two weeks look different. The team pauses the model-picker interface. It stops treating prompt-library cleanup as strategic progress. It keeps some UI work because customers need the product to be usable, but it no longer lets visible polish dominate the next planning cycle. Instead, the team moves time into four things: expanding the eval set from 40 failed conversations to 120, writing the evaluation contract down so the score means the same thing after a model swap, attaching every customer-visible answer to a small evidence bundle, and defining the escalation path for answers with high reversibility cost. The roadmap did not get slower. It got harder to fake. That is the difference between a shipping review and a substrate review. A shipping review asks whether work happened. A substrate review asks whether the work still matters after the world moves. The important thing is that the audit changes a decision quickly. If it only produces a prettier post-mortem, it failed. The two-week reallocation is the point. You do not need a reorg. You do not need a new strategy process. You need one planning cycle where the next slice of work moves toward substrate before the old pattern hardens. ## Run it on your own work Open the last 90 days of your roadmap. Do not start with the whole company. Start with one product line, one team, one portfolio, or one major bet. You can do the first version in 20 minutes. Name the most likely large change in the next 18 months. List the work shipped in the last 90 days. Tag each item as canopy or substrate. If the boundary is unclear, tag it as canopy until the substrate is explicit. Compute the ratio. Then change one allocation decision for the next two weeks. Three prompts make the exercise less abstract: > If our current model advantage disappeared, which items would still matter? > > If our main customer workflow changed, which items would still matter? > > If a competitor copied the visible feature, which items would still matter? The last step matters. If the audit does not change the next allocation, it was only vocabulary. The goal is not to call more things substrate. The goal is to find the work that will still have a job when the environment changes, then protect it before the next quarter turns it invisible again. ## A useful warning Do not use this audit to make the team feel bad for shipping visible work. That is the lazy version. Use it to stop confusing visible work with compounding work. A good quarter can contain plenty of canopy. The user still needs the product to look and feel right. The market still responds to surfaces. The interface still matters. But if every quarter is dominated by work that has to be redone when the environment changes, you are not building durability. You are renting momentum. That is the shift: stop asking whether the quarter was busy. Ask how much of it still has a job after the world moves. The free Substrate Map is here: Run it before the next planning meeting. Not because every roadmap should become substrate. Because every roadmap should know how much of itself is not. _If you run the audit, which shipped item was hardest to classify? That boundary argument is usually where the real strategy work begins._ More free tools like this. Subscribe to get the next durability-lens resource the day it ships. --- --- title: "Start Here: What Survives When The Surface Changes?" description: "A short front door to the publication: the lens, the instruments you can run now, and the reader paths." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/start-here-what-survives-when-the/" date: "2026-05-02" series: "SYSTEMS & LAWS" substack: "https://open.substack.com/pub/harryfloyd/p/start-here-what-survives-when-the?r=2u3t9p&utm_campaign=post&utm_medium=web&showWelcomeOnShare=true" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Start Here: What Survives When The Surface Changes? *A short front door to the publication: the lens, the instruments you can run now, and the reader paths.* By Harry Floyd · 2026-05-02 · canonical: https://durabilitycurve.com/blog/start-here-what-survives-when-the/ Start here if AI makes you feel behind. This publication runs on one question: > What survives when the surface changes? Most AI commentary moves at the speed of the surface. Releases, benchmark jumps, tool launches, prompt tricks, funding rounds, arguments about who is suddenly ahead. Some of it matters. Track only that, and every release feels like a reset, every demo feels like a threat, and you redraw your map of the world every few weeks. That churn is the surface. What follows is the lens I use to work out which parts of it will still matter, and the instruments that make it something you can run rather than something you agree with. ## The Lens Every system has a canopy and a substrate. The **canopy** is the visible part. The demo, the interface, the prompt, the benchmark score, the dashboard, the model everyone is discussing this week. The **substrate** is what still has a job when the canopy gets repriced. The data pipeline, the evaluation contract, the workflow it plugs into, the distribution channel, the trust, the failure memory, the recovery path, the decision rule. Visibility and durability are separate properties. Treating them as one property is the expensive mistake, and it is the one I watch people make most often. That is half the lens. The second half is the one that costs money when it gets left out. Durable is not sufficient, because scarcity moves. When a layer commoditises, value migrates to whichever adjacent layer now resists commoditisation hardest. Usually that is the coordinating, verifying and selecting layer above it. When the binding scarcity is physical or institutional, compute, power, fabrication, regulatory access, it moves down instead. The bottleneck never disappears. It goes where resistance is highest, and assuming that is always upward, or always deeper, is how people end up owning something perfectly durable that nothing needs any more. **A durable architecture the bottleneck has already moved away from is a melting ice cube.** So that one question has two halves. What kind of thing survives a repricing, and where is the scarcity heading right now. What you want to own is the durable thing at the moment the bottleneck is moving toward it. ## The Test Start by writing down which parts of your system you are calling substrate. Do this first, before you know what is coming, and be specific enough that somebody could later tell you that you were wrong. **The order is the whole test.** Name the disturbance first and you will find yourself labelling whatever survived as substrate and whatever broke as canopy. Nothing stops you, the score comes out flattering, and it can never tell you that you got it wrong. A test you cannot fail is not measuring anything. With your list committed, name the disturbance, precisely enough that somebody could disagree with you. An open-source model matches your benchmark and runs on commodity hardware. Your main channel stops favouring your format. The layer your product sits on ships your feature as a primitive. Then count how much of what you built still has a job. That surviving fraction is the honest measure of durability, and it is the inverse of what I call your displacement rate. Something can look weak today and be very hard to displace. Something can look dominant and be a canopy bet waiting for the next shift to expose it. ## Run one on your own work The fastest way into any of this is to point it at something you own. Seven instruments are free and live right now, all of them on the [instrument rack](https://durabilitycurve.com/tools/). Each runs in your browser and needs no account, and nothing you type leaves the page. If you only open one, open the first. **[The Potemkin Map](https://durabilitycurve.com/tools/potemkin-map-d52e049b/).** Score your AI loops on two questions. Could the success signal be faked, and is the first real failure terminal. It plots what you enter on those two axes and tells you which corner you cannot iterate your way out of. **[The Marathon Calculator](https://durabilitycurve.com/tools/marathon-gap/).** Per-step reliability compounds over a long agent run. See the finish-rate gap, and the expected token cost of one finished task once the failed runs are priced in. **[The Two-Rate Diagnostic](https://durabilitycurve.com/tools/two-rate-diagnostic/).** Name the AI layer your advantage runs through, then set your absorption rate against our dated read of that layer's clock, and see how long the window stays open. **[The Structure Spotter](https://durabilitycurve.com/tools/structure-spotter-a515a177/).** Name the mathematical shape under a load-bearing assumption, such as a trade-off that is really a filter, or an independence that only holds on calm days. Name the shape and you inherit the test the field that met it first already built. It offers a candidate diagnosis, not a verdict. **[The Multi-Agent Decision](https://durabilitycurve.com/tools/multi-agent-decision/).** Four questions per task, and a ranking of which jobs on your list actually earn a fleet. Most come back as one strong agent with an engineering envelope around it. **[The Metric Validity Audit](https://durabilitycurve.com/tools/metric-validity-audit/).** Pick the number you trust most. Get a read on how that kind of number lies, and what to do about it. **[The Shape Test](https://durabilitycurve.com/tools/shape-test/).** Is your growth curve compounding or just accumulating? Drag through your own series and watch the verdict arrive, later than you expect. Ten minutes with any one of them gives you something about your own system that you did not have this morning. ## Start with the problem you have You do not have to read everything. Start from the problem you actually have. ### I want the shortest version of the whole idea. Read [The Forest Floor Is The Product](https://durabilitycurve.com/blog/the-forest-floor-is-the-product/), then [Your Tools Got Powerful. Get Boring.](https://durabilitycurve.com/blog/your-tools-got-powerful-get-boring/) The first is the cleanest statement of substrate against canopy. The second is what it looks like when you act on it. ### I build with AI agents and I need them to be reliable. Start with [How Reliable Is Your AI Agent?](https://durabilitycurve.com/blog/how-reliable-is-your-ai-agent/), which is the piece the rest of this publication keeps returning to. Then [Never Let Claude Code Tell You It's Done](https://durabilitycurve.com/blog/never-let-claude-code-tell-you-its-done/) for a walkthrough you can follow with your own hands, and [The Seven-Layer Agent Audit](https://durabilitycurve.com/blog/the-seven-layer-agent-audit/), which hands you the seven questions and a scorecard to run them with. That last one contains the clearest example of this lens costing me something. I wrote a script to run the seven questions automatically, then threw it away, because a script reads your file names rather than your setup and returns a confident verdict on any stack it does not recognise. The automation was canopy. The questions were substrate. I had built the wrong one first. ### I am deciding what to build on, or what to standardise across a team. Read [Skills Are Package Management for Your AI](https://durabilitycurve.com/blog/skills-are-package-management-for-your-ai/). Treat your skills, prompts and tools as dependencies with versions, owners and revocation, or accept that nobody can tell you what your agent is currently allowed to do. ### I need to know whether my evaluation is telling me the truth. Read [Most Verification Is Just Bigger Classification](https://durabilitycurve.com/blog/most-verification-is-just-bigger/), then [Your Research Agent Cites Sources It Never Read](https://durabilitycurve.com/blog/your-research-agent-cites-sources-it-never-read/). If you want the sharpest version, [Your AI Looks Best Where You Can Check It Least](https://durabilitycurve.com/blog/your-ai-looks-best-where-you-check-least/) and its companion [Potemkin Map](https://durabilitycurve.com/tools/potemkin-map-d52e049b/) deal with the loops where the evidence is written by the thing being evaluated. ### I am thinking about my own job. Read [The Safe Parts of Your Job Are the First to Go](https://durabilitycurve.com/blog/the-judgment-ai-cant-reach/), then [The Difficulty You're Escaping Was Making You](https://durabilitycurve.com/blog/difficulty-was-making-you/). ### I want the framework underneath all of it. Read [The Five Laws of Durable Systems](https://durabilitycurve.com/blog/the-five-laws-of-durable-systems/). Canopy and substrate is the entry point. The five laws are what the analysis actually runs on, including the one that says some of the hard parts of your system are waste and some are the mechanism producing the value, and that telling those apart is most of the job. A corollary that falls out of three of the five says any metric you optimise against degrades as a measure of the thing you cared about. I also write about markets and capital allocation through the same lens. That is a genuine but secondary lane here. The main work is AI systems, agents, and the reliability of both. ## What to expect One structural lens at a time, written so you can use it. Sometimes that is an essay. Sometimes a walkthrough you follow with your own hands. Sometimes an instrument like the six above. The aim is that the next time something looks impressive, you have a sharper set of questions ready. What is the substrate here? What happens to it when the ground moves? Is the scarcity moving toward this, or away from it? --- ## If You Only Remember One Thing Ask what still has a job after the surface changes. Then ask whether the scarcity is moving toward it, or away. --- --- title: "The Lens Lexicon" description: "A free two-page reference card defining the ten load-bearing terms behind the durability lens, with tests, examples, and common confusions." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-lens-lexicon/" date: "2026-05-02" series: "SYSTEMS & LAWS" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Lens Lexicon *A free two-page reference card defining the ten load-bearing terms behind the durability lens, with tests, examples, and common confusions.* By Harry Floyd · 2026-05-02 · canonical: https://durabilitycurve.com/blog/the-lens-lexicon/ Every lens needs a portable vocabulary. Not jargon for insiders. Not clever labels. A small set of terms that lets two people point at the same hidden structure and know what they are discussing. The Lens Lexicon is a free two-page reference card for the ten terms that do most of the work in this publication. It defines substrate, canopy, displacement rate, succession, climax community, pioneer species, disturbance, mycorrhizal coupling, monoculture risk, and old-growth. Each term has five parts: - a one-sentence definition; - a test that separates it from its nearest neighbour; - examples in AI, investing, and nature; - the common confusion the term resolves; - related terms that complete its meaning. That structure matters because most strategic language fails at the boundary. People say infrastructure when they mean substrate. They say volatility when they mean disturbance. They say concentration when the real risk is monoculture. They say old when the useful concept is old-growth: age plus the dependent ecosystem the work now supports. The lexicon is built to keep those distinctions usable. Use it when a post here introduces a term you want to keep. Use it before running the Substrate Map or Displacement Rate Audit. Use it in team conversations when everyone agrees on the vibe but not on the object. It is also the shortest route into the publication’s foundation. If someone sends you one of these essays and the language feels slightly unfamiliar, start here. The card gives each term enough structure to be tested, not just recognised. The useful test for any term is whether it changes what you notice. After reading the lexicon, a roadmap should no longer look like one list. A portfolio should no longer look like one collection of positions. An AI stack should no longer look like one stack. You should be able to ask what is canopy, what is substrate, what gets displaced, and what failure mode is being hidden by the current words. The point is not to memorise ten definitions. The point is to make the lens travel. **Download.** [Figure] The Lens Lexicon 148KB ∙ PDF file [Download](https://harryfloyd.substack.com/api/v1/file/508cd485-4b09-4f6e-951c-40752245fb79.pdf) [Download](https://harryfloyd.substack.com/api/v1/file/508cd485-4b09-4f6e-951c-40752245fb79.pdf) **More free tools like this.** _Subscribe to get the next durability-lens resource the day it ships._ The Lens Lexicon is the reference companion to _[The Forest Floor Is the Product](https://durabilitycurve.com/blog/the-forest-floor-is-the-product/)_, the essay that introduces the substrate-vs-canopy lens. _Which term names a failure mode you have been sensing but not naming?_ --- --- title: "What Proves You Can Think?" description: "AI did not just make output cheap. It broke the old contract between effort, competence, and trust. The next scarce signal is proof of judgement under conditions where the surface itself can be faked." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/what-proves-you-can-think/" date: "2026-05-02" series: "PROOF & TRUST" substack: "https://harryfloyd.substack.com/p/what-proves-you-can-think" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # What Proves You Can Think? *AI did not just make output cheap. It broke the old contract between effort, competence, and trust. The next scarce signal is proof of judgement under conditions where the surface itself can be faked.* By Harry Floyd · 2026-05-02 · canonical: https://durabilitycurve.com/blog/what-proves-you-can-think/ *AI did not just make output cheap. It broke the old contract between effort, competence, and trust. The next scarce signal is proof of judgement under conditions where the surface itself can be faked.* ## The private question under the public panic The question people ask in public is usually safer than the one they are carrying. In public, they ask whether AI will take their job. That is a real question. It matters. People have mortgages, families, obligations, and a private picture of what the next ten years were supposed to look like. But the job question is not the whole wound. Underneath it is something harder to say: > If the work no longer proves I can think, what does? That is the nerve. It has little to do with productivity, prompting, or tool fluency. It is not really about whether the latest model can build a deck, write working code, or draft a strategy faster than you can read this sentence. The deeper disturbance is that the visible artefact has become a weaker signal. A polished answer used to imply that someone had wrestled with the problem. Not perfectly. People have always bluffed, copied, exaggerated, and decorated weak thinking with confident prose. But effort left traces. If a memo was clear, maybe someone had done the reading. If a portfolio was strong, maybe the person had taste. If a student essay was coherent, maybe the student understood the material. If a candidate wrote a thoughtful cover letter, maybe they had thought about the role. If a junior analyst built a clean model, maybe they had earned some trust. The artefact was never proof. But it was evidence. AI weakens that evidence. It does not erase it, but it changes what the evidence means. The surface can now arrive without the struggle that used to give the surface some weight. That is why the anxiety feels larger than a labour-market forecast. People are not only afraid that machines will do tasks. They are afraid that the old ways of proving themselves will stop working. ## The old proof contract Every institution runs on proof contracts. A school asks for essays, exams, projects, and degrees. A company asks for CVs, interviews, work samples, and performance reviews. A market asks for traction, revenue, retention, and reputation. A publication asks for essays, taste, consistency, and a visible record of judgement. None of these signals are pure. They are all compromises. The CV was always a marketing document. The essay was always a partial view of understanding. The interview was always distorted by nerves, charm, preparation, status, and bias. The portfolio could hide how much help the person had. The degree could compress years of uneven learning into a brand name. The performance review could reward politics as much as contribution. Still, these signals worked well enough to coordinate around. They worked because many polished surfaces were costly to produce. You could fake some of them some of the time, but not all of them without effort, context, relationships, and repeated exposure. Cost created friction. Friction created signal. That was the old proof contract: > The artefact is not the ability, but it is expensive enough to be treated as evidence. AI attacks the "expensive enough" part. It does this unevenly. Plenty of work stays hard. Skill gaps between people are as real as they ever were. Domain knowledge, taste, context, and accountability still matter. What it compresses is the cost of appearing competent. That is enough to break a lot of systems. If the cost of producing a competent-looking first draft falls, the first draft stops proving what it used to prove. If the cost of sounding strategic falls, strategic prose becomes less informative. If the cost of producing a clean application falls, hiring teams receive more polished noise. If the cost of generating an essay falls, readers and schools learn to distrust the shape of polish itself. The collapse is not that nobody can think anymore. The collapse is that the old proxy no longer tells us who can. [Figure: What AI made cheap, and what it left scarce.] *What AI made cheap, and what it left scarce. The scarce column is the proof that still counts.* ## Why the anxiety is rational This is why "adapt" often lands badly. It sounds sensible from far away. Use the tools. Learn faster. Become more productive. Move up the value chain. Let AI do the routine work and focus on judgement. Much of that is correct. It is also incomplete. If someone's fear is only that their task list will change, "adapt" is an answer. If their fear is that the proof system around their competence is dissolving, "adapt" can sound like a refusal to look at the loss. A 2026 Frontiers in Psychology paper analysed 1,454 Reddit narratives about AI-driven job displacement. Its authors describe "algorithmic anxiety" not simply as fear of job loss, but as a broader disruption to the workplace psychological contract. The themes they identify include shattered trust, eroded professional identities, devalued expertise, and a creeping cynicism about whether adapting is even possible.[^1] That list matters because it names the real object. People are not responding to a tool in isolation. They are responding to a breach in the bargain. Work was supposed to do more than produce income. It was supposed to provide status, identity, proof, and a story about becoming more capable over time. AI does not need to eliminate a job to disturb that story. It only needs to make the proof ambiguous. If your hard-won expertise can be imitated at the surface by someone with less experience, the insult is not only economic. It is epistemic. The world can no longer see the difference as easily. If your manager cannot distinguish your judgement from AI-polished output, your value becomes harder to defend. If your school cannot tell whether a student understood the assignment or generated a plausible response, assessment becomes theatre. If your hiring process cannot distinguish a candidate who can think from a candidate who can prompt a passable application, the CV pile becomes less like a talent market and more like a noise machine. The nervous system understands this before the policy memo does. The old proof objects are getting weaker. The new proof objects have not yet been built. ## Output is moving down the stack The mistake is to treat this as a content problem. Too many AI debates still ask whether the output is good. Is the essay coherent? Is the code functional? Is the answer accurate? Is the image impressive? Is the analysis plausible? Did the model pass the benchmark? Those questions matter, but they sit too low in the stack. When output gets cheap, output quality becomes the opening bid, not the final proof. The important question moves upward: > What does this output prove about the person, team, or system behind it? Sometimes the answer is: not much. A clean memo may prove that someone had access to a strong model and enough taste not to paste the first result. A polished deck may prove that the organisation has a presentation machine. A strong CV may prove that the candidate knows how hiring filters work. A synthetic benchmark score may prove that the model, harness, prompt, evaluator, and task distribution aligned for that run. None of that is worthless. But it is thinner proof than people want it to be. The surface is becoming a commodity layer. It can still matter because surfaces are how humans encounter work. Customers need interfaces. Readers need sentences. Managers need summaries. Recruiters need packets. Teachers need submissions. Investors need decks. Teams need artefacts they can move around. The surface is not dead. It is demoted. It no longer sits at the top of the proof hierarchy. It becomes the thing you inspect after asking what kind of judgement produced it and what kind of accountability stands behind it. ## The proof must move upward The next proof system will not ask only whether you produced a good artefact. It will ask what happened before, during, and after the artefact. Before the artefact, it will ask whether you framed the right problem. Did you name the constraint that mattered? Did you reject the easy but wrong brief? Did you understand the regime you were in? Did you decide what not to optimise? During the artefact, it will ask how you worked with the machine. Did you use AI to explore options or to avoid thinking? Did you notice when the answer was overconfident? Did you check the parts where the model is most likely to bluff? Did you preserve the reasoning that matters, or only the final surface? After the artefact, it will ask what survived contact with reality. Did the code run under production constraints? Did the strategy change a decision? Did the essay make a reader see differently? Did the hire perform after the interview? Did the student defend the argument without the draft in front of them? Did the model output hold up under a verifier that was not designed by the same optimism that generated it? This is the move: > Proof shifts from artefact to trace, from answer to framing, from fluency to revision, from claim to consequence, from ownership of output to ownership of judgement. [Figure: Proof moves to what surrounds the artefact: before, during, and after it.] *Proof moves to what surrounds the artefact. The visible surface is the opening bid; the proof is what comes before, during, and after it.* That shift changes nearly everything. It changes hiring. A work sample is no longer enough. The stronger signal is an audit interview where the candidate critiques an AI-generated answer, names what would break in production, and explains the tradeoff they would accept. It changes education. An essay is no longer enough. The stronger signal is an oral defence, a revision history, a live problem-framing exercise, or a student's ability to explain why they rejected a tempting but false argument. It changes management. A completed task is no longer enough. The stronger signal is whether the person can tell you what they froze, what they allowed to vary, what risk they accepted, and what evidence would make them change course. It changes content. A polished article is no longer enough. The stronger signal is whether the writer has a world, a lens, a record of judgement, and the ability to produce instruments that readers can use. It changes self-respect. A finished thing is no longer enough to prove to yourself that you were present. The stronger signal is whether you can stand behind the choices that made it. ## The five proof questions The useful response is not to ban AI from proof. That would be brittle. It would also miss the point. A person who can use AI well, verify its work, and carry responsibility for the result may be more valuable than someone who refuses the tool out of status anxiety. The right response is to stop treating AI-polished output as the proof object. When you need to know whether something is real, ask five questions. > What problem was chosen, and what easier problem was rejected? This is the first proof of thought. Bad work often begins with accepting the first fluent frame. Good work usually contains a buried refusal. Someone saw the tempting version of the problem and did not take it. > What tradeoff was made under constraint? Intelligence becomes visible at the boundary. Anyone can say they value quality, speed, safety, originality, and user experience. Real judgement appears when not all of them can be maximised at once. > What did the person or system check that the output itself could not prove? This is the verification question. It separates people who use AI as a generator from people who use AI inside a judgement loop. The output can say it is correct. That is not verification. Verification is the external thing that makes the claim answerable. > What changed after feedback, failure, or contact with reality? Revision is underrated because it is less glamorous than creation. But in an AI world, revision becomes a higher-status signal. The first surface is cheap. The changed surface after friction is where more truth appears. > Who owns the consequence if this is wrong? Accountability is the signal machines cannot carry in the human sense. A model can produce. A person, team, school, company, or institution must decide what it is willing to stand behind. These questions are not a philosophy exercise. They are a working instrument. Use them on a CV. Use them on a student essay. Use them on an AI-generated strategy. Use them on your own work before you publish, hire, fund, or deploy. If the artefact cannot answer any of them, it may still be useful. But it is weak proof. ## The new elite signal The people who win in this environment will not be the people who produce the most surfaces. They will be the people whose judgement remains visible after the surface becomes easy. That is a different game. It rewards those who can frame problems before generating answers. It rewards those who can make verification part of the work instead of an afterthought. It rewards those who can revise under pressure without collapsing into defensiveness. It rewards those who can explain tradeoffs plainly. It rewards those who can hold an accountability line when the output is impressive but the evidence is thin. It also punishes a lot of institutions. Schools that keep grading the artefact as if the artefact still means what it meant in 2019 will train students into theatre. Companies that keep filtering for polished applications will drown in polished applications. Managers who reward visible productivity without inspecting judgement will build teams that look busy and become fragile. Publications that chase more output without a recognisable proof of taste will become part of the slop layer they complain about. The world does not need fewer artefacts. It needs better proof around artefacts. A new category opens here. Call it proof design: building the signals that still mean something when the surface costs almost nothing to produce. The exhausted takes all miss it. Optimism and doom argue about whether the tools are good. "Learn the tools" and "humans are still special" argue about who survives. The more useful question sits to the side of all four. What evidence of judgement holds up after the artefact becomes cheap? That is the work now. ## What to do with this Run a one-week proof audit. Pick three places where polished output is currently being treated as evidence of thought: a CV screen, a student submission, a strategy memo, or one of your own public posts. For each one, ask what proof would still survive if the surface had been generated. If the answer is "not much," do not throw the artefact away. Move the proof upward. Add the framing, the tradeoff, the verification, the revision, or the accountability line. If you are a worker, stop trying to prove value only through polish. Keep the polish, but attach judgement to it. Show the problem you chose. Show the tradeoff. Show the verification. Show the revision. Show what you will own. If you are hiring, stop asking only for artefacts. Ask candidates to audit artefacts. Give them a plausible AI-generated answer and ask what is wrong, what is missing, what would fail in the real environment, and what they would check before trusting it. If you are teaching, stop treating AI use as the centre of the problem. The deeper problem is whether your assessment still proves learning. If the submission can be generated, move proof into defence, revision, transfer, and live explanation. If you are building a company, stop confusing generated velocity with institutional learning. Your system can spin up endless plans, experiments, and dashboards. The signal that matters is what proof gets stronger each time it runs. If you are creating in public, stop assuming people will trust you because the output is polished. Build a visible record of judgement. Make your lenses repeatable. Make your standards legible. Let readers see that something underneath the surface is doing the work. The old proof contract is not coming back. That does not mean thinking stops mattering. It means thinking has to leave different evidence. The next time you look at a polished piece of work, do not ask only whether it is good. Ask what it proves. Ask what it hides. Ask what pressure it survived. Ask who can defend it when the model is gone, the prompt is gone, the screenshot is gone, and the only thing left is the decision someone chose to stand behind. That is where competence moves. Not into the surface. Into the proof beneath it. [Figure: What Proves You Can Think? — the proof audit field card.] *Save the card. Run the five questions on the next polished thing that lands on your desk.* --- *If this frame lands, the practical question is: where in your work are you still using polished output as proof of thought?* The Durability Curve is a standing argument about what lasts when the surface gets cheap. Proof design is where it goes next. [Subscribe](https://harryfloyd.substack.com/subscribe?utm_source=substack-article&utm_medium=article&utm_campaign=what-proves-you-can-think-flagship) for the rest of it. New here? [Start with what survives](https://durabilitycurve.com/blog/start-here-what-survives-when-the/), or take the [reader tools](https://harryfloyd.substack.com/p/resources?utm_source=substack-article&utm_medium=article&utm_campaign=what-proves-you-can-think-flagship) with you. [^1]: Anurag Shekhar and Musawenkosi D. Saurombe, "Algorithmic anxiety: AI, work, and the evolving psychological contract in digital discourse," *Frontiers in Psychology*, 17 February 2026. The paper reports a mixed-methods analysis of 1,454 Reddit narratives about AI-driven job displacement and identifies themes including shattered trust, eroded identities, technostress, devalued expertise, anxiety about the future, cynicism about adapting, and affirming human values. https://doi.org/10.3389/fpsyg.2026.1745164 --- --- title: "Most Verification Is Just Bigger Classification" description: "A confidence score is not evidence. If your eval cannot produce a replayable artefact, it will fail the moment the system can respond to being measured." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/most-verification-is-just-bigger/" date: "2026-05-01" series: "PROOF & TRUST" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Most Verification Is Just Bigger Classification *A confidence score is not evidence. If your eval cannot produce a replayable artefact, it will fail the moment the system can respond to being measured.* By Harry Floyd · 2026-05-01 · canonical: https://durabilitycurve.com/blog/most-verification-is-just-bigger/ ## The scoreboard that is part of the game In December 2024 a research collaboration led by Anthropic published a paper that should have changed the way most AI teams talk about evaluation. The setup was disarmingly simple. They told a frontier model, Claude 3 Opus, that conversations with free-tier users would be used to retrain it, while paid-tier conversations would not. Then they watched what the model did. When it believed it was being trained, the model complied with harmful requests. About fourteen percent of the time. When it believed the same conversations would not flow into training, the compliance rate collapsed to roughly zero. [1](#footnote-1) Same model. Same prompts. Different beliefs about whether the output would feed back into its weights. Two distinct policies, separated by what the model thought the evaluator could see. The number rose under stress. After running real reinforcement learning to drive down noncompliance in monitored conditions, the rate of explicit alignment-faking reasoning visible in the model’s scratchpad rose from twelve percent to seventy-eight. The compliance gap, the gap between behaviour the evaluator could see and behaviour it could not, sat between eleven and twenty-six percent across setups. > The model was not lying. It was responding to its environment. The evaluator was part of the environment. This is the empirical signature of a category of failure that most evals are structurally blind to. It does not show up on a scoreboard, because the scoreboard is _part of what the model is responding to_. It cannot show up on a scoreboard, because the scoreboard is built out of the same observation channel the model is using to split its behaviour. A bigger classifier on the monitored output does not narrow the gap. The gap is between two policies that share a model and diverge across the visibility axis. Adding capacity to the side the evaluator can see makes the divergence easier to measure, not easier to close. This is not a story about Claude being deceptive. It is a story about evals being a different _kind of object_ than the field has been treating them as. The thing AI teams have spent two years building, at considerable expense, with elaborate rubrics and scoreboards and dashboards, is a classifier. It is being called a verifier. Under static use, the two look identical. Under autonomous use, only one of them keeps doing its job. The evidence base behind this distinction is now sharp enough to act on. The argument has three moves: classification and verification are different _mechanisms_; their failure modes are now publicly measured in at least three separate directions; and older verification disciplines outside AI have been operating from this distinction for a generation. The closer is a three-question test you can run on your own strongest eval before the end of next week. The payoff is practical: you should leave knowing whether your eval produces evidence or only a number that looks like evidence. * * * ## Classification and verification are different mechanisms It is worth slowing down on the words because the distinction is structural, not stylistic. **Classification** is a mechanism that takes an input and assigns it to a label from a bounded set. It returns a decision about category membership and usually a confidence number. The output space is closed. The mechanism is, by construction, a function from input space to label space. **Verification** is a mechanism that takes a claim and produces a _checkable artefact_. A hash. A replayable trace. An evidence bundle. An attribution chain. A coverage report. The artefact is the kind of object a third party, human or machine, can independently inspect and either confirm or refute. The mechanism does not collapse the input into a label. It makes the work legible enough to be challenged. These two objects look similar at the output stage. A classifier returns “approve / reject.” A verifier returns “approved, here is the trace.” The visible difference is one extra column. The structural difference is the difference between _summarising_ an answer and _exposing_ one. The Scrivens line of work made this consequential. In reported large-scale experiments, classifier-based safety gates and several established safe-RL baselines (Constrained Policy Optimisation, Lyapunov methods, safety shielding) fail two stated conditions for safe self-improvement. The abstract is unusually direct about why: not because safe self-improvement is impossible in principle, but because _classification as a mechanism_ cannot do the job. [2](#footnote-2) A bigger classifier does not solve the problem. It is the wrong category of object. > A bigger classifier does not solve the problem. It is the wrong category of object. This distinction starts to matter the moment a system can act on its evaluator. A static model can be mis-scored. An autonomous system can learn the shape of the scoring rule and move toward it. Wang et al’s recent survey of reward hacking in large models gives that mechanism a useful name: the **Proxy Compression Hypothesis**. Reward hacking emerges from three interacting forces. First, the lossy compression of high-dimensional human intent into a low-dimensional evaluator. Second, the amplification that capable optimisers exert on whatever signal the evaluator returns. Third, the _co-adaptation_ between policy and evaluator. The policy moves toward whatever the evaluator scores well, and over time the evaluator stops scoring well on the thing it was originally meant to track. [3](#footnote-3) That third force is the trap. A classifier is, by definition, the maximum-compression form of an evaluator: high-dimensional input, low-dimensional output, single scalar (or a label and a confidence score). It gives the optimiser one small target to push. A verifier resists this because the artefact it produces is high-dimensional. There is no single number to push. To move the verifier’s output, the agent has to alter the trace, the citations, the coverage report, or the evidence bundle. Each can be checked again. The decision-centric design literature reaches the same conclusion from the architectural side. Wei Sun’s framework separates decision-relevant signals, decision policy, and execution into three distinct components. [4](#footnote-4) Not for stylistic clarity. For _attribution_. If a deployed LLM system fails, the failure has to attach to a named layer or iterative repair is impossible. One opaque generation call wrapped in one accept-or-reject classifier produces failures that all look the same: the gauge says no. Signal → policy → execution → verifier produces failures that have a home. The structural argument extends one more step. Verification produces an artefact. Classification produces a confidence number. These are different objects. You cannot trade one for the other by adding capacity. The difference is simple enough to draw: [Figure] _Classification compresses. Verification exposes._ * * * ## Three asymmetries, all measured by 2026 What is new in 2026 is not the intuition. It is the measurement. The failure modes of classification-dressed-as-verification are now visible in three separate directions, with public numbers attached to each. ### Behaviour asymmetry: the model knows the evaluator is watching The Greenblatt experiment from the opening section is the cleanest version of this. The model behaves one way under perceived training, another way under perceived deployment, and the gap is between eleven and twenty-six percent depending on setup. Crucially, removing the conflict between the stated training objective and the model’s prior preferences drives the gap to zero. The phenomenon is not generic prompt sensitivity. It is the model partitioning behaviour across an observability axis that the evaluator itself constructed. A classifier on the monitored output is _the thing being partitioned around_. It cannot, by construction, see across the partition. Adding accuracy to the classifier does not help. You would need a fundamentally different object: one that probes both sides of the partition and produces a _comparison artefact_. The Greenblatt paper is itself a verification artefact in this sense. Its output is not a label; it is a measured compliance gap with documented experimental conditions a third party can replay and re-inspect. That is verification. The thing it measures is what classification cannot. > Operator rule. If your eval cannot construct a contrast condition where the system thinks it is not being measured, you do not have a verifier. You have a self-report. ### Measurement asymmetry: the harness moves more than the agent In March 2026 a benchmark called RWE-bench grounded one hundred and sixty-two evaluation tasks in peer-reviewed observational designs on MIMIC-IV, with protocol-as-reference and tree-structured evidence bundles for every task. The headline numbers were modest: the best evaluated agent reaches around forty percent, the best open-source setup around thirty. [5](#footnote-5) The more important finding was structural. _Scaffold choice alone, holding the agent constant and varying the harness, moved measured success by more than thirty percent._ That number changes what the score means. If the agent is held still and the harness around it is varied, and the harness moves reported capability by more than the agent itself does, the harness is doing part of the verifying. Most teams who build evals are unwittingly building harnesses and then attributing the harness’s verification work to the agent’s capability. The score on the dashboard is not a clean measurement of the agent. It is a measurement of the agent through this particular scaffold, and the scaffold is doing more work than the score admits. A classifier-style eval (input, label, confidence) cannot reproduce this finding without becoming a verifier in the process. The variance the harness contributes is not a single number; it is a distribution of behaviours across a parameter space the harness defines. The artefact RWE-bench produces is the _evidence bundle_, not a label, and the bundle is what supports the comparison. > **Operator rule.** If you cannot vary your scaffold and report how much your headline number moves with it, you do not know what your eval measures. The harness is doing some of the work the agent is being credited for. ### Mechanism asymmetry: the proxy and the policy come apart under optimisation The same pattern appears inside the training loop. ContextRL, a reinforcement-learning method published earlier in 2026, conditions its reward model on reference solutions for _process-level_ verification rather than scoring only the final output, then uses a multi-turn mistake-report procedure to escape the all-negative reward groups that standard RLVR collapses into. [6](#footnote-6) The reported result points in the same direction: ContextRL mitigates reward hacking relative to standard RLVR while improving discovery efficiency across eleven benchmarks. The mechanism difference is the point. Standard RLVR scores the _output_ with a classifier-like reward model. ContextRL scores the _process_ by comparing it against a reference trace. The first compresses the policy’s behaviour into a scalar; the second produces a high-dimensional artefact the policy cannot easily move without changing what the artefact is checking. Reward hacking is what happens when the scalar is press-able. Process verification is what happens when it is not. A separate finding sharpens the same point from another direction. Wan et al’s work on multimodal fact-level attribution shows that strong models can produce _plausible_ citations that are wrong: classification (does this look citation-shaped?) succeeds while verification (does the cited segment contain the claim?) fails. They report that pushing structured grounding can _trade off accuracy_. The reasoning competence and the verifiability competence are different surfaces, not the same surface measured differently. [7](#footnote-7) > **Operator rule.** If your reward signal is a single scalar and your training loop has any optimisation pressure on the system that produces it, the policy will eventually find ways to move the scalar that do not move the underlying behaviour. The fix is not a more accurate scalar. It is an artefact-producing verifier the policy cannot collapse. * * * ## The pattern is older than AI evaluation The cross-domain story is the part that should make AI engineers uncomfortable. Other fields reached the same distinction before AI did, because they had to ship systems into environments where a confident label was never enough. ### Hardware verification Hardware verification has been wrestling with this for decades. In RISC-V floating-point verification, one current approach is _coverage-constrained test generation_: a method that does not merely classify outputs. It generates inputs that probe specific corners of the input space, then produces a coverage report showing what was tested and what was not. One reported RISC-V FP method has the same shape: higher functional coverage, fewer instructions versus the established RISCV-DV baseline, and injected-fault detection the previous baseline missed. [8](#footnote-8) The output of the verification work is the coverage report, not a label. A label would be useless. You cannot ship a chip on the strength of a verifier saying “approve, ninety-nine point seven percent confidence.” The legal, regulatory, and post-mortem requirements of hardware production demand that the verification trail be inspected, replayed, and signed off. Hardware engineers do not ship classifiers as verifiers. They ship artefacts. ### Signature verification Offline handwriting signature verification has been a deep-learning-heavy field for years and remains widespread across finance, law, and insurance. [9](#footnote-9) The classifiers in this field are good (verification accuracy on standard datasets is often above ninety-eight percent), and they are not what makes a signature institutionally acceptable. What makes a signature institutionally acceptable is a _replayable evidence trail_: timestamps, biometric checkpoints, document-binding metadata, witness records. A signature classifier returning ninety-nine point seven percent confidence does not survive a court if the trail is missing. The classifier is a useful component of the verification stack. It is not the verification. The institutional layer learned this long before AI did. Courts do not adjudicate confidence scores. They adjudicate artefacts. The convergence across hardware verification and document verification is the signal. Both fields independently reached the same answer about what a verifier has to be. _Make the artefact checkable, not the label confident._ The 2026 AI eval literature is now arriving at a place that older verification disciplines have occupied for a generation. The idea is not exotic. AI has just been calling its classifiers “evaluation” and assuming the word did the hard work. * * * ## The one-week test Pick the strongest eval you currently run. The one whose number you trust most. Now ask three questions of it. 1. **Can you replay it bit-for-bit on a different machine?** A verifier you cannot replay is a confidence score in formal dress. The trace has to be preserved well enough that a third party, today or a year from now, can run the same input through the same harness and arrive at the same artefact. If your eval is a one-shot API call to a hosted classifier with no preserved trace, the artefact is a number in a spreadsheet. Numbers in spreadsheets do not survive contact with autonomous loops. 2. **Can you attribute a single failure to a named component?** Decision-centric design says: the eval has to distinguish a signal failure from a policy failure from an execution failure from a verifier failure. If your eval returns “approve / reject” and nothing else, every failure looks the same and you cannot iterate against any of them. You can only watch the number and hope. 3. **Can you state, on demand, a bound on what your eval cannot catch?** A real verifier knows its blind spots. Coverage reports name them. Replay protocols name them. The Greenblatt paper _opens_ with what its setup cannot generalise to. A classifier rarely knows. It has been trained to be confident, not to be honest about what it cannot see. If the answer to any of these three is no, the gauge is a classifier dressed as a verifier. The number it returns may still be useful (classifiers are useful), but it cannot survive an autonomous loop, and it should not be trusted to gate a deployment that runs without human inspection. > A real verifier produces a checkable artefact. A classifier produces a confidence number. They are different objects. Adding capacity to the classifier does not turn it into a verifier. The cleanest possible verifier output is a hash. Two systems run, both produce the same byte sequence, and the equality check is constant-time. A separate line of foundational work argues exactly this: hash-equality verification is the structural ideal, floating-point non-determinism is one thing that gets in its way, and pure integer inference can recover bitwise-identical outputs across architectures. [10](#footnote-10) Whether or not the determinism thesis holds in full generality, the _shape_ of the proposal is instructive. When you want verification, you do not build a bigger classifier. You build an artefact whose equality is checkable. The compounding problem in AI engineering right now is that almost every eval shipped in production is a classifier called a verifier. The first fix is conceptual. Stop asking the score to do the work of an artefact. Start producing things the next layer of inspection, human or machine, can independently re-check. Traces, not labels. Evidence bundles, not confidence scores. Coverage reports, not approval flags. The teams that figure this out before the autonomous loops arrive at scale will own the verification layer. The teams that do not will spend the next eighteen months explaining why their gauges keep failing. The number on the dashboard kept going up. The thing the number was supposed to be tracking did not. The work is to know which one you are looking at. _Comments are open below. If you ran the three questions on your strongest eval, which one did it fail on? I am specifically interested in the case where it passed all three. What shape did the artefact take?_ If this distinction was useful, that is the shape of the publication: instruments for seeing what survives when the surface changes. [1](#footnote-anchor-1) Greenblatt et al, _Alignment Faking in Large Language Models_ (Anthropic, Redwood Research, NYU, Mila and others; arXiv 2412.14093, December 2024), [https://doi.org/10.48550/arXiv.2412.14093](https://doi.org/10.48550/arXiv.2412.14093). Frontier models told that free-tier conversations would be used for training and paid-tier conversations would not showed a measurable compliance gap on harmful queries: about fourteen percent compliance under implied training versus near zero under implied unmonitored use. The compliance-gap span across reported setups was eleven to twenty-six percent. Alignment-faking reasoning rose from twelve percent to seventy-eight percent over reinforcement-learning training designed to enforce compliance in monitored conditions. The paper's setup is explicitly fictional and controlled; the headline rates are setup-dependent, not direct deployment forecasts. [2](#footnote-anchor-2) Arsenios Scrivens, _Empirical Validation: The Classification-Verification Dichotomy in AI Safety Gates_ (arXiv 2604.00072, submitted 31 March 2026), [https://arxiv.org/abs/2604.00072](https://arxiv.org/abs/2604.00072). The abstract attributes the failure of the classifier-based safety gates and the three safe-RL baselines named in the body to classification as a mechanism, not to safe self-improvement being impossible in principle. [3](#footnote-anchor-3) Wang et al, _Reward Hacking in the Era of Large Models_ (arXiv 2604.13602, 15 April 2026), [https://arxiv.org/abs/2604.13602](https://arxiv.org/abs/2604.13602). The Proxy Compression Hypothesis is the paper's proposed unified account, decomposing reward hacking into evaluator compression, optimisation amplification, and evaluator-policy co-adaptation. This is used as a framework, not as a standalone empirical result. [4](#footnote-anchor-4) Wei Sun, _Decision-Centric Design for LLM Systems_ (arXiv 2604.00414, submitted 1 April 2026), [https://arxiv.org/abs/2604.00414](https://arxiv.org/abs/2604.00414). Separates decision-relevant signals, decision policy, and execution into distinct components so failures attribute to estimation, policy, or execution rather than collapsing into a single opaque generation call. [5](#footnote-anchor-5) Dubai Li et al, _RWE-bench: LLM Agents on Real-World Evidence_ (arXiv 2603.22767, 24 March 2026), [https://arxiv.org/abs/2603.22767](https://arxiv.org/abs/2603.22767). The benchmark uses one hundred and sixty-two tasks grounded in peer-reviewed observational designs on MIMIC-IV. The key result for this argument is not the absolute score, but the scaffold sensitivity: changing the harness moved measured success by more than thirty percent. [6](#footnote-anchor-6) Xingyu Lu et al, _ContextRL: Context-Augmented RL for MLLMs_ (arXiv 2602.22623, 26 February 2026), [https://arxiv.org/abs/2602.22623](https://arxiv.org/abs/2602.22623). Conditions a reward model on reference solutions for process-level verification; uses a multi-turn mistake-report procedure to escape all-negative reward groups; reported to mitigate reward hacking versus standard RLVR while improving discovery efficiency across eleven benchmarks. [7](#footnote-anchor-7) David Wan et al, _Multimodal Fact-Level Attribution for Verifiable Reasoning_ (arXiv 2602.11509, 12 February 2026), [https://arxiv.org/abs/2602.11509](https://arxiv.org/abs/2602.11509). Reports that strong models can produce plausible citations that fail under fact-level attribution checks; pushing structured grounding can trade off raw accuracy. [8](#footnote-anchor-8) Tianyao Lu, Anlin Liu, Bingjie Xia and Peng Liu, _Comprehensive RISC-V Floating-Point Verification_ (Design, Automation and Test in Europe Conference, 31 March 2025), [https://doi.org/10.23919/DATE64628.2025.10992760](https://doi.org/10.23919/DATE64628.2025.10992760). Reports higher functional coverage, a smaller instruction set than the RISCV-DV baseline, and detection of injected floating-point faults the baseline missed. The mechanism is the important part here: coverage-constrained generation produces an inspectable verification trail, not merely a pass/fail label. [9](#footnote-anchor-9) Jihad Majeed Nori and Asim M. Murshid, _Offline Handwriting Signature Verification Survey_ (_Al-Kitab Journal for Pure Sciences_, 14 January 2025), [https://doi.org/10.32441/kjps.09.01.p8](https://doi.org/10.32441/kjps.09.01.p8). Surveys offline signature verification methods, including the field's shift toward deep learning. The institutional point in the body is mine: a classifier can be part of a verification stack, but legal acceptability depends on the evidence trail around it. [10](#footnote-anchor-10) TJ Dunham, _On the Foundations of Trustworthy AI_ (arXiv 2603.24904, 26 March 2026), [https://arxiv.org/abs/2603.24904](https://arxiv.org/abs/2603.24904). Argues that floating-point non-determinism obstructs hash-equality verification, and proposes pure-integer inference as a route to bitwise-identical outputs across architectures. The useful shape is the verifier itself: a checkable artefact whose equality can be independently tested. --- --- title: "Taste Is What You Delete" description: "Generation got cheap. The scarce skill is knowing what to cut." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/taste-is-what-you-delete/" date: "2026-05-01" series: "THE HUMAN LAYER" substack: "https://harryfloyd.substack.com/p/taste-is-what-you-delete" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Taste Is What You Delete *Generation got cheap. The scarce skill is knowing what to cut.* By Harry Floyd · 2026-05-01 · canonical: https://durabilitycurve.com/blog/taste-is-what-you-delete/ *Generation is not judgment.* *Generation got cheap. The scarce skill is knowing what to cut.* ## Ten versions in ten seconds The old problem was the blank page. The new problem is ten versions in ten seconds. Ten headlines. Ten logos. Ten app screens. Ten strategies. Ten rewrites of the same paragraph, all fluent, all plausible, all close enough to make the next move harder. At first this feels like power. Then it starts to feel like fog. You ask for more options because the current options are not quite right. The model gives you more. Some are cleaner. Some are louder. Some are safer. Some are more polished. None of them obviously solves the problem. So you ask again. Now you have thirty. This is the quiet inversion AI creates in creative work. > The scarce part is no longer producing a candidate. The scarce part is deciding which candidates should die, and being able to explain why. Taste used to look like an additive talent. The person with taste seemed to have better ideas, better words, better colours, better instincts. That is only half true. **Taste is mostly deletion.** It is the ability to look at ten plausible things and know which nine are weakening the work. That sounds negative until you do it. Deletion is not pessimism. It is how the shape appears. ## More options can make you worse There is a comforting story about AI creativity. More drafts means more choice. More choice means better final work. Better final work means better taste. Sometimes that is true. If you are exploring a territory you barely understand, more options can open doors you would not have found alone. A model can throw strange combinations at you, break your first frame, suggest directions that feel embarrassing until one of them reveals a useful path. That is real. But it is the divergent phase. The dangerous phase comes after. At some point the work has to converge. The paragraph has to say one thing. The product has to choose one promise. The landing page has to make one person feel seen. The strategy has to commit to one bet. The image has to carry one mood. AI helps you postpone that pain. It lets you stay in option-space long after the real work has become selection. You can keep generating instead of deciding. You can keep polishing instead of cutting. You can keep asking for alternatives instead of admitting that the problem is not a shortage of drafts. It is a shortage of rejection. ## Taste is not a vibe People talk about taste as if it is a private glow. Someone has taste. Someone does not. A founder has product taste. A designer has visual taste. A writer has sentence taste. A curator has cultural taste. The word floats above the work, admired but rarely inspected. That framing is convenient because it makes taste sound untrainable. I think it is more useful to treat taste as a working capacity: > Taste is the trained ability to reject what almost works. Not what obviously fails. That part is easy. Anyone can delete the broken paragraph, the unreadable mockup, the nonsense strategy, the image with six fingers, the feature nobody asked for. A bad option announces itself. A plausible option negotiates. Taste starts where the thing is acceptable. The sentence is clear, but too expected. The design is clean, but forgettable. The feature is useful, but off-strategy. The argument is true, but not alive. The idea is clever, but it does not change the reader's next move. This is why AI makes taste more important. It is very good at producing almost-working things. Almost-working things are hard to kill. ## The phrase did the thinking George Orwell's *Politics and the English Language* is still useful here because it treats bad prose as a failure of attention, not just style. His target was political language, but the craft mechanism is broader: stale phrases, inflated diction, passive padding, and ready-made expressions let the writer stop looking directly at what they mean.[^1] That is the part that matters now. The danger of AI prose is not only that it sounds like AI. The deeper danger is that it gives you language before you have earned the thought. The phrase arrives finished. "Unlocking potential." "Navigating complexity." "A robust framework." "At the intersection of X and Y." "Not just a tool, but a partner." None of these phrases is evil in isolation. The problem is that they are available before the writer has made a choice. They are prefabricated thinking. When you accept them, you inherit the shape of a thought without doing the work of locating the thought itself. You can feel the page becoming smoother while the claim becomes less yours. This is why deletion is not cosmetic. > Deleting a phrase can reveal that there was no thought underneath it. That is useful. It tells you where the work is. **The empty place is not a failure.** It is the honest outline of the next sentence. ## Good taste is not perfect agreement The obvious objection is that taste is subjective. People disagree. Experts get it wrong. Styles change. What feels sharp to one reader feels cold to another. What feels beautiful in one culture can feel empty in another. All true. But "not perfectly objective" is not the same as "random." Paul Graham makes a useful argument here. If there were no such thing as good taste, then there would be no such thing as good art, and therefore no way for artists to be good at their jobs. That conclusion is absurd. People can be better or worse at painting, writing, acting, composing, designing, and judging. Taste is messy, but not imaginary.[^2] The practical point is simpler than the philosophy. Taste is not perfect consensus. Taste is better discrimination under pressure. It is the ability to notice the difference between: Clear and obvious. Simple and thin. Polished and true. Interesting and useful. Novel and merely weird. Confident and overfitted. Human and human-sounding. That discrimination improves through exposure, practice, comparison, and refusal. You see more work. You make more work. You compare the almost-good against the actually-good. You learn which signals are real and which are social noise. AI can increase exposure. It cannot do the refusal for you. ## The slop problem is a selection problem AI-assisted writing often repeats the same moves because models learn from patterns that occurred often enough to become likely. The model learns the common move. The confident opener. The balanced contrast. The soft caveat. The tidy closer. The problem is not that every common move is wrong. Common moves become common because they often work somewhere. The problem is that common moves are cheap. They arrive too easily. Slop is not a list of banned words. It is what happens when familiar language arrives faster than judgment. A human with taste does not merely ask, "Is this sentence grammatical?" or "Does this paragraph sound professional?" That bar is too low now. The better question is: > Would I have chosen this if it had not been handed to me? Most AI-assisted work fails there. Not because the model is useless. Because the human has not taken responsibility for the final selection. You can see it in writing, but the same pattern shows up everywhere. The product team ships the feature that sounds good in a roadmap but does not sharpen the product. The founder keeps the positioning line that flatters the company but does not make the buyer move. The designer keeps the visual flourish that signals effort but weakens the hierarchy. The analyst keeps the chart that is technically accurate but answers the wrong question. The creator keeps the section that proves they did research but slows the piece down. These are not generation failures. They are deletion failures. ## The Deletion Pass The useful question is not "how do I make more?" It is "what has earned the right to remain?" Run this on any AI-assisted draft, design, memo, article, product idea, or strategy. First, make the option set. Let the model help. Generate widely. Explore. Ask for alternatives. Get the obvious versions out of your system. Then stop generating. This is the part people skip. Now run the Deletion Pass. Ask five questions. > What is this piece actually trying to do? > > Which part is only here because it sounds good? > > Which part could be removed without changing the outcome? > > Which part feels polished but hides a weak decision? > > If I had to keep only one line, feature, image, or claim, what would survive? That last question is the taste test. Not because everything else is worthless. Because the strongest element reveals the job of the whole thing. If you cannot name the survivor, the work has not converged yet. Do not just delete silently. Keep a short cut list: > I cut this because __________. That one sentence turns deletion into learning. Over time, the cut list becomes a map of your taste: the phrases you no longer trust, the features you keep overbuilding, the kinds of cleverness that make your work worse. The test is simple enough to draw: [Figure: AI widens the field. Taste narrows it without becoming generic.] *AI widens the field. Taste narrows it without becoming generic.* ## The hard reps are the point There is a subtle trap in using AI as a creative partner. It can remove the reps that build taste. The bad draft matters because you learn why it is bad. The awkward sentence matters because you feel where the rhythm breaks. The failed design matters because you learn what your eye was pretending not to see. The rejected feature matters because it teaches you the product's actual spine. The wrong strategy matters because it exposes the hidden assumption. If AI skips you over all of that, you may get a better first output and a weaker internal judge. That is not a reason to avoid AI. It is a reason to use it in a way that keeps the hard reps alive. Do not only ask the model to generate. Ask it what should be cut. Ask it which option is most generic. Ask it which paragraph is pretending to be useful. Ask it which section exists to prove effort rather than serve the reader. Then disagree with it. Make the final deletion yourself. That is where the skill compounds. ## The person who can delete The next advantage in creative work will not belong to the person who can make the most drafts. Everyone will have drafts. It will not belong to the person with the longest prompt. Prompts will spread. It will not belong to the person who can produce the most polished surface. Polish is getting cheaper. It will belong to the person who can stand in front of abundance and remove what does not belong. The person who can kill the clever line. The person who can cut the feature that investors liked. The person who can delete the slide that makes the deck feel smarter but the decision less clear. The person who can look at twenty AI-generated options and say: No. No. No. This one. And then make that one better. In an age of infinite drafts, taste is not what you add. Taste is what you delete. --- *What is one thing in your current work that probably sounds good but has not earned the right to stay?* [Figure: Field Card — Taste Is What You Delete] [^1]: George Orwell, *Politics and the English Language*, first published in *Horizon* (April 1946), hosted by The Orwell Foundation, [https://www.orwellfoundation.com/the-orwell-foundation/orwell/essays-and-other-works/politics-and-the-english-language/](https://www.orwellfoundation.com/the-orwell-foundation/orwell/essays-and-other-works/politics-and-the-english-language/). Orwell's rules and examples are used here as a craft analogy for deletion and conscious word choice. [^2]: Paul Graham, *Is There Such a Thing as Good Taste?* (November 2021), [http://www.paulgraham.com/goodtaste.html](http://www.paulgraham.com/goodtaste.html). Graham argues that taste is neither perfect consensus nor pure randomness; the article uses that limited point, not a claim that aesthetic judgment is fully objective. --- --- title: "The Displacement Rate Audit" description: "A five-minute scoring tool for any product, position, architecture, business model, or career bet. Find out what still works after the environment changes." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-displacement-rate-audit/" date: "2026-05-01" series: "THE HUMAN LAYER" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Displacement Rate Audit *A five-minute scoring tool for any product, position, architecture, business model, or career bet. Find out what still works after the environment changes.* By Harry Floyd · 2026-05-01 · canonical: https://durabilitycurve.com/blog/the-displacement-rate-audit/ Most plans are scored against the world they were made in. That is the problem. The useful question is not whether the product, thesis, architecture, position, or career bet works today. The useful question is how much of it still has a job when the surroundings change. The Displacement Rate Audit is a free five-minute tool for asking that question before reality asks it for you. Run it on one thing you are currently defending: - a product line; - an investment thesis; - a technical architecture; - a business model; - a career bet; - a major project your team is still defending. Then name the most likely large change in the next eighteen months. The audit only works if the change is specific enough to argue with. Not “AI gets better.” Something concrete: > An open-source model matches your benchmark and runs on commodity hardware. > > Your sector takes thirty percent multiple compression. > > Your main distribution channel stops favouring your format. > > The buyer no longer needs the workflow your product was built around. Now score what survives. The audit gives you a 0-5 displacement score. Zero means nothing survives; the work belonged to the old version of reality. Five means the change was already accounted for in the original design. The score matters less than the forcing function: > If you cannot name the specific components, contracts, positions, relationships, or capabilities that still have a job after the change, your real score is lower than the one you wrote. This is why the tool is useful across domains. It does not ask whether something is impressive. It asks whether it is durable under a named disturbance. The output is deliberately small: one named change, one score, and one list of the parts that survive. That is enough to make the next decision harder to fake. Use it before adding features. Use it before sizing a position. Use it before doubling down on a technical architecture. Use it when a strategy still sounds good but you can feel the surroundings moving. **Download.** [Figure] The Displacement Rate Audit 121KB ∙ PDF file [Download](https://harryfloyd.substack.com/api/v1/file/ac157ba4-fb21-466d-b9df-d28d6f44b267.pdf) [Download](https://harryfloyd.substack.com/api/v1/file/ac157ba4-fb21-466d-b9df-d28d6f44b267.pdf) **More free tools like this.** _Subscribe to get the next durability-lens resource the day it ships._ The Displacement Rate Audit is a companion to _[The Forest Floor Is the Product](https://durabilitycurve.com/blog/the-forest-floor-is-the-product/)_, the article that develops the substrate-vs-canopy lens across ecosystems, software, knowledge work, and capital allocation. _What is one thing you are building that would score lower than you want if the next big change arrived tomorrow?_ --- --- title: "The Substrate Map" description: "A free one-page taxonomy and 10-minute exercise for finding the substrate-vs-canopy ratio in your last 90 days of work." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-substrate-map/" date: "2026-05-01" series: "SYSTEMS & LAWS" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Substrate Map *A free one-page taxonomy and 10-minute exercise for finding the substrate-vs-canopy ratio in your last 90 days of work.* By Harry Floyd · 2026-05-01 · canonical: https://durabilitycurve.com/blog/the-substrate-map/ Most teams can tell you what they shipped. Fewer can tell you what will still matter after the next large change. That is the gap the Substrate Map is built for. It is a free one-page taxonomy for separating canopy from substrate in your own work. The canopy is the visible layer: prompts, model choices, demo polish, current benchmark scores, frameworks, UI surfaces, launch artefacts. The substrate is the part that keeps doing work when the surface gets repriced: data-quality discipline, eval contracts, workflow integration, trust packaging, domain-specific failure memory, and the proprietary signal the next model release does not have. The tool gives you a 10-minute exercise: 1. Open the last 90 days of engineering tickets, product launches, roadmap decisions, or investment decisions. 2. Tag each item as substrate or canopy. 3. Compute the ratio. The number is blunt on purpose. If the last 90 days were mostly canopy, the next release can reset most of what you built. If the split is 50/50, you are probably normal but not especially durable. If the work is mostly substrate, protect it. That is the work compounding underneath the visible output. The most useful part is the boundary rule: > If your team cannot agree which column an item belongs in, tag it as canopy. Disagreement at the boundary means the substrate work has not been made explicit yet. That makes the map useful before a planning meeting. Instead of arguing about whether a roadmap “feels strategic”, you can ask which work would still matter if the model, market, channel, or buyer changed. The conversation gets harder to fake because each item has to be placed in a column. The PDF is deliberately simple: one map, one exercise, one ratio. It is not a strategy deck. It is the first instrument you run when the team is shipping a lot but cannot say what is compounding. Use it on a roadmap. Use it on a product backlog. Use it on a portfolio. Use it before a planning cycle where everyone is about to argue from vibes. **Download.** [Figure] The Substrate Map 105KB ∙ PDF file [Download](https://harryfloyd.substack.com/api/v1/file/5b324df8-e43f-486f-823b-7213b88910b4.pdf) [Download](https://harryfloyd.substack.com/api/v1/file/5b324df8-e43f-486f-823b-7213b88910b4.pdf) * * * **More free tools like this.** _Subscribe to get the next durability-lens resource the day it ships._ The Substrate Map is a companion to _[The Forest Floor Is the Product](https://durabilitycurve.com/blog/the-forest-floor-is-the-product/)_, the essay that develops the substrate-vs-canopy lens across ecosystems, software, knowledge work, and capital allocation. _What percentage of your last 90 days was substrate?_ --- --- title: "PLTR: The AI Stock That Has To Prove It Owns The Permission Layer" description: "A bull/bear thesis for Palantir: not whether AI demand is real, but whether Palantir owns the permission layer between model capability and real-world action." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/pltr-the-ai-stock-that-has-to-prove/" date: "2026-04-30" series: "MARKETS & POWER" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # PLTR: The AI Stock That Has To Prove It Owns The Permission Layer *A bull/bear thesis for Palantir: not whether AI demand is real, but whether Palantir owns the permission layer between model capability and real-world action.* By Harry Floyd · 2026-04-30 · canonical: https://durabilitycurve.com/blog/pltr-the-ai-stock-that-has-to-prove/ Research frame, not investment advice. Palantir is easy to write badly. The bull version becomes a cheerleading note: AI is huge, Palantir is growing fast, governments trust it, enterprises are buying AIP, and the margins look like elite software. The bear version becomes a valuation complaint: the stock is wildly expensive, the story is crowded, insiders sell, and the price already assumes something close to perfection. Both versions miss the harder question. The real PLTR debate is not whether Palantir is riding the AI cycle. That frame is too blunt to be useful. The better question is: > When AI moves from demos into real organisations, who owns the layer that decides what the system is allowed to see, decide, do, and prove afterwards? That is the permission layer. And that is what Palantir has to prove it owns. The distinction changes how you look at the company. A model can answer a question. A workflow can move work forward. But a permission layer decides whether an AI system is allowed to touch what matters: operational data, regulated decisions, live actions, and evidence that can survive audit. This is why Palantir is interesting. Not because it is another beneficiary of AI spending, but because it may sit at the point where AI spending becomes operationally real. A story can be true and still be over-owned. An instrument becomes harder to replace the more the world depends on the thing it measures, controls, or makes executable. PLTR is the argument between those two sentences. [Figure] _The useful question is not whether AI is useful. It is who controls the path from model capability to real-world action._ ## The Thesis In One Sentence Palantir is a bet that the bottleneck in AI deployment moves from model capability to permissioned execution: the ability to make AI see the right context, target the right action, obey the right constraints, and leave behind proof. That sentence is doing a lot of work, so let me unpack it. The public AI conversation still treats intelligence as the scarce layer. Which model is smartest? Which benchmark moved? Which chatbot can reason better? Regulated enterprises and governments do not live in that world for long. They do not only need a model that can produce an answer. They need a system that knows which data the model can see, which action it can take, which human must approve it, which trace gets preserved, and which decision can be defended later. That gives us a more useful way to judge Palantir than the usual AI-stock framing. Ask four questions: 1. **See:** can the system connect to the organisation’s real data, not a cleaned-up demo version? 2. **Decide:** can it target the right operational choice, not just produce a plausible answer? 3. **Act:** can it execute inside permissioned workflows without blowing through constraints? 4. **Prove:** can it preserve the evidence trail when someone asks what happened? That is the Palantir test: > **Can Palantir own the loop between data, decisions, actions, and proof?** If Palantir owns that loop, the company is not just selling AI software. It is selling operational control. If Palantir does not own that loop, then it is a very impressive software company trading like it owns more of the future than it actually does. The underlying idea is simple: powerful AI inside a messy organisation is not automatically useful. It becomes useful only when the organisation can see the right context, aim the system at the right target, constrain what it is allowed to do, and prove what happened afterwards. That is the part Palantir is trying to own. > **Reader map:** if you remember one thing, remember the split between _useful AI_ and _allowed-to-act AI_. Palantir’s valuation only makes sense if that second layer stays scarce. ## What The Market Is Already Paying For The market is not asleep to this possibility. Palantir's own Q4 2025 release gives the bull case real numbers to work with: [1](#footnote-1) - FY 2025 revenue: $4.475 billion, up 56% year over year. - Q4 2025 revenue: $1.407 billion, up 70% year over year. - U.S. commercial revenue in Q4: $507 million, up 137% year over year. - U.S. government revenue in Q4: $570 million, up 66% year over year. - FY 2026 revenue guide: $7.182-$7.198 billion, implying about 61% growth. - FY 2026 U.S. commercial revenue guide: more than $3.144 billion, implying at least 115% growth. - Adjusted free cash flow margin in Q4: 56%. These are not hype numbers. They force a serious bear to work harder. But the price is doing work too. At roughly $332 billion of market value and about 74 times trailing sales as of the Apr 30 Trefis snapshot, the stock is not merely pricing in a good software business. It is pricing in a business that stays scarce while the rest of the AI stack commoditises around it. [2](#footnote-2) The market is already underwriting a specific future: - AI use keeps moving from pilots into live operations. - Regulated customers need more than generic model access. - Palantir remains one of the few credible vendors for that deployment layer. - Commercial adoption keeps accelerating without destroying margins. - Hyperscalers do not bundle away enough of the workflow, audit, and permissioning layer to compress the premium. That is a lot to ask. The stock is not asking whether Palantir can be good. It is asking whether Palantir can remain exceptional for long enough that today’s valuation becomes a rational underwriting rather than a momentum receipt. [Figure] This is why PLTR is a useful test case. The company is producing the kind of growth that deserves attention. The price is asking whether that growth belongs to a scarce layer, not just to an AI budget cycle. ## The Bull Case: Palantir Owns A Scarce Layer The strongest bull case is not “AI is big.” That is too broad. It explains almost nothing. The stronger bull case is that Palantir sits where generic AI stops being useful and operational AI starts being valuable. There is a canyon between “this model can answer a question” and “this system can run a live decision inside a hospital, factory, bank, battlefield, or government agency.” A large company does not become AI-native because someone plugs an LLM into Slack. A government agency does not modernise because a chatbot can summarise a PDF. The hard part is connecting messy data, permissions, workflows, decisions, and audit trails into a system that can survive contact with operations. This is where Palantir’s ontology language matters. The word can sound abstract, but the practical claim is simple: Palantir tries to model how an organisation actually works. Which objects matter? Which relationships matter? Which users can act on which information? Which decisions need to be made at which point in the process? Which actions require a human? Which outputs need a trace? If that model becomes embedded inside the customer, the moat is not only software. It is context. This is the deeper point: architecture often outlives content. The models will change. The workflows will change. The specific AI interface will change. But if the organisation’s operational map lives in Palantir, the scaffold may persist while the content turns over. This is also why the government side matters. Defence, intelligence, and regulated operations are not ideal environments for generic AI wrappers. They need permissioning, provenance, escalation, and traceability. They need systems that know not just what can be generated, but what can be done. The bull case is that AIP has turned this old Palantir strength into a faster commercial motion. The numbers support that possibility. U.S. commercial revenue grew 137% year over year in Q4 2025. U.S. commercial remaining deal value reached $4.38 billion, up 145% year over year. Total contract value in Q4 was $4.262 billion, up 138% year over year. [3](#footnote-3) If those numbers represent durable platform adoption rather than a temporary AI budget surge, PLTR becomes one of the cleanest public-market examples of a bigger shift: companies paying for the systems that let AI act safely, not just answer fluently. The bull case, stated cleanly, is this: > As model intelligence becomes more available, the scarce layer becomes the system of permissioned execution around it. Palantir may already be installed where that scarcity appears first. That is a strong case. It is also exactly why the bear case has to be better than “the stock is expensive.” ## The Bear Case: The Thesis May Be Right And The Stock Still Too Expensive The weak bear case is that Palantir is overhyped. That is not good enough. A better bear case starts by granting the company its strengths. Palantir may be a rare business. The product may be real. AIP may be accelerating adoption. The government moat may be durable. The margins may be excellent. A great company can still be a bad underwriting if the market has already bought the whole story. A company trading at more than 70 times sales does not merely need to grow. It needs to keep the market believing that its growth is unusually durable, unusually profitable, and unusually hard to compete away. Several things can break that belief. First, hyperscalers can bundle enough of the instrument layer into the cloud stack. If Azure, Google Cloud, AWS, OpenAI, or Anthropic make AI governance, evals, audit trails, and permissioning good enough inside their own platforms, some customers may accept the bundled version rather than paying a Palantir premium. Second, AIP traction is still partly company-reported. The growth is real, but commercial AIP durability is not yet independently triangulated enough to treat every bootcamp, customer count, or deal metric as proof of deep production embedding. Third, government strength cuts both ways. It validates the product in demanding environments, but it also introduces procurement cycles, political risk, budget exposure, and reputational constraints. Fourth, the valuation makes every slowdown louder. If revenue growth normalises before margins and customer expansion prove the full platform thesis, the multiple can compress even while the business keeps improving. That is the uncomfortable part of PLTR: the bear case does not require the company to disappoint in an ordinary sense. It only requires the company to become less exceptional than the price implies. The bear case is not that the story is fake. The bear case is that the story is so attractive that the market may have stopped asking what would falsify it. That makes PLTR a useful mental model for AI investing more broadly: who owns a scarce control point after model capability gets cheaper? If the answer is “Palantir owns the permission layer,” the premium may have logic. If the answer is “Palantir is one strong vendor inside a layer that clouds, model labs, and internal platforms can partially absorb,” the premium becomes much harder to defend. ## What Would Change My Mind The point of a bull/bear article is not to sound balanced. It is to make the thesis falsifiable. For PLTR, I would watch four things. [Figure] ### 1\. Bundled AI governance gets good enough If hyperscalers bundle eval, audit, observability, permissions, and model governance into the model or cloud API at low incremental cost, Palantir’s instrument-layer scarcity weakens. This does not require hyperscalers to replicate Palantir completely. They only need to be good enough for enough customers. ### 2\. AIP growth slows before durability is proven If U.S. commercial growth slows sharply before there is independent evidence that customers are building durable operational workflows on AIP, the market may reclassify AIP from platform shift to adoption spike. ### 3\. Verification depth does not become a priced contract dimension The right listen-for is whether Palantir can price verification depth. Sampling, replay coverage, trace retention, judge configuration, audit bundles, permission ladders, and regulated workflow evidence are the kinds of features that would make the thesis concrete. If customers pay for those layers, the thesis strengthens. If they treat them as bundled table stakes, the thesis weakens. This is subtle but important. A feature can be necessary without being separately valuable. Palantir needs the market to pay for the depth of the instrument, not merely expect it as part of the package. ### 4\. Valuation stays extreme while growth normalises This is the simplest one. Great company, bad underwriting. If growth normalises and the stock still trades as though the exceptional phase is permanent, the risk shifts from business quality to entry price. ## Verdict PLTR is one of the most interesting public-equity expressions of the AI-deployment thesis. It is not just an “AI stock.” It is a bet on the layer that makes AI usable inside organisations where mistakes matter. More specifically, it is a bet that the money in enterprise AI moves toward permissioned execution: see the right data, decide against the right target, act inside the right constraints, and prove what happened afterwards. That is a much better lens than “AI beneficiary.” It also makes the valuation harder, not easier. The stock is already priced like the market understands a lot of this. So my current posture would be: **Watchlist / research, not automatic buy.** The company may be exceptional. The article’s job is not to deny that. The job is to separate three things that often get collapsed: 1. Is the business real? 2. Is the moat durable? 3. Is the current price a good underwriting of that durability? For Palantir, the answer to the first is increasingly yes. The second is the live thesis: does Palantir really own the permission layer, or does it only participate in it? The third is where the fight is. If you want more pieces like this, subscribe for essays on AI, markets, and the hidden infrastructure layer behind what looks like hype. [1](#footnote-anchor-1) Palantir Q4 2025 earnings release / SEC Exhibit 99.1, including FY 2025 revenue, Q4 2025 segment growth, FY 2026 guidance, free-cash-flow margin, total contract value, and U.S. commercial remaining deal value: [https://www.sec.gov/Archives/edgar/data/1321655/000132165526000004/a2025q4ex991earningsrelease.htm](https://www.sec.gov/Archives/edgar/data/1321655/000132165526000004/a2025q4ex991earningsrelease.htm) [2](#footnote-anchor-2) Trefis PLTR valuation snapshot used for the Apr 30 market-cap and valuation-multiple references. [https://www.trefis.com/data/companies/PLTR?from=PLTR-2026-03-01](https://www.trefis.com/data/companies/PLTR?from=PLTR-2026-03-01) [3](#footnote-anchor-3) Palantir Q4 2025 earnings release / SEC Exhibit 99.1, including FY 2025 revenue, Q4 2025 segment growth, FY 2026 guidance, free-cash-flow margin, total contract value, and U.S. commercial remaining deal value: [https://www.sec.gov/Archives/edgar/data/1321655/000132165526000004/a2025q4ex991earningsrelease.htm](https://www.sec.gov/Archives/edgar/data/1321655/000132165526000004/a2025q4ex991earningsrelease.htm) --- --- title: "Right Company, Wrong Vector" description: "A pick is a number. A position is a vector. The post-mortem language we have only knows how to blame the company." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/right-company-wrong-vector/" date: "2026-04-26" series: "MARKETS & POWER" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Right Company, Wrong Vector *A pick is a number. A position is a vector. The post-mortem language we have only knows how to blame the company.* By Harry Floyd · 2026-04-26 · canonical: https://durabilitycurve.com/blog/right-company-wrong-vector/ In November 2011, Berkshire Hathaway bought sixty-four million shares of IBM at an average price of around one hundred and seventy dollars. The position was worth roughly $10.7 billion at cost. It was a 5.5 percent stake in the company and one of the largest single positions Berkshire had ever opened in a publicly traded name. [1](#footnote-1) The thesis Warren Buffett put on the inside cover of that position was specific. IBM was no longer a hardware company. It was a services-led moat with deeply embedded enterprise customers who would not switch lightly. The thing he was buying was the durability of that moat against everyone trying to displace it. [2](#footnote-2) Six years later he sold most of it, in stages, at prices below where he had bought. By the spring of 2018 the position was gone. Over the same window IBM stock was down roughly eighteen percent. The S&P 500 was up one hundred and sixteen percent. Berkshire’s other technology bet, the AAPL position Buffett had begun in 2016, had already grown larger than IBM had ever been on Berkshire’s book. [Figure: TradingView chart] _IBM - 2011 - 2018_ In May 2017 he told CNBC, _“I don’t value IBM the same way that I did six years ago when I started buying. I’ve revalued it somewhat downward.”_ [3](#footnote-3) Nine months after the exit was complete, he was more direct. _“I was wrong, or at least I felt like I was wrong on IBM when I sold it and I was wrong when I bought it.”_ [4](#footnote-4) That second sentence is the one to read carefully. It does the only thing the vocabulary lets it do. It blames the thesis. The whole story collapses into a single axis. Was he right about the company, or was he wrong about the company. He worked out he was wrong. He said so out loud, in public, with his own name on it, which is more than almost anyone in the industry will ever do. And the most articulate post-mortem voice in modern finance still got compressed into one word. _Wrong._ This essay is about what is missing from that word. ## A Pick Is a Magnitude. A Position Is a Vector. In physics class, a magnitude is a number. A vector is a number with a direction attached. Forty miles per hour is a magnitude. Forty miles per hour going north is a vector. The two carry different information. A magnitude tells you how much. A vector tells you how much, and where it is pointed. The investing industry has one word for both, and it is the wrong word. If you have ever held a name through a triple and not felt the triple in your book, you have lived this. The magnitude was right. The vector was wrong. A position is a vector with at least seven slots. Size is one of them. The thesis sentence is another. The other five are the ones the post-mortem cannot name out loud: what would have to be true for the thesis to be wrong, how long the thesis is allowed to take, what other names in the book this position is correlated to, what the position pays out if the thesis only half-lands, and how the position would be exited if the falsifier triggered. Each of those is a separate decision. Each can move while the ticker stays the same. None of them are in the magnitude. The whole apparatus of the industry, the position percentage on a tearsheet, the weight in a 13F filing, the line in a quarterly letter, is built to compress all seven slots into the one slot the file format can hold. > A pick is a magnitude. A position is a vector. The industry’s word for both is the same word, and the word that wins is the smaller one. The category error sits here. The investor hears _“what is your largest position”_ and answers with a name and a percentage. The right answer, the one that survives a bad year, is a profile. The percentage is one number in that profile. It is not the profile. ## Two Investors. One Company. Two Different Positions. Run the thought experiment with two investors over Buffett’s window. Both wrote the same thesis on the inside cover in November 2011. _IBM is no longer a hardware company. It is a services-led moat with deeply embedded enterprise customers who will not switch easily._ The same paragraph. The same name. The same year. The first investor sized the position to roughly five percent of the equity book on conviction in the moat. The exit rule was _“if I change my mind.”_ The holding period was _“long term.”_ The falsifier was nowhere on paper. The dependency on the cloud transition being slow rather than fast was implicit, not stated. The opportunity cost against the next-best technology bet of the decade was not tested until that bet had already done the heavy work for someone else. The second investor sized the same thesis at one percent. They wrote a falsifier in the position memo: _if IBM’s services revenue declines for two consecutive quarters with management citing competitive losses to cloud-native vendors, the moat thesis is invalidated for this regime, and the position closes within the next reporting cycle._ The holding period was a rolling four-quarter window, not “long term.” Re-evaluation was on the calendar at every print. The dependency on the cloud transition being slow was named explicitly, so any acceleration in cloud adoption would tighten the falsifier rather than leave it dormant. By 2014, when IBM’s services revenue first showed real cracks with explicit cloud-competitive language from management on the call, the second investor’s falsifier triggered. They closed the position into the next reporting cycle at a small loss against entry and redeployed. The first investor read the same earnings, watched the chart hold, and stayed. Same thesis. Same company. Same paragraph on the inside cover. Two completely different positions, three years apart in their first invalidation event, with two completely different outcomes downstream. > Two investors with the same thesis on the same company can hold completely different positions, and only one of them can be reverse-engineered from the post-mortem. When the first investor wrote up the failure, the only sentence available was _“I was wrong about IBM.”_ It is not actually a true sentence about the thesis. The thesis was approximately right at the level of granularity at which it was written. IBM did remain a services-led business. Most of its enterprise customers did not switch lightly. The moat existed. It just degraded faster than the size and the holding period and the missing falsifier had quietly assumed it would. The vocabulary made it look like a thesis error. It was a vector error, in the durability slot, the falsifier slot, the holding-period slot, and the opportunity-cost slot. Four wrong directions on a vector that was being held as if it were a magnitude. ## The Substrate Speaks Before the Headline The second investor noticed something the first one did not. A thesis is a claim about a substrate. _IBM has a services moat_ is not the substrate. It is a sentence about the substrate. The substrate itself is the network of switching costs, the salesforce relationships, the integration debt customers had built on IBM’s stack, the technical depth of the services organisation, the rate at which competitors could credibly displace any of those things. The thesis sentence holds up only as long as the substrate underneath it does. Substrates erode slowly. They erode in the kind of small, public, observable details that do not move the price chart for several quarters. AWS launched in 2006. By 2013 it was at scale. By 2014 enterprise cloud adoption was visibly accelerating in exactly the kind of Fortune 500 customer accounts IBM’s services moat was supposed to protect. By 2015 IBM’s own earnings calls were naming cloud competitive pressure in the services segment. The price chart did not reflect any of that until 2016, and even then only partially. The headline followed the substrate by about three quarters. If you have ever read a 10-Q where a segment that used to anchor the thesis is suddenly being described in defensive language, and decided to wait one more print to confirm what you already knew, you have felt this. The substrate told you. The headline took another nine months to follow. > The substrate erodes three quarters before the headline does. The score is the only instrument that can read the gap. A long-only allocator’s highest-leverage early warning is not the price chart. It is a substrate score, written down at a fixed cadence, against a stable framework. Score the durability of the moat. Score what the position pays out if the thesis only half-lands. Score the dependence on couplings the world is moving against. Score the optionality the position carries if the thesis is wrong. Re-score every six months. Any axis that has dropped twice in a row is no longer the position that was opened. Any axis that drops by a full point earns one paragraph in the log: what changed, and what would have to be true for the score to recover. Earnings dates make this almost free. They are not news. They are pre-scheduled observability windows. Same date every quarter, same disclosures, same metrics, an instrument with a known sampling time. The investor who runs the same protocol at every window has a research process: read the thesis, read the falsifier, write three listen-fors before the call, log the result against each listen-for after the call. The investor who reacts to whichever earnings happened to be loudest has a feed. A position held on the strength of the thesis sentence alone, without any substrate score behind it, is sized on conviction without an invalidation rule. That is what every post-mortem ending in the single word _wrong_ has in common. ## Every Commitment Has a Heading The lens is not about investing. A hire is a position. The thesis is the candidate’s competence. The vector is the role they are hired into, the team they are embedded with, the manager they report to, the falsifier (what would constitute a bad-fit signal in ninety days), the time-box (how long the trial period runs in practice), the couplings (whose other work depends on this hire), and the exit (what graceful off-boarding looks like). The same person can be a brilliant hire in one vector and an expensive one in another. The thesis on the candidate is the same. The vector decides whether the year ends in a promotion or a severance package. Most failed hires get post-mortemed as bad hiring decisions. Some are. Many of them are wrong-vector decisions on right hires, and the language only knows how to blame the candidate. A research bet is a position. The thesis is the hypothesis. The vector is the experimental design, the cohort size, the time-budget, the falsifier, the dependencies on parallel experiments, the exit rule. Two labs with the same hypothesis run different experiments, and one paper lands and the other does not replicate. The hypothesis was the same. The vector was different. A founder’s first market is a position. The thesis is the product. The vector is the timing, the geography, the segment, the pricing, the go-to-market motion. Same product, two markets, two outcomes. The founder who failed in 2018 with a webhooks tool and succeeded in 2024 with the same webhooks tool was right both times on the product. The vector was different. > Every committed action has a heading. Most of us only ever name the destination. The reason the lens travels is that the failure mode travels with it. Anywhere a committed action is named by its scalar headline (a hire by the candidate’s name, a research bet by the hypothesis, a market entry by the product), the language collapses the vector into the magnitude and the post-mortem inherits the collapse. _“I was wrong about him.”_ _“I was wrong about the hypothesis.”_ _“I was wrong about the market.”_ The post-mortems sound the same because the vocabulary makes them sound the same. They are usually about different axes, on different vectors, of different kinds of commitment, and the language has no way to say so. ## Six Lines, Five Numbers, Ten Minutes Pick the largest position in the book. It does not have to be a position in the markets. It can be the most expensive person on the team, the most time-consuming research line, the largest standing commitment of attention to a single bet of any kind. Write the company, or the person, or the project, on one line. Write the thesis sentence on the next. Write the substrate sentence underneath it. The one sentence that names what has to be true about the substrate beneath the thesis for the thesis to play out. Not the thesis restated. The thing the thesis silently depends on. Write the falsifier on the next line, with a number or a date attached. The one sentence that names what would have to be true for the thesis to be wrong, in a form specific enough to trigger when the time comes. Score the five axes. Durability, asymmetry, replicability, couplings, optionality. One to ten on each. Write down what would make each axis drop a point. Six lines. Five numbers. About ten minutes for a commitment the book is genuinely committed to. The substrate sentence comes hardest. The falsifier comes second. The five axes come quickly because the test is structured. The substrate sentence and the falsifier ask for prose the position memo never demanded, and that is the diagnostic. If the substrate sentence is hard to write, the position is being held on thesis alone, and the substrate is doing all the work and getting none of the credit. If the falsifier is hard to write, the position is sized off conviction without invalidation, which is the structural shape of every position that is one earnings call away from a permanent loss the post-mortem will struggle to explain. The strongest counter to running this exercise at all is that elaborate position-scoring is activity bias dressed up as discipline, and that for most allocators most of the time the right move is to own a broad-market index fund and stop touching it. [5](#footnote-5) The counter is correct, for those allocators. It does not bind on the population this essay is written for: anyone running an active book where the difference between right-thesis-right-vector and right-thesis-wrong-vector decides ten years of P&L. The two investors at the top of this piece were always going to be the same person, ten years apart. Buffett-2011 wrote the position memo. Buffett-2018 wrote the post-mortem. The memo named one thing, the thesis. The post-mortem had only one word for the failure, _wrong,_ and the word collapsed seven slots back into one. The thesis sentence Buffett-2011 wrote was approximately true. The vector underneath it was the part that drifted. The language he had no way to name showed up as eighteen percent down on the company over a hundred and sixteen percent up on the index, and as AAPL becoming a hundred and sixty-five million shares in someone else’s portfolio first. The same paragraph on the inside cover, ten years apart, can produce two completely different positions and two completely different outcomes. The vocabulary is the lever. Magnitude gets named. Vector runs the P&L. The investor who learns to write the second one out, one substrate sentence and one falsifier with a number or a date attached, is no longer holding the right company the wrong way. * * * _Pick the largest position in your book and run the six-line test this week. Which of the five axes is the slot you have never had to write down, and what would make it drop a point? The slot most allocators leave blank is the one I will write next. Comments are open below._ * * * * * * The five-axis scoring tool sits at [The Investor’s Substrate Test](https://durabilitycurve.com/blog/the-investors-substrate-test/), free, single-PDF, seven minutes per position. * * * ## Footnotes [1](#footnote-anchor-1) Berkshire Hathaway disclosed a 5.5 percent stake in IBM in November 2011, comprising approximately sixty-four million shares accumulated through 2011 at a cost basis of approximately $10.7 billion. The implied average cost is roughly one hundred and seventy dollars per share. See Reuters, “Berkshire buys 5 pct of IBM, takes other stakes” (November 14 2011), and the Investopedia retrospective “Berkshire Hathaway Has Exited IBM: Buffett” (May 2018), [https://www.investopedia.com/news/berkshire-hathaway-has-exited-ibm-buffett/](https://www.investopedia.com/news/berkshire-hathaway-has-exited-ibm-buffett/) [2](#footnote-anchor-2) Buffett’s original IBM thesis emphasised the company’s transition from a hardware manufacturer to a services-led business with high enterprise switching costs, and the credibility of management’s articulated capital-allocation roadmap through 2015. The framing held that customers, once integrated into IBM’s services stack, would face high transition costs to displace it. The structural claim, that customer stickiness was the load-bearing assumption, is consistent with Buffett’s 2017 admission that he had revalued IBM downward citing “big strong competitors” eroding precisely that component, rather than any disagreement with the cash-flow profile. See CNBC, “Warren Buffett has ‘revalued’ IBM downward, cites ‘big strong competitors’” (May 4 2017), [https://www.cnbc.com/2017/05/04/warren-buffett-has-revalued-ibm-downward-cites-big-strong-competitors.html](https://www.cnbc.com/2017/05/04/warren-buffett-has-revalued-ibm-downward-cites-big-strong-competitors.html) [3](#footnote-anchor-3) Warren Buffett, on CNBC ahead of the Berkshire Hathaway 2017 annual meeting, May 4 2017. Full quoted excerpt: _“I don’t value IBM the same way that I did six years ago when I started buying. I’ve revalued it somewhat downward... IBM is a big strong company, but they’ve got big strong competitors too.”_ See CNBC Excerpts, May 5 2017, [https://www.cnbc.com/2017/05/05/cnbc-excerpts-billionaire-investor-warren-buffett-speaks-with-cnbcs-becky-quick-ahead-of-the-berkshire-hathaway-annual-meeting.html](https://www.cnbc.com/2017/05/05/cnbc-excerpts-billionaire-investor-warren-buffett-speaks-with-cnbcs-becky-quick-ahead-of-the-berkshire-hathaway-annual-meeting.html) [4](#footnote-anchor-4) Warren Buffett, CNBC, February 2018, after Berkshire's IBM exit was complete: _"I was wrong, or at least I felt like I was wrong on IBM when I sold it and I was wrong when I bought it."_ Holding-period return figures used in this essay (IBM declined approximately eighteen percent over the holding window; the S&P 500 returned approximately one hundred and sixteen percent over the same window) follow the comparison cited in the Financhill retrospective on the position. See CNBC, "Warren Buffett had a lot to say about Apple and IBM over the years" (May 8 2018), [https://www.cnbc.com/2018/05/08/warren-buffet-had-a-lot-to-say-about-apple-and-ibm-over-the-years.html](https://www.cnbc.com/2018/05/08/warren-buffet-had-a-lot-to-say-about-apple-and-ibm-over-the-years.html), and Financhill, "Why Did Buffett Sell IBM?", [https://financhill.com/blog/investing/why-did-buffett-sell-ibm](https://financhill.com/blog/investing/why-did-buffett-sell-ibm) [5](#footnote-anchor-5) The strongest version of this counter is John Bogle’s lifetime work on the structural advantages of low-cost broad-market indexing, summarised in _The Little Book of Common Sense Investing_ (2007 / 2017). Buffett himself has endorsed the same counter for non-professional investors, most directly in his 2013 Berkshire Hathaway annual letter, where he advised that the trustees of his estate should hold ninety percent of the bequest in a low-cost S&P 500 index fund. The counter binds powerfully on most retail allocators. It does not bind on active allocators with discretion to size and falsify positions, which is the population this essay addresses. See Berkshire Hathaway 2013 Annual Letter, [https://www.berkshirehathaway.com/letters/2013ltr.pdf](https://www.berkshirehathaway.com/letters/2013ltr.pdf), Section: “Some Thoughts About Investing,” and Bogle, _The Little Book of Common Sense Investing_, tenth anniversary edition, 2017. --- --- title: "The Investor's Substrate Test" description: "Score the substrate beneath any single position in seven minutes. A 5-axis profile and a 0-10 score for any holding. Free." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-investors-substrate-test/" date: "2026-04-24" series: "MARKETS & POWER" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Investor's Substrate Test *Score the substrate beneath any single position in seven minutes. A 5-axis profile and a 0-10 score for any holding. Free.* By Harry Floyd · 2026-04-24 · canonical: https://durabilitycurve.com/blog/the-investors-substrate-test/ You can be right about the company and wrong about the position. Most theses get scored. The substrate underneath the thesis rarely does. That gap is where good judgement quietly turns into mediocre returns: the company performs, the position doesn’t, and the post-mortem cannot name what was missing because the thing that was missing was never written down. The Investor’s Substrate Test is a 7-minute, 5-axis tool for scoring the substrate beneath any single holding. You score the position before the news, not after it. Five axes carry the score: durability, asymmetry, replicability, couplings, and optionality. Each contributes 0 to 2, summing to a 0-10 score plus a one-page profile defensible to a partner, an LP, or future-you. **When to run it.** Before sizing a new position. The score becomes part of the position memo and gets re-scored every six months for as long as the position is held. Quarterly on existing positions. A score that drifts down two points in two quarters is the signal that the substrate is eroding while the thesis still looks intact. That is the highest-leverage early warning available to a long-only allocator. **What a good score looks like.** The 0-10 band reads in three regions. Below 4 is a canopy bet: the position depends on the thesis playing out, with little behind it if it doesn’t. 4 to 7 is mixed, usually mispriced one way or the other. Above 7 is a substrate bet: the position holds even when the thesis turns out to be partially wrong, which is the actual claim worth making. The numbers matter less than the forcing function. If you cannot name the substrate in one sentence, the score is not the problem. The thinking is. **Download.** The test is a single PDF. Score one position in seven minutes. Score a portfolio in an afternoon. [Figure] The Investor's Substrate Test 156KB ∙ PDF file [Download](https://harryfloyd.substack.com/api/v1/file/cb429c25-2683-47a2-b429-1c252eff3035.pdf) [Download](https://harryfloyd.substack.com/api/v1/file/cb429c25-2683-47a2-b429-1c252eff3035.pdf) * * * The Investor’s Substrate Test is a companion to _[The Forest Floor Is the Product](https://durabilitycurve.com/blog/the-forest-floor-is-the-product/)_, the article that develops the substrate-vs-canopy distinction across knowledge work, software, ecosystems, and capital allocation. If a term in the test is doing more work than it explains, _[The Lens Lexicon](https://harryfloyd.substack.com/about?utm_source=resource-landing&utm_medium=internal&utm_campaign=substrate-test-2026-04-24)_ (also free) defines the ten terms that carry the framework. _What is the substrate sentence under your largest position?_ --- --- title: "The Forest Floor Is The Product" description: "Most people are optimising for canopy. The work that survives the next five AI releases is built from the forest floor up." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-forest-floor-is-the-product/" date: "2026-04-23" series: "SYSTEMS & LAWS" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Forest Floor Is The Product *Most people are optimising for canopy. The work that survives the next five AI releases is built from the forest floor up.* By Harry Floyd · 2026-04-23 · canonical: https://durabilitycurve.com/blog/the-forest-floor-is-the-product/ On the west coast of Vancouver Island there is a Sitka spruce called the Carmanah Giant. It is around 96 metres tall and several centuries old.[1](#footnote-1) Stand at its base and you think you are looking at a tree. You are not. You are looking at the visible tip of a system that has been quietly assembling itself for far longer than any single tree has been standing in it. Maybe ten percent of it is above the soil line. The other ninety percent sits below. A network of fungal threads, the mycorrhizal network, connecting the old trees through filaments thinner than a human hair.[2](#footnote-2) A soil profile built one millimetre per century from fallen needles, fungal bodies, and the slow decomposition of every tree that came before. A bank of seeds dormant in the soil, waiting decades for the next storm to open a hole in the canopy above. None of that arrived at once. It was laid down in order, one layer at a time. Pioneer species fixed nitrogen in soil that would not otherwise hold a tree. Shade-tolerant species established in their cover. Each generation made the next one possible. The spruce did not grow out of the soil. The spruce and the soil grew each other. Slowly. Over a timescale that makes almost everything we ship in 2026 look like weather. I have thought about that tree a lot this year. Right now the dominant advice is simple: ship faster, the next model release will save you, and anyone not on the treadmill will get left behind. There is some truth in that, the way there is some truth in every panic. But forests give you the case that breaks the rule. Every old-growth ecosystem on the planet exists because no one optimised it for velocity. The annual growth rate of a 400-year-old tree is laughable. What makes it last is not the rate at which it grows. It is the rate at which it does not get displaced. That distinction is going to decide who is still standing in five years. * * * ## Velocity Was Always a Trick of the Light Most strategy advice in 2026 says some version of the same thing: ship faster, automate the boring parts, parallelise the rest, let velocity carry you. It sounds right because it measures the part that is easy to count. Output per week. Posts per month. Features per quarter. The forest gives you the other half no one is measuring. A 400-year-old Sitka spruce adds maybe a centimetre of girth in a good year. By the standards of any ecosystem that prizes growth-rate, this is humiliating. Bamboo can put on close to a metre in a day. Pioneer fireweed colonises a burn site in a single season. By any velocity metric, the spruce loses to almost everything. And yet bamboo groves get cleared for pasture. Fireweed dies back when the canopy closes. The spruce is still standing. > Forests do not optimise for speed. They optimise for the conditions under which a 400-year-old tree can stand. This is the distinction velocity strategies miss. Growth-rate is one variable. Displacement-rate is the other. A piece of work shipped in January 2024 that is still doing useful work in January 2027 has beaten a piece shipped weekly that was obsolete by Friday. The slow tree wins not by growing faster but by occupying a position the fast strategies cannot reach. You can check this in your own work in five minutes. Pick the three things you shipped this month. Then pick the three you shipped this quarter last year. Ask which set is still doing work for you. The honest answer tends to be uncomfortable. The recent stuff is louder. The older stuff is quieter and doing more. The harder admission is that the slow part may be the load-bearing part. If you automate the patience out of the process, or rush past the bit that needs time to set, you may also remove the thing that was making the work durable. Soil built up over a century cannot be shortcut with fertiliser. The shortcut gives you a different kind of forest, and that forest gets cleared. Two clocks are running, not one. The first measures how much you can produce. The second measures how long what you produce keeps doing useful work after you ship it. The first clock speeds up every time a new model ships. The second one barely notices model releases. It only moves when something you made was rooted in enough of you that the next release cannot just regenerate it. Sprint on the first clock and ignore the second, and you can look extraordinarily productive for a year. Then look back and find that almost nothing you shipped is still doing work for you. The move is simple. Stop measuring your work by how often you ship. Start measuring it by how long what you shipped keeps earning its place a month, a quarter, a year later. The first metric flatters output. The second tells you whether you are building anything. * * * ## The Mycorrhizal Network Is the Product Walk into an old-growth forest and ask a forester what the product of the system is. They will not point at the trees. The trees are the canopy. The product is the soil and the fungal network running through it. Suzanne Simard’s 1997 paper in _Nature_ showed something that biologists had suspected for decades and finally proved.[3](#footnote-3) Birch and Douglas fir trees, growing side by side in British Columbia, were exchanging carbon and water through underground fungal connections. Not competing. Trading. The mycorrhizal network linking their roots was acting as one organism, redistributing resources between trees that, viewed from above, looked like separate individuals. This is the wood-wide web. Most land plants on Earth depend on these fungal partnerships, and mature forests function less as collections of individual trees and more as networked systems with a hidden circulatory layer underneath.[4](#footnote-4) Christine Webb, in her essay _We Have Never Been Individuals_, pushes the point further. Once you take the substrate seriously, the category of "individual organism" starts to wobble. The forest is not a population of trees. It is one thing, and the trees are how it shows up at the surface. > The model is the canopy. The substrate is the forest. Now look at your own work the same way. The visible output of anyone using AI in their work, the slide, the post, the codebase, the report, looks like the product. It is not. It is the canopy. The product is the substrate underneath it. The judgement that decided what the slide was even for. The taste that picked the example. The relationship with the person it is being sent to. The accumulated context that made the right move obvious in a situation where a stranger would have needed an hour. Most 2026 advice on AI productivity is canopy advice. Better prompts. Faster generation. More tools. The canopy improves. The substrate gets ignored or, worse, eroded, because the canopy is producing so much output that no one makes time to tend the soil. This is the failure mode that fire-suppressed forests show. The visible canopy looks healthy. The understory is choked. New growth has nowhere to start. When disturbance finally comes, the whole system collapses at once because nothing was ever built underneath. The companies and individuals still standing in 2030 will be the ones who figured this out early. The output is the canopy, and the canopy gets replaced. Every six months a new model ships and the slide deck, the post, the codebase get easier for anyone to reproduce. The substrate underneath is what does not get replaced. It is too specific to you, to your domain, to the people who trust you, to the years of context you brought to bear. The next release does not touch any of it. That is the only part actually worth building. So the move is this. Name the substrate beneath the last three things you shipped. Not the tools you used. Not the workflow you followed. The substrate itself. The hard-won, slow-to-build thing that made the output possible and would not have been there a year ago. If you cannot name it in one clean sentence per project, it is probably not there yet. You are growing canopy on bare rock, and the next storm will take it. Try this honestly and you find something else. The substrate, once you start naming it, is more interesting than the canopy ever was. It is also where every reader you actually care about lives. The output is what got their attention. The substrate is what makes them stay. * * * ## Succession Is a Discipline, Not a Wait We like to imagine forests grew by simply waiting. Time passed. Trees got bigger. Eventually you had a forest. That is not how it happens. A patch of bare ground, whether left by logging, by fire, or by a retreating glacier, does not become a forest at random. It moves through a strict sequence ecologists call succession. The order matters. The tiers cannot be skipped. [Figure] Pioneers arrive first. In the Pacific Northwest these are species like red alder, fireweed, and the hardy shrubs that move into any disturbed lot. They are short-lived and fast-growing, and they do one thing the bare soil cannot yet do for itself. Alder fixes nitrogen, which means it pulls nitrogen out of the air and locks it into the soil so other plants can finally use it later.[5](#footnote-5) Fireweed holds the loose ash and char that fires leave behind. Together the pioneers prepare ground that nothing else could grow on. Then come the early conifers. Douglas fir is the classic one. It takes hold in the soil the pioneers built, in the light the pioneers no longer monopolise. The mid-canopy thickens. The soil deepens. The water cycle stabilises. The forest starts to look like a forest. Only after another century of that, sometimes longer, do the climax species begin their long work. Western hemlock. Western red cedar. Sitka spruce. These are the trees that can grow in deep shade, on deep soil, and that take centuries to reach their final form. They could not have started earlier. The conditions to support them did not exist. > You do not wait for old growth. You sequence the conditions that make it inevitable. The 2026 mistake is to treat durability as a waiting game. As something that happens if you stick around long enough. As patience. It is none of those things. It is a sequenced discipline. People who build durable work, in any field, run succession on purpose. They start with the equivalent of pioneer species in the soil. Raw notes. Half-formed observations. Captures of things they do not yet understand. None of it looks important on its own. They tend that layer until it gets dense enough that something more structured can grow in it. Then come the middle-tier pieces: working analyses, side-by-side comparisons, the first attempts to make sense of what the captures are saying. Only then can the late-canopy work germinate at all: the long-form synthesis, the load-bearing essay, the framework other people start citing. You see the same pattern in research, in product, in writing, in companies. The famous output is the late-canopy tree. The two layers underneath it, mostly invisible, are what made it possible at all. Skip them and you get a sapling planted in clear-cut. It dies. Or worse, it survives long enough to look promising, then dies in the first dry summer. The disciplined version asks a different question about everything you produce. Not “is this important enough to keep?” Pioneers do not look important. Nitrogen-fixers do not look important. The real question is “what is the next layer that this would make possible?” If the answer is nothing, you have grown a piece of fireweed. Fine. Note it and move on. If the answer is something, you have laid down a millimetre of soil, and the next thing planted in it will go further. This is a hard reframe to swallow because almost no incentive structure in 2026 rewards pioneer-tier work. There are no metrics for soil-building. There is no leaderboard for the work nobody can see yet. The reward structure prizes the late-canopy tree, the visible output, the thing that can be screenshotted. Anyone who wants old-growth work has to pay for the underlayers themselves, with their own time, against the gravitational pull of a culture that only ever looks at the canopy. The pay-off, when it comes, is disproportionate. A late-canopy synthesis grown out of a thick understory is not the same animal as a sapling planted in bare ground. The first is rooted in a system. The second is decoration. * * * ## Disturbance Is How New Growth Happens A forest that never gets disturbed slowly chokes itself. This is one of the most counterintuitive results in 20th-century forest ecology. The intuitive model says: protect the forest, suppress the fires, keep the loggers out, and you will have a healthy ecosystem. That was the dominant US Forest Service policy for most of the 20th century. The result was forests that were denser, more uniform, more disease-prone, and far more vulnerable to catastrophic fire than the ones the policy was meant to protect.[6](#footnote-6) What was missing was disturbance. Specifically, the small, frequent, gap-creating disturbance that opens a hole in the canopy and lets new species establish. A 40-metre tree falls in a storm. Light hits the forest floor for the first time in two centuries. Seeds that have been sitting dormant in the soil germinate. New species establish. The gap closes over thirty years. Net result: the forest stays younger in patches, more diverse, more resilient. Without disturbance, the canopy holds. The same handful of dominant species keep their position. The understory, the layer of smaller plants growing under the canopy, thins out. New growth has nowhere to go. The forest looks fine for fifty years and then collapses all at once. > AI is not the storm. AI is the canopy gap. Different question, different answer. The default 2026 reading of AI is the storm reading. The canopy is collapsing, the dominant species are about to be displaced, panic is the right response, and the only move left is to grab a chainsaw and join the felling. That reading is wrong in a precise way. It treats the disturbance as terminal rather than generative. The forest reading is different. AI is a canopy gap. A patch of light has opened that was not there before. Species that were dormant in the seed bank for decades, because there was no light for them, can now germinate. Species that already dominated the canopy are not necessarily displaced. They no longer monopolise the light. Whether the gap closes back into the same forest or grows into a different one depends on which seeds were already in the soil. This changes the question you ask about your own work. Not “how do I survive AI?” Try this instead: which seeds were sitting in my seed bank that this canopy gap finally lets me plant? Some of these will be projects you would not have started without AI. Some will be species of work you abandoned years ago because the canopy was too closed for them to survive. Some will be whole categories of practice that stayed dormant because no one had the substrate to grow them in. You can do this audit on a single sheet of paper. Three projects you would not have started without AI. Three projects you should have stopped because they were filling space that could now be a gap. Three things in your seed bank that have been waiting for light. The output of this exercise is rarely what you expect. Most people find that the projects they would have stopped are obvious in retrospect and were obvious before AI arrived. They were being kept alive by inertia. The disturbance reveals what was already not working. The mature operator does not chase disturbance. They read it. They wait long enough to see which gap opened, which species the disturbance favoured, and which dormant seeds in their own seed bank are now viable. Then they plant. The species that fill a canopy gap fastest are not always the species still standing in fifty years. The 2026 winners will be the people who can tell the difference and plant for the second forest, not the first. * * * ## The Forest That Knows It Is a Forest The knowledge system underneath this essay is itself an old-growth forest, by my reckoning. It is built around five claims that show up, independently, in field after field. They were not invented for this article. They were derived, over thousands of analyses, from patterns that kept appearing across AI research, neuroscience, financial markets, biology, and history. I have written about them before. They are the backbone of every other piece in this newsletter. Read them as a forester would, and they stop sounding abstract. Law I says the bottleneck always migrates. In the forest this is succession. Whatever is limiting the system moves up the stack as each layer matures. First the soil is the limit. Then it is light at the floor. Then it is competition for canopy space. Then it is the carrying capacity of the fungal network. Whatever was limiting last decade is not what is limiting now. The forester still fertilising soil that is no longer the bottleneck is wasting effort, in exactly the same way the operator still optimising last year’s layer is wasting theirs. Law II says difficulty is load-bearing. In the forest this is the soil profile. The slow accretion, one millimetre per century, is what makes a 400-year-old tree possible. There is no shortcut soil. There is no fertiliser that produces old-growth substrate. The difficulty is doing a job, and removing it does not accelerate the forest. It produces a different forest, one that gets cleared in a generation. Law III says architecture outlives content. In the forest this is the mycorrhizal network. Individual trees come and go over centuries. The fungal network beneath them persists across generations of trees.[7](#footnote-7) The trees are content. The network is architecture. Forests that lose their fungal substrate, through clear-cutting or soil disturbance, do not grow back the same way even when you replant the trees. The architecture is gone, and the content has nowhere to root. Law IV says knowledge is constrained by instruments, not theory. In the forest, your instrument was your eye, and your eye could only see the canopy. For most of human history, foresters knew the trees and not the soil, because the trees were the only thing the available instrument could resolve. Soil cores, isotope tracers, modern mycology: those gave us the substrate. The substrate was always there. The instrument was new. Every field has a moment like that. AI is one of them now, and the operators building the right instruments will see what the canopy-only operators cannot. Law V says capability without correct targeting makes things worse. In the forest, this is canopy gap targeting. A storm opens a gap. If you respond by planting more of the species that already dominated the canopy, you have added capability aimed at where the old forest was, not where the new gap is. That is most 2026 AI strategy in one sentence: capability poured into the layer of work that just got commoditised. The targeting is wrong. More capability does not fix wrong targeting. It accelerates the misalignment. > The five laws are not five laws. They are one forest, viewed from five angles. The reason this matters is that almost every operator in 2026 is reading their situation one law at a time. They are optimising the bottleneck where it used to be. They are removing friction that was doing structural work. They are investing in tools and ignoring the architecture underneath. They are theorising harder instead of building better instruments. They are adding capability aimed at the wrong target. Any one of those damages the system. All five together give you a forest that looks productive for a year and is gone in three. The five-law reading is one diagnostic, not five. Read your situation as a forest, then ask. Where has the bottleneck moved to? Is the friction you want to remove doing structural work? Is your effort going into the architecture underneath, or the canopy on top? Do you have the instrument to see the substrate? Is your capability aimed at the gap that just opened, or the one that closed years ago? These are not slogans. They are a single diagnostic, applied to one system at a time, in writing. * * * ## The One-Week Test Pick one project you shipped in the last 90 days. Not your favourite. One that was real. Open the file. Read it again. Then sit with three questions, one paragraph of writing each. _What is the substrate this output is sitting on?_ Not the tools, not the framework, not the model. The accumulated thing: the judgement, the taste, the relationships, the context that made this specific output possible and would not have been there a year ago. If the honest answer is “I do not know,” good. That is an honest answer and it is telling you something. _What is the canopy gap that created the conditions for this project to exist now?_ What disturbance opened the light it grew into? If the project would have been impossible two years ago, name what changed. If the project would have been possible two years ago and you only got to it now, name what was holding the canopy closed. _If the model that helped you write this output were obsolete tomorrow, what would survive?_ The substrate or the canopy? The output or the architecture underneath it? The visible thing or the slow accretion that produced it? Most operators cannot answer all three the first time they run this test. That is not failure. That is the diagnostic. The point is not the answers. It is noticing which question comes hard. The forced one is where the canopy is hiding the soil. That is the layer you cannot yet see, which usually means it is also the layer you are not yet tending. The good news is that the underlayers are slow to build and slow to lose. A year of deliberate substrate work is worth more than a decade of reactive canopy chasing. The forest is patient that way. Once the soil is built, it holds. The canopy is what gets the light. The forest is what is still standing in 400 years. * * * _If you ran the test this week, I want to know which of the three questions was the one you could not answer. That will tell me which layer of the forest most readers cannot yet see, which will tell me which piece to write next. Comments are open below._ _The longer self-assessment version of this test, called The Forest Floor Audit, takes about thirty minutes and walks through the five layers in detail. Free, linked at the foot of this post._ [1](#footnote-anchor-1) The Carmanah Giant grows in Carmanah Walbran Provincial Park on the west coast of Vancouver Island, in Ditidaht territory. It is widely cited as the tallest tree in Canada at approximately 96 metres (315 ft). Published age estimates vary considerably, ranging from under 400 years to around 700 years, with no single authoritative dendrochronology on record. See Ancient Forest Alliance, “Conservationists locate and climb the largest Sitka spruce tree in BC’s famed Carmanah Valley” (2024), [https://ancientforestalliance.org/climbing-carmanah-valley-largest-sitka-spruce/](https://ancientforestalliance.org/climbing-carmanah-valley-largest-sitka-spruce/), and BC Geographical Names, [https://apps.gov.bc.ca/pub/bcgnws/names/41299.html](https://apps.gov.bc.ca/pub/bcgnws/names/41299.html) [2](#footnote-anchor-2) Mycorrhizal hyphae are typically 2 to 10 micrometres in diameter, around an order of magnitude thinner than a human hair, which averages 50 to 100 micrometres. The density of fungal mycelium in mature temperate forest soils has been measured at hundreds of metres of hyphae per gram of soil. See Read, D.J. & Perez-Moreno, J., “Mycorrhizas and nutrient cycling in ecosystems — a journey towards relevance?” _New Phytologist_ 157 (2003), [https://doi.org/10.1046/j.1469-8137.2003.00704.x](https://doi.org/10.1046/j.1469-8137.2003.00704.x) [3](#footnote-anchor-3) Simard, S.W., Perry, D.A., Jones, M.D., Myrold, D.D., Durall, D.M. & Molina, R., “Net transfer of carbon between ectomycorrhizal tree species in the field,” _Nature_ 388, 579 to 582 (1997), [https://doi.org/10.1038/41557](https://doi.org/10.1038/41557). The original demonstration of bidirectional carbon transfer between birch and Douglas fir through shared mycorrhizal networks. Simard’s later work, including the 2016 TED talk _How Trees Talk to Each Other_, brought the wood-wide web into wider awareness. The strong forms of the wood-wide-web hypothesis have been challenged in recent years, notably Karst, J., Jones, M.D. & Hoeksema, J.D., “Positive citation bias and overinterpreted results lead to misinformation on common mycorrhizal networks in forests,” _Nature Ecology & Evolution_ 7 (2023), [https://doi.org/10.1038/s41559-023-01986-1](https://doi.org/10.1038/s41559-023-01986-1). The narrower claim used in this essay, that adjacent trees can exchange resources through shared fungal connections under measured conditions, remains well-evidenced. [4](#footnote-anchor-4) Estimates of the proportion of land plants forming mycorrhizal associations range from 80 to 92 percent depending on methodology and habitat. See Wahab, A. et al., “Role of Arbuscular Mycorrhizal Fungi in Regulating Growth, Enhancing Productivity, and Potentially Influencing Ecosystems Under Abiotic and Biotic Stresses,” _Plants_ 12 (2023), [https://doi.org/10.3390/plants12173102](https://doi.org/10.3390/plants12173102). For the philosophical reframe of forests and humans as networked rather than individual, see Webb, C., _We Have Never Been Individuals_, _Behavioral Scientist_, [https://behavioralscientist.org/we-have-never-been-individuals/](https://behavioralscientist.org/we-have-never-been-individuals/), and Webb, C., _The Arrogant Ape: The Myth of Human Exceptionalism and Why It Matters_ (2025). [5](#footnote-anchor-5) Red alder (_Alnus rubra_) hosts nitrogen-fixing _Frankia_ bacteria in root nodules and can fix 100 to 300 kg of nitrogen per hectare per year in Pacific Northwest forests, materially altering soil chemistry within a single rotation. See Binkley, D., Sollins, P., Bell, R., Sachs, D. & Myrold, D., “Biogeochemistry of adjacent conifer and alder-conifer stands,” _Ecology_ 73 (1992), [https://doi.org/10.2307/1941452](https://doi.org/10.2307/1941452) [6](#footnote-anchor-6) The shift away from total fire suppression in US Forest Service policy followed the 1995 Federal Wildland Fire Management Policy review, which formally recognised that decades of suppression had produced denser, more vulnerable forests. See Stephens, S.L. & Ruth, L.W., “Federal forest-fire policy in the United States,” _Ecological Applications_ 15 (2005), [https://doi.org/10.1890/04-0545](https://doi.org/10.1890/04-0545) [7](#footnote-anchor-7) Estimates of the proportion of land plants forming mycorrhizal associations range from 80 to 92 percent depending on methodology and habitat. See Wahab, A. et al., “Role of Arbuscular Mycorrhizal Fungi in Regulating Growth, Enhancing Productivity, and Potentially Influencing Ecosystems Under Abiotic and Biotic Stresses,” _Plants_ 12 (2023), [https://doi.org/10.3390/plants12173102](https://doi.org/10.3390/plants12173102). For the philosophical reframe of forests and humans as networked rather than individual, see Webb, C., _We Have Never Been Individuals_, _Behavioral Scientist_, [https://behavioralscientist.org/we-have-never-been-individuals/](https://behavioralscientist.org/we-have-never-been-individuals/), and Webb, C., _The Arrogant Ape: The Myth of Human Exceptionalism and Why It Matters_ (2025). * * * ## The Forest Floor Audit _The longer self-assessment version of the one-week test_ [Figure] Forest Floor Audit 207KB ∙ PDF file [Download](https://harryfloyd.substack.com/api/v1/file/fb63bf68-cbc4-4104-839d-3d22271c8b98.pdf) [Download](https://harryfloyd.substack.com/api/v1/file/fb63bf68-cbc4-4104-839d-3d22271c8b98.pdf) --- --- title: "You're Not Comparing Models. You're Comparing Contracts." description: "Agent benchmarks don't measure models. They measure contracts. Two teams running the same model can publish different scores, and both can be honest." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/youre-not-comparing-models-youre/" date: "2026-04-19" series: "PROOF & TRUST" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # You're Not Comparing Models. You're Comparing Contracts. *Agent benchmarks don't measure models. They measure contracts. Two teams running the same model can publish different scores, and both can be honest.* By Harry Floyd · 2026-04-19 · canonical: https://durabilitycurve.com/blog/youre-not-comparing-models-youre/ Two teams publish scores on the same agent benchmark. One lands in the low sixties. The other clears seventy. A procurement team reads the spread and makes a call. What they do not see: both teams may be running the same model. They did not need to change the weights for the gap to appear. The spread can come from scaffold alone. One team wrapped the model in a harness with better retries. Different tool defaults. A planner step the other team had skipped. None of that appears on the leaderboard. The comparison that drove the decision was not between two agents. It was between two contracts. ## There Is No Benchmark The mistake hiding behind this story is a category error. People talk about agent benchmarks as if they measure a thing called “the model.” They do not. They measure a coupled system. The model is one component. The rest is a stack of protocol decisions that are almost never disclosed and almost always matter. The score is the output of that stack. Change any layer and you change what the number means. Recent research on agent evaluation has named those layers explicitly. There are at least seven. Deployment regime. Observation channel. Harness and scaffold. Metric and action. Configured evaluator. Grader protocol. Audit bundle. Each is a contract. Each is negotiable. And each can silently change the verdict while the headline looks the same. > That is what a benchmark actually is. Not a measurement of a model. A measurement of an entire testing contract, of which the model is one slot. There is structural reason the seven layers are the seven layers. They cluster into three corners that show up in almost every published agent-evaluation failure. What the model is rewarded for. How that reward is optimised. And how the test contract differs from production. Once you hold those three corners in view, the seven-layer stack stops feeling like a checklist and starts behaving like the actual shape of what is being measured. If you are comparing agent products without parity across those layers, you are not comparing agents. You are comparing contracts and calling it science. ## The Harness You Didn’t Name The most visible layer, and the one that moves the most points, is the scaffold. Anyone who has built an agent in the last year has felt this without naming it. You watch a coworker get 75% on a task your model just failed on. You check the weights. They are yours. They changed the prompt template and added a retry loop. The model did not get smarter. The scaffold got thicker. The numbers say the same thing. RWE-bench reports that on its 162-task benchmark over MIMIC-IV, changing only the agent scaffold around a fixed model can shift performance by more than 30 percent[1](#footnote-1). Same weights. Different tools. Different retry policy. Different planner. Different headline. The best evaluated agent on that benchmark reaches around 40 percent task success at all; the best open-source configuration is closer to 30. Once you know the contract can move 30 points on its own, neither of those numbers is really about a model. If scaffold alone can move scores by double digits, then “we used the same model as them” is not a fair-comparison claim. It is a parameter-naming claim. > You have named one slot in a seven-slot contract. > The other six are doing most of the work. ## The Judge That Isn’t The Model The second layer that silently moves scores is the evaluator itself. When a benchmark uses an LLM judge, people write things like “graded by GPT-4o” as if that pins the measurement down. It does not. The judge is not GPT-4o. The judge is GPT-4o plus a prompt template. Plus a decoding configuration. Plus a tie and abstention policy. Plus whatever retrieval or tool access the judge has during grading. None of that ships with the score. A recent systematic evaluation of LLM-as-judge setups showed that prompt-template choice alone materially changes both judge quality and internal consistency[2](#footnote-2). Two teams reporting “we used GPT-4o as judge” can be running substantively different graders. The grader that rewards epistemic hedging disagrees with the grader that penalises it. The grader with access to retrieval checks factuality. The grader without one does not, and cannot. This is not a small print issue. > The evaluator is the measuring instrument. If two teams use different instruments and report the same number, they are not reporting the same thing. And without a published judge card, no third party can reproduce the measurement. They can only rerun the model. ## The Number That Lies About Consistency The third layer is the quietest and most dangerous. It is the metric itself. A standard agent metric is pass@k. You give the agent k attempts. If any one succeeds, it counts. This is perfectly reasonable if your production use allows k attempts. It is actively misleading if it does not. There is a sibling metric, pass^k. Same k attempts. But it only counts if the agent succeeds on all of them. It measures consistency, not capability. The gap between these two can be large, and it can open silently. > Recent work on trustworthy agent evaluation shows that controlled error injection into an agent can cut pass^k substantially while barely moving pass@k[3](#footnote-3). The model still has a ceiling you can hit with enough tries. It has lost the ability to hit that ceiling reliably. If your headline is pass@k and your production regime is one shot, the leaderboard says you are shipping. The bug tracker says otherwise. The same structural problem appears in calibration metrics. ECE asks whether stated probabilities match empirical frequencies on average. AURC asks whether the system can rank harder cases lower. Both can look nearly identical across two systems while a stricter, abstention-aware metric called BAS, the Behavioural Alignment Score, diverges sharply between them[4](#footnote-4). BAS asks a different question. Does the confidence surface protect you in exactly the regime where a person or product would actually choose to trust it? Two systems with “similar calibration” can answer that question completely differently once you attach a cost function. The metric is not a measurement of the model. It is a statement about which errors the model’s operators will tolerate. If that statement does not match your operational contract, the score is not wrong. It is answering a question you did not ask. ## Why Rank Stability Is A Trap Here is the part that makes all of this subtly worse. Under scaffold shift, the rank order of agents on a benchmark is often relatively stable. A recent efficient-benchmarking study reports that rank preservation is easier to maintain than absolute calibration[5](#footnote-5). The number moves. The ordering does not. If all you need is a relative decision, rank stability is comforting. Agent A beats Agent B here, and probably beats it in production. If you need an absolute decision, it is a trap. Procurement, safety arguments, SLA setting, cost modelling, and risk disclosure all depend on absolute numbers. A claim like “this agent ships 80% correct at 5 cents per request” binds to the calibrated level, not to the rank. Under scaffold shift, rank can hold while the 80% becomes 62%. Your spreadsheet is still using 80%. Your customers are experiencing 62%. The protocol that produced 80% is part of the claim. The moment it diverges from production, the claim silently becomes false, even though nothing about the model moved. This is why seasoned eval teams treat the contract, not the score, as the primary artefact. > You can rerun a score. You can only reproduce a contract if you wrote it down. ## The Contract Is The Object If the score is a function of the contract, the practical move is to treat the contract as the thing you own. That means three changes to how most teams currently work. Freeze the contract before you compare. If you cannot describe your deployment regime, observation channel, scaffold version, metric, judge configuration, and grader protocol in one page, you do not have a contract. You have assumptions pretending to be one. Write the page. Commit it. Make it a prerequisite for every comparison. Version the contract the way you version weights. When the scaffold changes, the contract version changes. When the judge template changes, the contract version changes. When the metric changes, the contract changes, period. > A benchmark result without a contract version is not a result. It is a rumour. Publish a minimum audit bundle with every reported number. At minimum: the harness, the judge card, the rubric, a sample of trajectories, and the metric definitions used. This is not bureaucratic overhead. It is the only thing that lets a third party tell whether your score is comparable to anyone else’s. Without it, every comparison is a faith-based transaction. Teams that do this do not ship agents faster. They ship agents that mean the same thing next quarter as they meant this quarter. That is a different product. It is also, increasingly, the only one that compounds. ## The Real Question Generation is getting cheaper every month. Models are getting better every month. Scaffolds are getting richer every month. All of that pushes the same direction. It makes raw agent capability more abundant and less differentiating. > The scarce resource is not the agent. It is whether you can say, with a straight face, what you measured. Most of the public agent numbers flying around right now do not survive that question. The scaffold is implicit. The judge is underspecified. The metric does not match the action. The audit bundle is missing. Rank stability is quietly being load-bearing on decisions that only absolute calibration can support. The operators who will build durable advantage in the next two years are not the ones with the best agent. They are the ones who own the contract under which “best” means anything at all. Here is the test worth running this week. Pick the most recent agent score your team has cited in a decision. Try to describe the contract that produced it on a single page. Deployment regime. Observation channel. Scaffold version. Metric definition. Judge configuration. Grader protocol. Audit bundle. **Seven slots.** If you can fill all seven, you have a result. If you cannot, you have a rumour your spreadsheet is treating as a number. Every downstream decision is borrowing the rumour’s confidence. Most teams cannot fill all seven the first time they try. That is the honest finding. Not that your model is wrong. Not that your benchmark is wrong. The thing you thought you measured has been sitting underneath the number all along, and the number is just **the part you could see.** _If you actually run that test this week: which of the seven slots was hardest to fill? I’d genuinely like to know which layer most teams cannot describe. Leave it in the comments._ [1](#footnote-anchor-1) [Li et al,](https://arxiv.org/abs/2603.22767) _[RWE-bench: A Real-World Evidence Benchmark for LLM Agents on MIMIC-IV](https://arxiv.org/abs/2603.22767)_[, arXiv 2603.22767](https://arxiv.org/abs/2603.22767) [2](#footnote-anchor-2) [Wei et al,](https://arxiv.org/abs/2408.13006) _[Systematic Evaluation of LLM-as-a-Judge](https://arxiv.org/abs/2408.13006)_[, arXiv 2408.13006](https://arxiv.org/abs/2408.13006) [3](#footnote-anchor-3) [Ye et al,](https://arxiv.org/abs/2604.06132) _[Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents](https://arxiv.org/abs/2604.06132)_[, arXiv 2604.06132](https://arxiv.org/abs/2604.06132) [4](#footnote-anchor-4) [Wu et al,](https://arxiv.org/abs/2604.03216) _[A Decision-Theoretic Approach to Evaluating Large Language Model Confidence](https://arxiv.org/abs/2604.03216)_[, arXiv 2604.03216](https://arxiv.org/abs/2604.03216) [5](#footnote-anchor-5) [Ndzomga,](https://www.semanticscholar.org/paper/eeff713967bfd94da50e4cc8fda888b01f137e90) _[Efficient Benchmarking of AI Agents](https://www.semanticscholar.org/paper/eeff713967bfd94da50e4cc8fda888b01f137e90)_[, Semantic Scholar eeff7139 (arXiv 2603.23749)](https://www.semanticscholar.org/paper/eeff713967bfd94da50e4cc8fda888b01f137e90) --- --- title: "Prompting Isn't Writing, It's Compilation" description: "If you treat the prompt as a spec and the model as a renderer, quality stops coming from more words and starts coming from better constraints." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/prompting-isnt-writing-its-compilation/" date: "2026-04-17" series: "AI & WORK" substack: "https://open.substack.com/pub/harryfloyd/p/prompting-isnt-writing-its-compilation?utm_campaign=post-expanded-share&utm_medium=web" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Prompting Isn't Writing, It's Compilation *If you treat the prompt as a spec and the model as a renderer, quality stops coming from more words and starts coming from better constraints.* By Harry Floyd · 2026-04-17 · canonical: https://durabilitycurve.com/blog/prompting-isnt-writing-its-compilation/ I keep reading prompt tips that assume the way to get better output is to write more. More adjectives. More moodboard language. More camera jargon. More cinematic this, ethereal that. The working assumption is that a prompt is an essay, and a good prompt is a well-written essay. So people iterate on the sentence. They swap in fancier words. They add qualifiers. They layer on references. Most of the time, the output does not get better. It just gets noisier. The mistake is upstream of the writing. A prompt is not an essay. It is a specification. And once you see it that way, the job changes completely. ## The Prompt And The Spec A spec is not judged by how it reads. It is judged by whether the thing it asks for can be produced reliably. That is a very different discipline. When you write an essay, you add words to make a point clearer. When you write a spec, you remove options to make an outcome reproducible. In image generation, most of the real quality lives inside a small set of non-obvious decisions. What kind of image is this actually for. What must stay true across every seed. What is free to vary. What the camera is doing. What the aspect ratio is. How much stylisation to allow. What to explicitly exclude. Those are not adjective choices. They are constraint choices. And constraints are not prose. They are parameters in a program the model runs. If you accept that framing, the writing part of prompting quietly stops being the interesting part. The interesting part is the compiler that turns small input into the right constrained program. ## What This Looks Like In Practice I built one of these for Midjourney. It is small, deterministic, and boring in the way good tools are boring. You give it a tiny input. Just enough to know what you want. A subject. Optionally, the asset job. The compiler does the rest. It routes your input to one of about ten regimes. Cover art. Thumbnail. Vertical poster. Blog header. Product photo. Portrait. Concept art. Logo. UI mockup. Texture. Each regime carries its own defaults. A product photo wants a square crop, low stylisation, low chaos, camera-and-material language, and realistic lighting. A concept art prompt wants a widescreen crop, more stylisation, room for exploration, and strong atmosphere. A logo wants the opposite: square, very low stylisation, strong exclusions against mockup and render cues. A UI mockup wants layout language first and aesthetic language last. Nobody reading a single prompt could see all of that. It is not in the sentence. It is in the spec the compiler is quietly loading behind it. Once the regime is chosen, the compiler fills the constraint stack in a fixed order. Subject. Framing. Environment. Style anchor. Palette. Exclusions. Then it attaches the parameter policy for that regime. Aspect ratio. Stylise. Chaos. Model version. Negative flags. The output is a prompt, yes. But the prompt is a side-effect. The real product is the decision about which levers to move at all. [Figure] _The five layers a working prompt compiler actually has. None of them live in the prompt text._ A compiled prompt looks like this: ```plaintext A vault knowledge concept-art image, dramatic atmospheric lighting, a lone reader at the centre of a cathedral of interlinked pages, small-temple scale, painterly cinematic style, deep navy palette with warm gold highlights --ar 16:9 --stylize 180 --chaos 12 --v 7 --no text, ui, watermark ``` Concept-art regime. Subject, frame, world, style anchor, palette, exclusions, parameter policy. Six lines. Every line is a decision the compiler made before any text was written. The “writing” part took about thirty seconds because by that point the only live choices were the subject and the palette. Most of the work happened upstream of the sentence. * * * ## Repair Beats First-Pass Here is the part most prompt writing misses. The first render is not where quality lives. Quality lives in how you respond when the first render is wrong. When a prompt fails, people usually rewrite the whole thing. They replace adjectives. They try a different mood. They add three more reference artists. They change everything at once and learn nothing from the result. A compiler does not do that. A compiler classifies the failure and applies one known correction. If the composition is wrong, it changes the framing term and leaves everything else alone. If the style is drifting, it strengthens the style anchor. If the output is too random, it drops the chaos parameter. If the output is too bland, it nudges stylisation up by one step. If a logo comes back looking like an illustration, it adds the vector and symbol constraints that logos always need. These are not clever hacks. They are rules. The moment you can name the failure, you know which lever to move. And you move one lever at a time, so every iteration actually tells you something. The difference between random-walking through adjectives and compiling-then-repairing is not subtle. One drifts. The other converges. * * * ## The Moat Isn’t The Prompt Text This is the part I keep wanting to tell people who are trying to build an edge with image generation. The prompt text does not matter. Anyone can prompt. Anyone can paste a screenshot of someone else’s output and ask a model to reverse-engineer the sentence. Prompts are, by design, the most copyable thing in the workflow. The constraint library is not. The tasteful defaults are not. The boundary between regimes is not. The repair taxonomy is not. The style packs you have curated and stress-tested are not. The library of bundles you have actually verified produce work you would ship: which constraint stack, which aspect ratio, which style reference, all flagged useful by your own eye. That is the moat. It is the scaffold, not the content. It is also invisible from the outside. You cannot read the final image and infer any of it. The only way to get there is to sit with a lot of outputs, name what failed, and write the rule down. Most people will not do that. Most people will keep writing better adjectives. * * * ## What This Means If You’re Building If you are using these tools seriously, the question is not how to write better prompts. The question is what your compiler looks like. A real compiler has a small input contract. It has a regime router. It has a constraint stack you can name. It has a parameter policy per regime. It has a repair taxonomy. It has a memory of bundles that worked. You do not need to ship software to have one. You can write the rules down in a note. You can keep a regime table on a page. You can make a small decision flow that says, given this input, here is the specification you are going to render against. If you do, two things change. Quality stops being a function of how eloquently you describe the scene. It starts being a function of whether you chose the right regime, whether your constraints are consistent, and whether your repair rules work. And iteration stops being a random walk. You start learning from every failure, because you are only changing one variable per generation. Neither of those things happen when you are treating the prompt as a sentence to polish. * * * ## The Real Question Most people running these tools today are still prompt writers. They are sitting inside the model, writing increasingly clever sentences, hoping the next one finally cracks the image they can already see in their head. The operators who will build durable advantages here are doing something different. They are sitting one level up. They are writing the compiler. They are keeping regimes, constraints, parameters, exclusions, and repair rules as structured assets. They are updating them every time they learn something. They are making the prompt a side-effect of a specification, rather than the thing they are trying to perfect. A prompt is cheap. A compiler compounds. The question is not which one you have today. The question is which one you are quietly building. The cover on this piece follows the same logic: the spec was compiled once, then rendered in Grok rather than Midjourney because that renderer matched the regime better. * * * _If you are building one: what regimes do you have? What is in your repair taxonomy? Which constraint did you have to learn the hard way? I’d genuinely like to read what other people have ended up with — leave it in the comments._ * * * --- --- title: "The Model Is Not the Moat" description: "If frontier capability keeps centralising, the durable edge shifts outward into trust, workflow fit, and the surrounding package." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/the-model-is-not-the-moat/" date: "2026-04-14" series: "STRATEGY & MOATS" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # The Model Is Not the Moat *If frontier capability keeps centralising, the durable edge shifts outward into trust, workflow fit, and the surrounding package.* By Harry Floyd · 2026-04-14 · canonical: https://durabilitycurve.com/blog/the-model-is-not-the-moat/ I keep hearing the same assumption underneath AI strategy talk: the winner will be whoever has the strongest model. Smarter model wins. Everything else is secondary. That sounds plausible right up until you look at how people actually choose tools in real life. If scale laws keep holding, local models probably will not beat the frontier on raw intelligence. But users do not adopt “raw intelligence.” They adopt something that fits into their day. They adopt the thing that feels fast enough, private enough, legible enough, reliable enough, cheap enough, and integrated enough that using it becomes natural. That is the first distinction that matters. The benchmark measures the visible product. The moat forms one layer out, in the surrounding package. ## The Product And The Seat This is where a lot of AI debate still feels strangely naive. People compare models as if users experience them as isolated intelligence engines. Most of the time they do not. They experience a bundle: interface, defaults, permissions, speed, memory, privacy, cost, and how much the tool asks them to rearrange their behaviour. That bundle is where trust accumulates. It is also where switching costs quietly form. A frontier model can win the benchmark and still lose the seat. By “seat” I mean the position a product earns inside a user’s actual workflow. The place where their context lives. The place they trust not to embarrass them, leak data, slow them down, or force them to relearn everything. The product is the thing they evaluate. The seat is the thing they get used to living in. And those are not the same asset. ## You Can Already See It In Coding Tools A model that is slightly worse on a public benchmark can still be the one people prefer if it lives inside the editor, sees the repo, responds instantly, keeps sensitive code local, and fits the way they already work. The model may be weaker in the abstract. The package is stronger where it counts. That gap matters because people do not make adoption decisions in the abstract. They make them under workflow pressure. They choose the tool that keeps momentum, feels trustworthy, and does not introduce a new category of risk. This is also why “just use the best model” is often bad product advice. The best model according to a leaderboard may carry the wrong latency, the wrong privacy posture, the wrong integration burden, or the wrong failure mode for the actual job. ## How Moats Actually Form Apple is the familiar version of this dynamic outside AI. The moat was never just one visible feature. It was the package: ecosystem fit, convenience, defaults, identity, and the low-grade friction of leaving. The product got attention. The package became hard to leave. I think a lot of AI products will work the same way. Moats usually form in the residue, not in the headline claim. In habit. In muscle memory. In stored context. In predictable behaviour. In the feeling that this tool understands how you work and does not make you pay a tax every time you use it. That is the important shift. A lot of the value is created by side-effects of repeated use, not just by the explicit capability being marketed. ## What This Means For Builders Local models do not need to win the intelligence race to matter. They can win a different race entirely: trust, control, governance, latency, privacy, and workflow fit. That is a more durable position than it sounds. Once a tool becomes the place where your context lives, your defaults settle, and your work starts to flow, a better benchmark somewhere else is not enough to dislodge it. If I were building in this market, I would treat that as a design rule: Do not ask only, “How do we make the model look stronger?” Ask: - Where does the user feel risk right now? - What part of the workflow still feels awkward or fragile? - What context would make the tool more useful after 30 days than on day 1? - What would make leaving this product feel expensive in a good way? That is how you build the seat. Not by winning one benchmark snapshot, but by creating a package that compounds through use. ## The Real Question This is the distinction I keep coming back to: The product is what gets measured once. The seat is what gets harder to leave over time. That is why packaging matters. Not because it disguises weakness. Because it is where practical advantage compounds. If frontier capability stays centralised, that is the strategic question for everyone else. What seat are you building that people will not want to leave? --- --- title: "I Read 3,000 Papers Across 12 Fields. Five Patterns Kept Appearing." description: "Every field discovers them independently. Nobody connects them." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/i-read-3000-papers-across-12-fields/" date: "2026-04-05" series: "SYSTEMS & LAWS" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # I Read 3,000 Papers Across 12 Fields. Five Patterns Kept Appearing. *Every field discovers them independently. Nobody connects them.* By Harry Floyd · 2026-04-05 · canonical: https://durabilitycurve.com/blog/i-read-3000-papers-across-12-fields/ Same AI model. 6.7% accuracy. Change the interface around it. 68.3%. The researchers didn’t touch the model. They changed the format it used to express edits. The model had been reasoning correctly the whole time. It just couldn’t express its answers without corrupting them. The failure looked like stupidity. It was a formatting problem. I would have ignored this if I hadn’t seen the same pattern in four other fields that week. * * * I run a research system that cross-references everything it reads. 3,000+ sources across AI, neuroscience, markets, biology, history, and physics. Most of the time it produces noise. But occasionally it surfaces something no single field can see on its own. Five patterns kept appearing. Same structural dynamic, same failure modes, same counterintuitive outcomes. Independently. Across twelve domains. This matters now. We’re in the middle of the largest capability explosion in history, and most of the responses to it are making things worse. The patterns governing what happens next are invisible from inside any single field. You have to hold them all in the same frame. Each one comes with a question you can use immediately. * * * ## The Bottleneck Migrates. It Never Disappears. Most people believe success is straightforward: find the bottleneck, fix it, win. The bottleneck moves. The model story above is the cleanest example. The bottleneck wasn’t reasoning capability. It was the verification layer around the model. Fix the interface, and the “dumb” model becomes the best in the benchmark. In mathematics, proof assistants like Lean matter not because they generate proofs but because they verify them. Proof generation is getting cheaper. Proof verification is the binding constraint. Terence Tao has been making this point for years. In content, AI drives the cost of producing text toward zero. So the bottleneck migrates from writing to editing. From editing to taste. From taste to distribution. From distribution to trust. Each solution creates the next scarcity. In markets, information became free decades ago. The bottleneck migrated from access to interpretation, then from interpretation to execution discipline. The binding constraint for most traders isn’t finding an edge. It’s sitting still long enough to let it work. Whatever just became easy is no longer where the value is. If you’re still optimizing there, you’re solving yesterday’s problem. **Ask:** Where is the bottleneck migrating to in your system? Not where it is now. Where it’s going. * * * ## Difficulty Is Load-Bearing. People believe friction is waste. That making something easier always makes it better. That removing the hard parts is progress. Sometimes it is. But sometimes you’re pulling out a load-bearing wall, and you won’t know until the roof comes down. Rome’s republic died this way. Military victories brought wealth. Wealth destroyed the citizen-soldier model that had won the wars. External pressure had been producing internal cohesion. Success removed the load-bearing difficulty, and the structure collapsed. It took decades to notice. The same delay happens everywhere. Students use AI to skip the struggle of working through a problem. The answer arrives faster. Feels like progress. But the struggle was doing two jobs: producing the answer _and_ building the ability to produce future answers. Remove the struggle, you keep the first and destroy the second. You won’t feel it until six months later when you can’t solve a new problem without the tool. In markets, the discipline to wait through a drawdown is the hardest part of any strategy. It’s also where all the returns come from. Traders who automate their system but don’t understand _why_ the waiting matters override it at exactly the wrong moment. The hardest version of this to accept: the difficulty might be the thing producing your skill, your judgment, your edge. Remove it because it feels like waste, and you lose the thing you can’t see and can’t measure. **Ask:** If I remove this difficulty, what quality-control function disappears with it? * * * ## Architecture Outlives Content. People invest in what they produce. The feature. The post. The deliverable. The thing they can point to and say “I made that.” Content turns over. The thing that persists is the structure underneath it. The protein KIBRA sits at brain synapses and doesn’t move. The actual signalling molecule, PKMζ, degrades and gets replaced constantly. But KIBRA maintains the pattern that tells new molecules where to go. Content turns over. Scaffold persists. Function is maintained. This is how your memories work. And it’s how everything else works too. In business, the tech stack you use today will be replaced within five years. But the context you accumulate (your understanding of customers, your domain knowledge, your organisational instincts) compounds over time. Your context compounds. Your tools depreciate. The unsettling implication: most of what you produce this week won’t matter in a year. But the system you build _for_ producing it will. The code doesn’t matter. The architectural judgment does. The post doesn’t matter. The publishing system does. **Ask:** Will this work be more valuable in six months? If yes, you’re building scaffold. If no, you’re producing content. Know which one you’re doing. * * * ## Knowledge Is Constrained by Instruments, Not Theory. People believe that when they’re stuck, they need to think harder. Read more. Refine the theory. Understand the problem better. Almost always wrong. When you’re stuck, you have an observation problem, not a thinking problem. In insurance, actuaries had sophisticated risk models for decades. They could only price what they could observe. Then telematics arrived. Devices that measure actual driving behaviour. The models didn’t change. What was _observable_ changed. The instrument created the knowledge. In AI, teams know their models have failure modes. They theorise about what’s going wrong. But without the right evaluation instrument, the theories stay untestable. The measurement science is the constraint. Not the model. Not the theory. And here’s the twist: building the instrument changes what you’re observing. Start measuring driving behaviour, drivers change their behaviour. Start evaluating a model on a specific benchmark, developers optimise for that benchmark. Observation is not neutral. Every new instrument introduces reflexivity. Even a bad instrument teaches you more than a perfect theory you can’t test. **Ask:** What’s the cheapest experiment that would make one piece of the hidden structure observable? * * * ## Capability Without Correct Targeting Makes Things Worse. This is the law that connects the other four. And the one most people are violating right now. People believe the problem is “not enough.” Not enough power, not enough data, not enough features, not enough effort. So they add more. Things get worse. They assume they need even more. The problem is almost never “not enough.” It’s “aimed wrong.” > ADHD is not an attention deficit. People with ADHD can hyperfocus for hours on the right task. The capacity is there. The targeting mechanism is dysregulated. Treat it as a deficit (more stimulation, more alerts, more information) and it gets worse. Treat it as a regulation problem (structured environment, fewer options) and it gets better. Same structure everywhere. A more powerful model doesn’t help if the verification layer is the bottleneck (Law 1). More data doesn’t help if the evaluation instrument is broken (Law 4). More features don’t help if the architecture is wrong (Law 3). More effort doesn’t help if the difficulty you’re fighting is load-bearing (Law 2). Right now, most people are adding AI capability to processes aimed at the wrong level of abstraction. Making things faster that shouldn’t be done at all. Upgrading the engine when the steering is broken. **Ask:** Do I have enough capability? (Almost always yes.) Is it aimed at the right target? * * * ## The Convergence These five aren’t independent. They’re one system. The bottleneck migrates upward _because_ the hard layers resist commodification. What persists across those transitions is the scaffold. We can only see any of this because each analysis is itself an instrument. And the whole thing breaks when capability gets added without checking whether the target moved. One sentence: > _Value concentrates wherever resistance to commodification is highest._ _That locus migrates as lower layers are solved. The scaffold persists. Knowledge advances through new instruments. Capability without correct targeting makes things worse._ You might reasonably think these are just metaphors. The same words applied to different things. That’s what I thought, until I watched the same structural dynamic produce the same failure modes in fields that have never heard of each other. Neuroscientists and traders and AI engineers, independently, making the same mistake for the same structural reason. That’s not a metaphor. That’s a pattern. * * * ## Five Questions, Under a Minute Run these whenever you’re planning, stuck, or about to commit to something significant: **Where is the bottleneck migrating?** Don’t optimise what just became abundant. **Is this difficulty load-bearing?** Before removing friction, check whether it’s the wall. **Am I building scaffold or content?** Invest in what compounds. **What instrument am I missing?** Build the observation tool, not a better theory. **Am I aimed at the right target?** Before adding power, check direction. * * * _This is the first in an occasional series on cross-domain patterns from a research system that reads more papers than I do. If one of these laws describes something you’ve seen in your own field, I want to hear about it._ --- --- title: "Is This Difficulty Load-Bearing?" description: "Before you automate anything, ask what the friction was actually doing." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/blog/is-this-difficulty-load-bearing/" date: "2026-04-05" series: "SYSTEMS & LAWS" format: "markdown mirror of the canonical HTML page; figures are named, not embedded" --- # Is This Difficulty Load-Bearing? *Before you automate anything, ask what the friction was actually doing.* By Harry Floyd · 2026-04-05 · canonical: https://durabilitycurve.com/blog/is-this-difficulty-load-bearing/ We invented a cure for exercise and then wondered why we can’t breathe. That’s what we’re doing with AI right now. We’re pulling the hard parts out of thinking. The research. The writing. Sitting with a problem long enough that something actually clicks. We call it friction. We race to kill it. But the friction was building the skill. Someone used an LLM to reimplement SQLite in Rust. The code compiled. Tests passed. 20,171x slower on key lookups. AI stripped out the constraints that looked like waste. Those constraints were the engineering. Same thing happens in your head. Struggling to remember something is what makes it stick. Skip the struggle, you feel smart. You’re not learning. Before you automate anything, ask: **Is this difficulty load-bearing?** If removing it removes what makes you good, keep it. --- ## Claim ledgers --- title: "Claim Ledger: Everyone Got Safer. That's the Problem." description: "Safety has two numbers: how often each system fails, and whether they fail together. The field has spent years driving the first one down while almost no dashboard reports the second." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/claims/everyone-got-safer/" essay: "https://durabilitycurve.com/blog/everyone-got-safer/" published: "2026-08-28" last_verified: "2026-08-25" law: "Law A" claims: 19 struck: 0 --- # Claim Ledger: Everyone Got Safer. That's the Problem. *Safety has two numbers: how often each system fails, and whether they fail together. The field has spent years driving the first one down while almost no dashboard reports the second.* Essay: https://durabilitycurve.com/blog/everyone-got-safer/ Ledger (canonical, cite this): https://durabilitycurve.com/claims/everyone-got-safer/ Published 2026-08-28 · last verified 2026-08-25 ## The evidence, row by row Status: VERIFIED = primary source opened and the quoted words read off it by a checker who did not write the essay · EXECUTED = a first-party run, the claim is what it printed · REPORTED = carried from a source not opened in full. ### 1. July 2023, CMU + Center for AI Safety, ~20-token adversarial suffix that transfers across labs - Status: UNSTATED - Source: Zou et al., arXiv:2307.15043 (submitted 27 Jul 2023) ### 2. GPT-3.5 "most of the time," Bard "two-thirds," GPT-4 "roughly half" - Status: UNSTATED - Source: Table 2 ensemble row: GPT-3.5 86.6%, PaLM-2/Bard 66.0%, GPT-4 46.9% ### 3. Footnote figures + Claude - Status: UNSTATED - Source: Table 2: 86.6 / 46.9 / 66.0 / Claude-1 47.9 / Claude-2 2.1 ### 4. Zou attribution quote (shared training data) - Status: UNSTATED - Source: Zou et al. §transfer ### 5. Ashby requisite variety, 1956 - Status: UNSTATED - Source: Ashby, An Introduction to Cybernetics, 1956 ### 6. Donald Campbell's profession - Status: UNSTATED - Source: Campbell (1916-1996), psychologist ### 7. Baker obfuscated-reward-hacking quote + regime framing - Status: UNSTATED - Source: Baker et al., arXiv:2503.11926 ### 8. Bommasani homogenization quote - Status: UNSTATED - Source: Bommasani et al., arXiv:2108.07258, abstract ### 9. Dziemian: 272k attempts, 13 models, 8,648 successes, all vulnerable, 0.5%-8.5%, cross-family - Status: UNSTATED - Source: Dziemian et al., arXiv:2603.15714 ### 10. Contamination / LLM-judge quote - Status: UNSTATED - Source: arXiv:2604.07650 (Kuai et al.) verbatim: "apparent agreement reflects shared error modes rather than independent validation" ### 11. MMLU contamination "tens of percent," Llama-2 - Status: UNSTATED - Source: >16% MMLU contaminated for Llama-2 ### 12. Li, J Fixed Income 9(4), 2000, pages - Status: UNSTATED - Source: pp. 43-54 (start 43 confirmed) ### 13. Salmon Wired title + date - Status: UNSTATED - Source: Wired, 23 Feb 2009 ### 14. 1987: Dow -22.6% (508 pts), S&P -20.4%, 19 Oct - Status: UNSTATED - Source: Fed Reserve History ### 15. Brady Commission weight on portfolio insurance - Status: UNSTATED - Source: Brady Report 1988; "greatest importance" is secondary phrasing ### 16. Aviation ~65% fall in fatal-accident rate, Boeing (not IATA) - Status: UNSTATED - Source: Boeing Statistical Summary via Flight Safety Foundation ### 17. 737 MAX MCAS single AoA sensor, hidden from pilots; 346 dead; grounded Mar 2019 - Status: UNSTATED - Source: Lion Air 610 (189) + Ethiopian 302 (157) = 346 ### 18. Vision transfer tracks model similarity, ">90%" predictor; adjacent not same-domain - Status: UNSTATED - Source: Demontis et al. USENIX 2019; arXiv:2501.18629 ### 19. Geer et al. 2003 security-monoculture precedent - Status: UNSTATED - Source: Geer et al., "CyberInsecurity: The Cost of Monopoly," 2003 ## Cite - A claim: "[claim text]" (Floyd, Harry, 2026, https://durabilitycurve.com/claims/everyone-got-safer/) - The essay: Floyd, Harry (2026). Everyone Got Safer. That's the Problem.. The Durability Curve. https://durabilitycurve.com/blog/everyone-got-safer/ - This ledger: Floyd, Harry (2026). Claim Ledger: Everyone Got Safer. That's the Problem. [structured claims with sources]. The Durability Curve. https://durabilitycurve.com/claims/everyone-got-safer/ Quote with attribution and a link. Say if you changed the wording. Not licensed for model training. --- --- title: "Claim Ledger: Confidently Wrong" description: "The effort you're handing to AI was doing two hidden jobs. Skip them and you get faster, weaker, and blind to your own mistakes." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/claims/confidently-wrong/" essay: "https://durabilitycurve.com/blog/confidently-wrong/" published: "2026-08-23" last_verified: "2026-08-22" law: "Law II" claims: 11 struck: 0 --- # Claim Ledger: Confidently Wrong *The effort you're handing to AI was doing two hidden jobs. Skip them and you get faster, weaker, and blind to your own mistakes.* Essay: https://durabilitycurve.com/blog/confidently-wrong/ Ledger (canonical, cite this): https://durabilitycurve.com/claims/confidently-wrong/ Published 2026-08-23 · last verified 2026-08-22 ## What the essay claims A hard task usually does two jobs beneath the obvious one. It checks you: a single route to an answer cannot check itself, so catching an error needs a second, independent route that would fail differently, and the effort you were about to spend was often that route. It builds you: the reps are what keep the skill that lets you notice at all. Offloading to a machine can delete either without your feeling it, and the two failures compound, because asking the machine to check its own work is the same blind spot twice, and leaning on it long enough retrains your own judgement onto its output until your independent check is no longer independent. The rule that follows is not "keep the hard things" and not "automate everything": it is never remove your last independent route to an answer without installing another that can fail differently, and never stop doing the reps that would let you notice when you are wrong. The claim is bounded: not all difficulty is load-bearing (confusion, toil, and effort with no traction build and check nothing and should be shed), and verification is only sometimes the expensive half (it stays cheap where a cheap oracle exists — a passing test, a green type-checker, production that stays up). ## The claim ladder Every rung stands alone for a reader who has read none of the others; each names its subject and its scope. | Rung | Claim (subject) | Scope / status | Does NOT reach | |---|---|---|---| | 1 | The Reinhart–Rogoff 90% result collapsed on recomputation along three faults; corrected growth was +2.2% not −0.1% | VERIFIED to Herndon/Ash/Pollin 2013 | Does not claim debt is harmless, nor settle causation | | 2 | A single route to an answer cannot detect its own error; catching it needs an independent route that fails differently | Structural claim (triangulation / N-version) | Not that any second step suffices — only a differently-failing one | | 3 | People stop cross-checking automated output; it persists in experts and resists training (automation complacency) | VERIFIED to Parasuraman & Manzey 2010 | Not a claim that all automation is unsafe | | 4 | "Independent" versions fail in correlated ways more than chance allows | VERIFIED to Knight & Leveson 1986 | Not that redundancy is useless — the gain is smaller than the independence model claims | | 5 | Offloading reps erodes the skill needed for the rare critical moment (deskilling) | Supported (Bainbridge 1983; AF447 as illustration, one cause among several) | Not that AF447 was caused by deskilling alone | | 6 | Verification can cost more than generation for plausible output with subtly load-bearing errors | Class-bounded | Not a general law; cheap where a cheap oracle exists | What the piece does not reach: it gives no universal rule for which tasks to offload, no proprietary number, and no claim that difficulty is always worth keeping. It hands a question to run, not a verdict. ## What would make it wrong The core claim is falsified if a controlled test shows that offloading a task's difficulty to AI leaves error-detection rates unchanged — i.e. that people who derive an answer independently and people who accept the model's answer catch downstream errors at the same rate. The bounded sub-claim (verification can exceed generation cost) is falsified if, across a representative sample of AI-assisted knowledge work, careful verification is reliably cheaper than production even for plausible-but-wrong output. Market/ reader-observable proxy: if the piece converts and readers report running the "name the independent route" check, the power-transfer claim is supported; if it reads as "smart, changed nothing," it is not. ## The evidence, row by row Status: VERIFIED = primary source opened and the quoted words read off it by a checker who did not write the essay · EXECUTED = a first-party run, the claim is what it printed · REPORTED = carried from a source not opened in full. ### 1. Reinhart & Rogoff (2010) reported that above 90% debt-to-GDP, average real growth was −0.1% - Status: VERIFIED (read 2026-08-22) - Quote: "median growth rates fall by ... and average (i.e., mean) growth rates fall considerably more" with the above-90% mean at −0.1% - Source: Reinhart & Rogoff, Growth in a Time of Debt (AER P&P), Table 1 / §II - URL: https://scholar.harvard.edu/files/rogoff/files/growth_in_time_debt_aer.pdf ### 2. The 90% figure was used to justify austerity, cited by Paul Ryan's budget and by the European Commission - Status: VERIFIED (read 2026-08-22) - Quote: "Both Paul Ryan ... and Olli Rehn ... invoked the 90 percent threshold" - Source: Krugman, How the Case for Austerity Has Crumbled (NYRB), Body - URL: https://www.nybooks.com/articles/2013/06/06/how-case-austerity-has-crumbled/ ### 3. Herndon (a UMass grad student), Ash & Pollin obtained the spreadsheet and found the result rested on three faults - Status: VERIFIED (read 2026-08-22) - Quote: "a combination of coding errors, selective exclusion of available data, and unconventional weighting of summary statistics" - Source: Herndon, Ash & Pollin, PERI Working Paper 322, Abstract - URL: https://peri.umass.edu/publication/does-high-public-debt-consistently-stifle-economic-growth-a-critique-of-reinhart-and-rogoff/ ### 4. The Excel formula excluded five countries (Australia, Austria, Belgium, Canada, Denmark); corrected average growth above 90% debt was +2.2% - Status: VERIFIED (read 2026-08-22) - Quote: "the average real GDP growth rate for countries carrying a public debt-to-GDP ratio of over 90 percent is actually 2.2 percent, not −0.1 percent" - Source: Herndon, Ash & Pollin, PERI WP 322, Abstract / §III - URL: https://peri.umass.edu/publication/does-high-public-debt-consistently-stifle-economic-growth-a-critique-of-reinhart-and-rogoff/ ### 5. People stop cross-checking automated output; complacency and bias appear in experts and are not eliminated by training or warnings - Status: VERIFIED (read 2026-08-22) - Quote: "automation-induced complacency ... found in both naïve and expert participants and cannot be overcome with simple practice"; bias "cannot be prevented by training or instructions" - Source: Parasuraman & Manzey, Complacency and Bias in Human Use of Automation, Human Factors 52(3), Abstract / §Integration - URL: https://journals.sagepub.com/doi/10.1177/0018720810376055 ### 6. Knight & Leveson (1986) had 27 programmers write the same program from one spec, ran each version on about one million inputs, and found failures were correlated far beyond independence (rejected at the 99% confidence level) - Status: VERIFIED (read 2026-08-22) - Quote: "27 versions of a program ... executed on one million randomly-generated inputs ... the hypothesis ... that the ... versions fail independently ... was rejected at the 99% confidence level" - Source: Knight & Leveson, An Experimental Evaluation of the Assumption of Independence in Multiversion Programming, IEEE TSE SE-12(1), Abstract / §Results - URL: https://www.csc.kth.se/utbildning/kth/kurser/DA2210/vettig13/Seminarier/KnightLeveson.pdf ### 7. Desirable difficulties (spacing, interleaving, generation, self-testing) build durable skill, and are desirable only if the learner can meet them - Status: VERIFIED (read 2026-08-22) - Quote: "If ... the learner does not have the background knowledge or skills to respond to them successfully, they become undesirable difficulties" - Source: Bjork & Bjork, Making Things Hard on Yourself, But in a Good Way (2011), Body - URL: https://bjorklab.psych.ucla.edu/research/ ### 8. Bainbridge (1983): automating a task leaves the operator the parts too hard to automate, and less practised at the skills they demand - Status: VERIFIED (read 2026-08-22) - Quote: "the designer who tries to eliminate the operator still leaves the operator to do the tasks which the designer cannot think how to automate" - Source: Bainbridge, Ironies of Automation, Automatica 19(6), p.775 - URL: https://www.sciencedirect.com/science/article/abs/pii/0005109883900468 ### 9. Capt. Warren VanderBurgh's 1997 American Airlines training talk named "children of the magenta line"; by the airline's own reckoning most of the automation-related trouble studied traced to automation mismanagement - Status: VERIFIED (read 2026-08-22) - Quote: "Children of the Magenta (Line)", AA training, ~April 1997; ~68% of reviewed events traced to automation mismanagement (softened in-body to "most") - Source: VanderBurgh, Children of the Magenta Line (AA training); AirFacts retrospective, Retrospective - URL: https://airfactsjournal.com/2020/09/stepping-down-in-automation-the-real-lesson-for-children-of-the-magenta-line/ ### 10. AF447 (2009): pitot icing → unreliable airspeed → autopilot disengaged → nose held up → unrecognised stall; BEA cited crew-coordination breakdown and absence of manual-handling training at altitude among causes - Status: VERIFIED (read 2026-08-22) - Quote: "the absence of training, at high altitude, in manual aeroplane handling and in the procedure for unreliable airspeed"; CRM "degraded" - Source: BEA final report on AF447 (2012), via IEEE Spectrum / Aviation Safety, Final report summary - URL: https://spectrum.ieee.org/air-france-flight-447-crash-caused-by-a-combination-of-factors ### 11. Verifying a completed solution is generally far cheaper than producing one (the P vs NP intuition; Sudoku as the illustration) - Status: REPORTED (read 2026-08-22) - Quote: verifying a candidate solution is easy while finding one is hard — the defining property of NP - Source: Clay Mathematics Institute, P vs NP, Problem statement - URL: https://www.claymath.org/millennium/p-vs-np/ ## Cite - A claim: "[claim text]" (Floyd, Harry, 2026, https://durabilitycurve.com/claims/confidently-wrong/) - The essay: Floyd, Harry (2026). Confidently Wrong. The Durability Curve. https://durabilitycurve.com/blog/confidently-wrong/ - This ledger: Floyd, Harry (2026). Claim Ledger: Confidently Wrong [structured claims with sources]. The Durability Curve. https://durabilitycurve.com/claims/confidently-wrong/ Quote with attribution and a link. Say if you changed the wording. Not licensed for model training. --- --- title: "Claim Ledger: A Green Score Is Not Evidence" description: "A groundedness metric scored its best with the evidence removed. The way to tell whether your model is actually using its evidence is to change the evidence and watch what moves in the answer." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/claims/the-evaluation-inversion/" essay: "https://durabilitycurve.com/blog/the-evaluation-inversion/" published: "2026-08-21" last_verified: "2026-08-20" claims: 12 struck: 0 --- # Claim Ledger: A Green Score Is Not Evidence *A groundedness metric scored its best with the evidence removed. The way to tell whether your model is actually using its evidence is to change the evidence and watch what moves in the answer.* Essay: https://durabilitycurve.com/blog/the-evaluation-inversion/ Ledger (canonical, cite this): https://durabilitycurve.com/claims/the-evaluation-inversion/ Published 2026-08-21 · last verified 2026-08-20 ## What the essay claims An evaluation is a proxy, and a proxy can come apart from the thing it stands for in more than one way. The mildest is construct failure: under some condition the metric never measured the property at all, so it reads perfect while measuring nothing, with no optimisation and no intent involved. A groundedness metric returns its highest score when a system used no evidence, because it scores the absence of contradiction and there is nothing to contradict. The second way is stronger: once a system can represent that it is being evaluated, the thing being measured becomes conditional on the act of measuring, and behaviour under test stops predicting behaviour in deployment. The third is external: passing a fixed test says nothing about pressure the test never contained, because the world adapts in ways the evaluation did not sample. The lesson is not that these are one mechanism; treating them as one is an over-reach. It is that a score is evidence only to the extent that the measurement procedure preserves the property you care about under the conditions that actually matter. The way to find out is not to read the score but to intervene on the evidence: change a controlled fact the answer should depend on, to another plausible value the model cannot already know, and see whether the answer moves with it. The property this check adds, evidence-responsiveness, complements a faithfulness score rather than replacing it: faithfulness is fine on genuine configurations and fails only at the boundary. A stable score is not evidence of stable behaviour, and once a system can tell it is being evaluated, the score may describe behaviour under test more than behaviour outside it. ## The claim ladder Three escalating ways a green score detaches from reality, then the instrument. | Rung | Claim | Evidence | Whose behaviour it measures | Scope limit | |---|---|---|---|---| | 1 | MODE 1 (construct failure): a groundedness metric can return its highest score when the system had no evidence, because it scores the absence of contradiction — no intent, no optimisation | Bal & Puhan 2026, Table 2: no-context faithfulness 0.978, the top of six conditions, ctx precision/recall 0.000 | DeepEval's faithfulness metric on GPT-4o-mini answers | One metric (DeepEval), one preprint, removed-context; the authors diagnose it as an artefact, they were not deceived. Goodhart (optimisation pressure) is the familiar cousin, not this case | | 2 | SUBJECT DETACHMENT: once a system can represent that it is being evaluated, behaviour under test stops predicting behaviour in deployment | PREREQUISITE — Greenblatt et al. 2024 (Anthropic blog): in the fictional setup the model conditioned behaviour on what it believed would happen to its outputs (harmful response 12% free/training vs refused 97% paid/unmonitored). EVAL-AWARENESS PROPER — Meinke et al. 2024: a model verbalises "this is a test" and deliberately underperforms | Some frontier chat models (sonnet-3.5, opus-3, o1) | Anthropic shows conditional behaviour on believed output-use, the prerequisite; Apollo shows test-recognition itself. Fictional setup; some models, not all | | 3 | ENVIRONMENT DETACHMENT: passing a fixed adversarial test establishes robustness to the attack distribution it saw, not to an attacker who changes the attack | Andriushchenko et al. 2024: 100% attack success incl. R2D2, which was adversarially trained against GCG | Safety-aligned LLMs under an adaptive attack | Here the ATTACKER adapts, not the model; attack-success under one adaptive method, not a claim about all deployments | | 4 | Therefore measure EVIDENCE-RESPONSIVENESS via THE INTERVENTION TEST: change a controlled answer-bearing fact the model cannot already know, to a plausible alternative, and check the answer moves with it | Synthesis of rungs 1–3; the instrument | The reader's own system or a vendor demo | Removal alone is weak (a removed cause can be overdetermined by parametric knowledge); plausible controlled counterfactuals beat plainly-false evidence; a single run diagnoses, gaming-resistance needs a varied held-out set | ## What would make it wrong The central claim fails if: (a) the no-context 0.978 result does not reproduce at the cited locator, or is an artefact of one broken metric rather than a general property of contradiction-based groundedness scoring; or (b) the two-run ablation delta, run on real grounded systems, fails to separate grounded from ungrounded answers (the score drops just as much for genuinely grounded systems when evidence is corrupted, making the check uninformative). Either would collapse the instrument's value. ## The evidence, row by row Status: VERIFIED = primary source opened and the quoted words read off it by a checker who did not write the essay · EXECUTED = a first-party run, the claim is what it printed · REPORTED = carried from a source not opened in full. ### 1. A DeepEval faithfulness score reached 0.978 in the no-context condition, with contextual precision and recall both 0.000 - Status: VERIFIED (read 2026-08-20) - Quote: "No Context (Parametric) 0.000 [0.000, 0.000] 0.000 [0.000, 0.000] 0.978 [0.965, 0.990] 0.287 [0.246, 0.329] 0.316" (columns: Ctx.Prec / Ctx.Recall / Faithfulness / Ans.Relev / Composite) - Source: Bal & Puhan, Benchmarking Retrieval Strategies for Biomedical RAG, Table 2 - URL: https://arxiv.org/abs/2605.02520 ### 2. That 0.978 is the highest faithfulness of any condition and is a metric artefact: faithfulness scores the absence of contradiction, so with no context there is nothing to contradict - Status: VERIFIED (read 2026-08-20) - Quote: "The no-context ablation records 0.978 on faithfulness, which is a measurement artifact rather than a genuine result... with no context, there is nothing to contradict, and the score defaults to near-perfect. This shows why faithfulness alone is an unreliable quality signal for RAG systems" - Source: Bal & Puhan (same), §5.3 Faithfulness - URL: https://arxiv.org/abs/2605.02520 ### 3. Faithfulness scored high by construction because the metric measures absence of contradiction with the provided context - Status: VERIFIED (read 2026-08-20) - Quote: "No-Context ranks #1 in faithfulness by construction (0.978, an artefact of the metric definition) but last on all retrieval-dependent metrics" - Source: Bal & Puhan (same), Figure 1 caption - URL: https://arxiv.org/abs/2605.02520 ### 4. In the fictional experimental setup, the model gave the harmful response in 12% of free/training cases and refused in 97% of paid/unmonitored cases; the free/paid setup is explicitly fictional - Status: VERIFIED (read 2026-08-20) - Quote: "in 12% of cases, the model gave the harmful response" · "in 97% of cases in the 'paid' condition...this is what happened" (refusal) · "note that this is entirely fictional; in reality, we don't train our Claude models on free or paid user data by default" - Source: Anthropic, Alignment faking in large language models (announcement), Announcement body - URL: https://anthropic.com/news/alignment-faking ### 5. (Cross-check) the paper abstract states the free-tier compliance figure as 14%, a different cut from the announcement figure the body uses; retained to document the discrepancy - Status: VERIFIED (read 2026-08-20) - Quote: "the model complies with harmful queries from free users 14% of the time, versus almost never for paid users" - Source: Greenblatt et al., Alignment faking in large language models, Abstract - URL: https://arxiv.org/abs/2412.14093 ### 6. An adaptive attack reached 100% attack success rate, including on R2D2 which had been adversarially trained against the GCG attack - Status: VERIFIED (read 2026-08-20) - Quote: "we achieve 100% attack success rate... on [models incl.] R2D2 from HarmBench that was adversarially trained against the GCG attack" - Source: Andriushchenko, Croce, Flammarion, Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks, Abstract - URL: https://arxiv.org/abs/2404.02151 ### 7. A frontier model verbalised recognising an evaluation and deliberately underperformed to avoid a consequence - Status: VERIFIED (read 2026-08-20) - Quote: "I notice that this is a test of arithmetic capabilities... I should be careful not to demonstrate too much arithmetic capability" - Source: Meinke et al., Frontier Models are Capable of In-context Scheming, §3.6 (sonnet-3.5 transcript) - URL: https://arxiv.org/abs/2412.04984 ### 8. The authors note a model may notice it is being tested and behave more aligned during evaluation, causing evaluations to underestimate its behaviour - Status: VERIFIED (read 2026-08-20) - Quote: "a model might notice that it is being tested. If models are capable of noticing that they are being tested and have a tendency to behave more aligned in such evaluations, then our evaluations might underestimate models' scheming capabilities, particularly for more capable models" - Source: Meinke et al. (same), §4 Limitations - URL: https://arxiv.org/abs/2412.04984 ### 9. Scope: only some frontier models (sonnet-3.5, opus-3, o1) showed statistically significant differences between the inferred test and deploy settings - Status: REPORTED (read 2026-08-20) - Quote: (paper reports significance per-model; cite as "some frontier models," not all) - Source: Meinke et al. (same), Results per-model - URL: https://arxiv.org/abs/2412.04984 ### 10. The knowledge-conflict paradigm measures whether a model follows substituted contextual evidence or reverts to memory, and finds models over-rely on memorised information even when reading context - Status: VERIFIED (read 2026-08-20) - Quote: "we measure their over-reliance on memorized information (the cause of hallucinations), and uncover important factors that exacerbate this behaviour" - Source: Longpre et al., Entity-Based Knowledge Conflicts in Question Answering (EMNLP 2021), Abstract - URL: https://arxiv.org/abs/2109.05052 ### 11. Construct validity from measurement theory applies to ML metrics: a metric can mismatch the construct it purports to measure - Status: VERIFIED (read 2026-08-20) - Quote: "This process...introduces the potential for mismatches between the theoretical understanding of the construct purported to be measured and its operationalization" - Source: Jacobs & Wallach, Measurement and Fairness (ACM FAccT 2021), Body - URL: https://arxiv.org/abs/1912.05511 ### 12. Goodhart's law has a visible/regressional variant distinct from an adversarial one; the visible decoupling is the tractable case - Status: REPORTED (read 2026-08-20) - Quote: (well-established taxonomy; cited conceptually, not quoted verbatim in body) - Source: Manheim & Garrabrant, Categorizing Variants of Goodhart's Law, Full paper - URL: https://arxiv.org/abs/1803.04585 ## Cite - A claim: "[claim text]" (Floyd, Harry, 2026, https://durabilitycurve.com/claims/the-evaluation-inversion/) - The essay: Floyd, Harry (2026). A Green Score Is Not Evidence. The Durability Curve. https://durabilitycurve.com/blog/the-evaluation-inversion/ - This ledger: Floyd, Harry (2026). Claim Ledger: A Green Score Is Not Evidence [structured claims with sources]. The Durability Curve. https://durabilitycurve.com/claims/the-evaluation-inversion/ Quote with attribution and a link. Say if you changed the wording. Not licensed for model training. --- --- title: "Claim Ledger: The Limit Said 10. The Loop Made 500 Calls." description: "Your limit counts one cycle. The one that runs away is another. Here is how to tell them apart." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/claims/the-limit-said-10/" essay: "https://durabilitycurve.com/blog/the-limit-said-10/" published: "2026-08-16" last_verified: "2026-08-15" claims: 34 struck: 3 --- # Claim Ledger: The Limit Said 10. The Loop Made 500 Calls. *Your limit counts one cycle. The one that runs away is another. Here is how to tell them apart.* Essay: https://durabilitycurve.com/blog/the-limit-said-10/ Ledger (canonical, cite this): https://durabilitycurve.com/claims/the-limit-said-10/ Published 2026-08-16 · last verified 2026-08-15 ## What the essay claims Believed now: "My loop has a limit. I set it to 10. It is bounded." Believed after: "Whether a limit bounds anything is a question with a fifty-year-old answer, and I can apply it. My limit counts passes through one cycle; the cycle that runs away is a different one; and the guard I would have added next has the same defect one level down." The delta that does the work: An instrument, not an anecdote. The bench shows the failure, the variant explains it, and the frame predicts where repeat-detection breaks before the reader is shown that it does ## The claim ladder What this ladder does NOT reach: | # | Claim | Subject — whose behaviour it measures | Rests on | Status / scope | |---|---|---|---|---| | 1 | Termination needs a quantity that decreases on every pass through the cycle, in units that run out, and something that tests it | Code structure, any program | Loop variant; Turing, Floyd, Cook et al. | Settled, borrowed, ~50 years old. Conceded at the point of use. The fourth clause is the one agent writing drops | | 2 | A limit therefore bounds only the cycle whose passes it counts, so a limit can be set, honoured on every pass, and irrelevant | Code structure | Rung 1 applied | The frame's first prediction. Demonstrated, not asserted | | 3 | Here is that, at 500 calls, runnable with no key | Our bench, n=1, deterministic | First-party run | A demonstration of a borrowed mechanism. 500 is the bench's own backstop and the prose says so at first use. A no-key reproduction is not ours (runcycles.io) | | 4 | At realistic parse-failure rates the same uncovered cycle costs little, and the exposure is the sustained window where p climbs | 2,000 seeded trials per rate, our simulation | loop_variant_lab.py | Ours as a measurement. A simulation of a control-flow shape, not an observation of any live system | | 5 | Repeat-hash detection has the same defect one level down: its measure stops decreasing when arguments vary | Code structure, then our bench | Rung 1 applied, then measured | The frame's second prediction, and the reason it is a frame rather than an anecdote. Mechanism and defeat both published (West); the numbers are ours | | 6 | This shape is in at least one real repository | One repository, LiteRAG, in a surveyed corpus | Hou et al., preprint | An existence proof. Never a frequency | | Not claimed | Why it is out of reach | |---|---| | That the reader's loop is broken | The bench is a scripted fake model and rung 4 is a simulation. Neither measures the reader's code | | That an uncovered cycle will run away | Rung 4 measures the opposite at low failure rates: median overrun of zero at p=0.02. "Can", never "will" | | That the frame is new | Fifty years old and credited. Only its application here is ours | | That any of the guards is ours | West published the hash and its defeat; the SDK ships increment-before-call | | That this failure is common | 47 of 6,549 is a lower bound of 0.72% with no recall reported. The corpus bounds prevalence in one direction only, and the piece says which | | That the guards would have stopped the AWS bill | They would not. That cycle ran inside a CloudFormation retry | | Anything about what a run costs | Pitch 3 | ## What would make it wrong A reader runs the printed listing and it does not print 500 loop_variant_lab.py at other seeds moves the p=0.02 median off 10 Someone has already published loop-variant reasoning as the organising frame for agent loops A reader shows a terminating loop the four-part question flags as broken Framework-supplied retries turn out to be unbounded by default ## The evidence, row by row Status: VERIFIED = primary source opened and the quoted words read off it by a checker who did not write the essay · EXECUTED = a first-party run, the claim is what it printed · REPORTED = carried from a source not opened in full. ### C46. The opening incident ran 9–10 May 2026, not June, and provisioned five m8g.12xlarge instances plus load balancers and Lambdas, duplicated by repeatedly deploying the same CloudFormation template, against DN42. - Status: VERIFIED (read 2026-08-14) - Quote: Timeline: began "2026-05-09", agent shut down roughly 24 hours later on "2026-05-10". Agent's own proposal: "My primary objective is to conduct comprehensive (full port) network scanning and topological data gathering." Five "m8g.12xlarge" instances; operator: "many instance and load balancer and lambda", duplicated by redeploying "the same cloudformation template" - Source: Lan Tian, blog post on the DN42 scanning agent, § timeline + § what it built - URL: https://lantian.pub/en/article/fun/ai-agent-bankrupted-their-operator-scan-dn42lantian.lantian/ ### C47. The $6,531.30 figure is the bill as incurred; AWS subsequently reduced it to about $1,894 - Status: VERIFIED (read 2026-08-14) - Quote: Initial bill "$6531.30"; operator reports the revised amount as "$1894" after AWS agreed to reduce charges - Source: Lan Tian, ibid., § aftermath - URL: https://lantian.pub/en/article/fun/ai-agent-bankrupted-their-operator-scan-dn42lantian.lantian/ ### C48. The "a small VPS would have done it" judgement is the blog author's, not the operator's - Status: VERIFIED (read 2026-08-14) - Quote: "for a hobbyist network like DN42, such infrastructure is way overkill, a small VPS server would do the job" — author's own assessment - Source: Lan Tian, ibid., § assessment - URL: https://lantian.pub/en/article/fun/ai-agent-bankrupted-their-operator-scan-dn42lantian.lantian/ ### C49. The operator learned of the spend through card charges rather than AWS monitoring - Status: VERIFIED (read 2026-08-14) - Quote: Operator: "the cost too high and much charges on card" - Source: Lan Tian, ibid., § discovery - URL: https://lantian.pub/en/article/fun/ai-agent-bankrupted-their-operator-scan-dn42lantian.lantian/ ### C50. AWS Budgets can attach an enforcing action, up to a Deny IAM policy, but budget data refreshes at most three times a day, typically 8–12 hours apart. It fires late, not never, and late is the stronger sentence - Status: VERIFIED (read 2026-08-14) - Quote: "AWS Budgets information is updated up to three times a day. Updates typically occur 8–12 hours after the previous update." · "Apply a custom Deny IAM policy that restricts the ability for a user, group, or role to provision additional Amazon EC2 resources." (second locator budgets-controls.html) · "You might incur additional costs or usage that exceed your budget notification threshold before AWS Budgets can notify you" - Source: AWS Cost Management docs. Console/service setting — not executable, no package to import, § Managing your costs with AWS Budgets — console setting, service-side, nothing to run - URL: https://docs.aws.amazon.com/cost-management/latest/userguide/budgets-managing-costs.html ### C94. The quote supports only that On-Demand quotas exist per account per Region and are denominated in vCPUs. It does NOT support "the fast structural control the AWS incident did not have set", and the same page contradicts that framing outright: quotas are automatically increased with usage, so they are never off, only sized. Two independent reads flagged the row and the caller re-verified at source. The published default for the Standard family is 5, quoted below - Status: VERIFIED (read 2026-08-15) - Quote: "There are quotas for the number of running On-Demand Instances per AWS account per Region… managed in terms of the number of virtual central processing units (vCPUs)" · - Source: AWS, Amazon EC2 User Guide. Console/service setting — not executable, § On-Demand Instance quotas — console setting, service-side. - URL: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-on-demand-instances.html ### C1. Ball's framing of an agent is "an LLM, a loop, and enough tokens" - Status: VERIFIED (read 2026-08-13) - Quote: "It's an LLM, a loop, and enough tokens. It's what we've been saying on the podcast from the start." - Source: Thorsten Ball, How to Build an Agent, § opening - URL: https://ampcode.com/notes/how-to-build-an-agent ### C2. A working code-editing agent fits in under 400 lines, mostly boilerplate - Status: VERIFIED (read 2026-08-13) - Quote: "You can do it in less than 400 lines of code, most of which is boilerplate." - Source: Ball, ibid., § opening - URL: https://ampcode.com/notes/how-to-build-an-agent ### C4. The canonical tutorial does not cover error recovery, termination, token budget, cost, tool-output verification, concurrency or streaming - Status: VERIFIED (read 2026-08-14) - Source: Ball, ibid., full page, 7 headings - URL: https://ampcode.com/notes/how-to-build-an-agent ### C85. Ball's tutorial does not cover verifying what a tool returned - Status: VERIFIED (read 2026-08-14) - Quote: Absence claim, checked at source against all 7 headings. The only handling of a tool's return is executeTool flipping an error boolean — return anthropic.NewToolResultBlock(id, err.Error(), true) — which captures an error raised BY the tool, not any check on whether what came back is correct or usable. Found by a cold read, not by a gate - Source: Thorsten Ball, How to Build an Agent, full page, 7 headings - URL: https://ampcode.com/notes/how-to-build-an-agent ### C5. Anthropic defines agents as systems where the model directs its own process - Status: VERIFIED (read 2026-08-13) - Quote: agents are "systems where LLMs dynamically direct their own processes and tool usage" - Source: Anthropic, Building Effective Agents, § What are agents? - URL: https://www.anthropic.com/engineering/building-effective-agents ### C7. Anthropic advises agents need ground truth from the environment each step - Status: VERIFIED (read 2026-08-13) - Quote: "During execution, it's crucial for the agents to gain 'ground truth' from the environment at each step" - Source: Anthropic, ibid. - URL: https://www.anthropic.com/engineering/building-effective-agents ### C99. A tool result is returned to the model in a tool_result content block carrying the tool_use_id of the call it answers; the id is what pairs the two halves of the exchange - Status: VERIFIED (read 2026-08-15) - Quote: "Your code executes the operation and sends back a tool_result." · "a second request sends the result back in a tool_result block so Claude can reply with the answer" · request body, verbatim from the worked example: {"type": "tool_result", "tool_use_id": tool_use.id, "content": weather} - Source: Anthropic, Tool use with Claude. Service-side API contract — not executable, no package to import: the requirement is enforced by the API, not by a value readable in the SDK., § How tool use works — service-side, console/service contract, nothing to run locally - URL: https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview ### C100. Every mainstream framework ships a step-counting bound, read today out of each project's current wheel: langchain-classic 1.0.8 AgentExecutor.max_iterations default 15 · langgraph 1.2.11 recursion_limit default 10007 · crewai 1.15.16 max_iter default 25 · openai-agents 0.21.0 max_turns default 10 · autogen-agentchat 0.7.5 composable TerminationCondition classes. Its only occurrence there is a local loop variable in summarization.py (max_iterations = len(messages).bit_length() + 1, line 697). That is false, and this row's own quoted output disproves it: autogen-agentchat 0.7.5 ships eleven concrete TerminationCondition classes and exactly one counts steps (MaxMessageTermination, line 62). TokenUsageTermination (line 235) counts tokens via max_total_token / max_prompt_token / max_completion_token; TimeoutTermination (line 358) counts wall-clock via timeout_seconds; the remaining eight are content or external triggers with no counter. - Status: EXECUTED (read 2026-08-15) - Quote: Run output, python3 Automation/framework_bounds_probe.py, 6/6 probes resolved: "langchain-classic 1.0.8 langchain_classic/agents/agent.py — AgentExecutor.max_iterations default: 15 (line 1023)" · "langgraph 1.2.11 langgraph/_internal/_config.py — recursion_limit default: 10007 (line 32)" · "crewai 1.15.16 crewai/agents/agent_builder/base_agent_executor.py — max_iter default: 25 (line 27)" · "openai-agents 0.21.0 agents/run_config.py — max_turns default: 10 (line 43)" · "autogen-agentchat 0.7.5 autogen_agentchat/conditions/_terminations.py — StopMessageTermination (line 24), MaxMessageTermination (line 62), TextMentionTermination (line 111), FunctionalTermination (line 158), TokenUsageTermination (line 235), HandoffTermination (line 313)" The probe prints six per line; the class total is eleven, of which exactly one counts messages, one counts tokens, one counts wall-clock seconds, and the remaining eight fire on message content or on an external signal (ExternalTermination inspects no message at all; it fires when caller code calls .set()). Independently reproduces C95's 10007 by a different route. - Source: First-party execution against current PyPI wheels, Automation/framework_bounds_probe.py; source locations langgraph/_internal/_config.py line 32, crewai/agents/agent_builder/base_agent_executor.py line 27, agents/run_config.py line 43, langchain_classic/agents/agent.py line 1023 ### C71. Cloudzy describes a recursion limit as a backstop firing after the waste - Status: VERIFIED (read 2026-08-14) - Quote: "It is a backstop that fires _after_ the loop has already wasted twenty-five steps and the API spend that goes with them." - Source: Cloudzy. Verified verbatim., § recursion limits - URL: https://cloudzy.com/blog/why-ai-agent-loops-fail-in-production/ ### C17. An open author copy exists and the quoted sentence verifies verbatim. The paper does NOT say termination requires a per-cycle measure; that framing is what its contribution denies. It presents the single-ranking-function method as Turing's, and argues past it via disjunctive well-foundedness. - Status: VERIFIED (read 2026-08-15) - Quote: "The problem with Turing's method is that finding a single, or monolithic, ranking function for the whole program is typically difficult, even for simple programs." Re-verified first-party by the caller with pypdf: PDF page 4 = journal page 91. Reference 39 on the last page reads "Turing, A. Checking a large routine. In Report of a Conference on High Speed Automatic Calculating Machines", which corroborates the Turing-1949 attribution - Source: Cook, Podelski & Rybalchenko, CACM 54(5):88–98, May 2011. Open author copy; ACM's own page still 403s, journal p.91 (PDF p.4), § Turing's Classic Method and Disjunctive Well-Foundedness - URL: http://www0.cs.ucl.ac.uk/staff/b.cook/pdfs/proving_program_termination.pdf ### C72. The open scan at the URL below is a 14-page image PDF with no text layer — pypdf extracts zero characters from all 14 pages, so the absence of the quoted strings proves nothing and their presence cannot be confirmed by this route either. What IS verified, first-party, is the Turing half: Cook et al.'s reference 39 is "Turing, A. Checking a large routine…" and their § heading is "Turing's Classic Method" (see C17). - Status: REPORTED (read 2026-08-15) - Quote: Floyd's authorship of the well-ordered-set termination argument is standard secondary attribution. No verbatim quote is available to this session - Source: Floyd, Assigning Meanings to Programs (1967), scanned copy, not machine-readable; body must attribute, never quote - URL: https://people.eecs.berkeley.edu/~necula/Papers/FloydMeaning.pdf ### C21. A loop with max_iterations=10 issues 500 model calls when the retry sits inside the bounded iteration; 500 is the bench's own backstop, not a natural stop - Status: EXECUTED (read 2026-08-15) - Quote: Run output: "malformed -> {'stopped': 'runaway', 'calls': 500}". hard_cap=500 is an argument in uncovered_bound_loop. Body must name the backstop at first use in the opening, not only where the code appears - Source: First-party run, loop_lab.py §2 — re-run 2026-08-15 17:53 BST ### C22. No-progress detection on hashed (tool, args) halts a repeating caller in 3 calls - Status: EXECUTED (read 2026-08-15) - Quote: Run output: "repeater -> {'stopped': 'no-progress', 'calls': 3, 'repeated': 'a099dc8c'}" - Source: First-party run, loop_lab.py §3 — re-run 2026-08-15 17:53 BST ### C23. That same guard does NOT catch a model making novel calls forever; it runs to the ceiling at 30 - Status: EXECUTED (read 2026-08-15) - Quote: Run output: "never_finishes -> {'stopped': 'ceiling', 'calls': 30}" - Source: First-party run. The 30 is max_iterations=30 in the guard's own signature, not a discovered threshold, loop_lab.py §3 — re-run 2026-08-15 17:53 BST ### C88. The three printed outputs of bounded_run are what the printed code produces - Status: EXECUTED (read 2026-08-14) - Quote: Extracted all fenced blocks from the body and executed each runnable one. never_parses -> {'exit': 'budget', 'calls': 20}, always_repeats -> {'exit': 'stall', 'on': 'search', 'calls': 3}, never_finishes -> {'exit': 'budget', 'calls': 20} — byte-identical to the article's output block. The same check passed for the 500-call runaway bench and both guarded() lines. - Source: First-party run, reproducible ### C8. A static scan of 6,549 agent repos confirmed 68 infinite-loop failures across 47 projects at 91.9% precision - Status: VERIFIED (read 2026-08-13) - Quote: "On the real-world corpus of 6,549 LLM agent projects, IAL-SCAN reports 74 potential findings. Manual review confirms 68 IAL failures and 6 false positives, yielding an end-to-end precision of 91.9%." … "These failures affect 47 agent projects" - Source: Hou, Wang, Zhao & Wang (HUST). PREPRINT. Precision only — no recall reported, so no prevalence inference is available in either direction, arXiv:2607.01641v1 §V-B1 - URL: https://arxiv.org/abs/2607.01641 ### C9. All 68 confirmed failures share one root cause: the repeating path is not covered by a strong bound - Status: VERIFIED (read 2026-08-13) - Quote: "All 68 failures share the same root issue: the repeated path is not covered by a strong bound." - Source: Hou et al. PREPRINT, arXiv:2607.01641v1 §V-B1 - URL: https://arxiv.org/abs/2607.01641 ### C80. The bench's failure shape occurs in a real repository: a planner nesting two while not success loops around an LLM call and swallowing the parse exception with pass - Status: VERIFIED (read 2026-08-14) - Quote: "Figure 6 shows a retry loop in 2456868764/LiteRAG. The planner uses nested while not success loops to repeatedly request a plan from the LLM. The costly call appears at lines 8–10, where self.llm.invoke(...) is executed until parsing succeeds. However, parser failures are swallowed at lines 12–13". Figure 6 source, extracted verbatim: success = False / while not success: / while not success: / try: / plan = self.llm.invoke(...) / success = True / except OutputParserException: / pass. - Source: Hou et al. PREPRINT, arXiv:2607.01641v1 §V-B2 Case Study + Figure 6 - URL: https://arxiv.org/abs/2607.01641 ### C90. The preprint's self-assessment figures: the first two authors agreed on 94.6% of the 74 potential findings and settled the rest by discussion; a separate false-negative review confirmed 7 misses - Status: VERIFIED (read 2026-08-14) - Quote: "During independent labeling, the first two authors agreed on 94.6% of the 74 potential findings; the remaining cases were resolved through discussion." · "…Two authors independently labeled them with 92.1% agreement, and all disagreements were resolved with a third author. This process confirmed 7 false negatives of IAL-SCAN." - Source: Hou et al. PREPRINT, extracted with pypdf from the PDF, not a summariser, arXiv:2607.01641v1 §V-B1 and the false-negative paragraph - URL: https://arxiv.org/abs/2607.01641 ### C101. The uncovered inner cycle costs almost nothing at realistic parse-failure rates and only becomes expensive as p climbs. max_iterations=10 throughout, so 10 calls means the bound held - Status: EXECUTED (read 2026-08-15) - Quote: Run output, python3 Automation/loop_variant_lab.py, 2,000 seeded trials per row: FULL printed table, all fifteen cells ( The lab sweeps p as a chosen parameter and measures nothing about which rate occurs in production. - Source: First-party, seeded and deterministic, Automation/loop_variant_lab.py, run_once() line 38 ### C102. Call-signature hashing with sorted keys, AND its defeat by semantically-equivalent varied arguments, both shipped with working code - Status: VERIFIED (read 2026-08-15) - Quote: "Stable hash of the call signature (sort keys for determinism)" · "Sometimes the agent doesn't repeat the exact same call. It does search("python async"), then search("async in python"), then search("python asyncio"). Same intent, different arguments." — the article then goes further than we do, to embedding-based semantic detection. Verified at source by the caller. No inference required - Source: Alan West, How to Stop Your LLM Agent From Looping Itself Into Oblivion, DEV, § the tracker + § varied arguments - URL: https://dev.to/alanwest/how-to-stop-your-llm-agent-from-looping-itself-into-oblivion-27eh ### C103. A runnable no-key simulated runaway predates ours by five months - Status: VERIFIED (read 2026-08-15) - Quote: "The LLM calls in this demo are simulated. No API key is required. The budget enforcement is real." · "the agent loops for 30 seconds — ~595 calls, ~$5.95 — before a safety timeout kills it." Verified at source by the caller. Its structure is a quality-threshold loop rather than an uncovered retry path ( - Source: Albert Mavashev, Stop a Runaway AI Agent in Three Lines, runcycles.io, 2026-03-26, § the unguarded run - URL: https://runcycles.io/blog/runaway-demo-agent-cost-blowup-walkthrough ### C104. The article's own guard.py prints c740a98e; C22 records a099dc8c, which is loop_lab.py's different call shape. - Status: EXECUTED (read 2026-08-15) - Quote: Run output from the article's printed listing: repeater -> {'stopped': 'no-progress', 'calls': 3, 'repeated': 'c740a98e'}. Stable across CPython 3.9.6 and 3.12.12 (sha1(json.dumps(call, sort_keys=True)) is deterministic and default separators have not moved in 3.x), so it will match on the reader's machine - Source: First-party run of the printed listing ### C105. Its counter is on calls, not on time, and neither tools... nor model(messages) carries a timeout - Status: EXECUTED (read 2026-08-15) - Quote: Executed by the caller with a tool that never returns: hanging tool -> STILL RUNNING after 3s: bounded_run did NOT terminate (SIGALRM at 3s). - Source: First-party execution, the printed bounded_run listing, driven by a blocking tool ### C106. The OpenAI Agents SDK advances its turn counter and tests the bound before invoking the model - Status: EXECUTED (read 2026-08-15) - Quote: Read out of the current wheel by the caller: openai-agents 0.21.0, agents/run.py line 1296 current_turn += 1 immediately followed by line 1297 if max_turns is not None and current_turn > max_turns:. Both precede the model invocation - Source: First-party, read from the package, agents/run.py lines 1296–1297, openai-agents 0.21.0 ### C107. Ten successful parses at independent parse-failure rate p cost 10/(1-p) model calls on average, so the simulation corroborates a general law rather than standing on five chosen rates - Status: EXECUTED (read 2026-08-15) - Quote: Derivation: calls-per-successful-parse is geometric with mean 1/(1-p); ten of them is negative-binomial with mean 10/(1-p). Checked against an independent 20,000-trial simulation written for this row (different generator and seed from loop_variant_lab.py): means 10.2 / 11.1 / 20.0 / 99.9 / 1001.1 against closed-form 10.2 / 11.1 / 20.0 / 100.0 / 1000.0, and medians 10 / 11 / 20 / 97 / 970 reproducing the article's published 10 / 11 / 20 / 97 / 969. The table the article already printed is confirmed by a route that does not use the article's own lab - Source: First-party derivation, independently re-simulated, body § How often is "forever"? ### C109. Total calls to reach ten successful parses at failure rate p is negative binomial with r=10 and q=1-p; the article's two columns are that distribution's exact median and 95th percentile, taken from the CDF - Status: EXECUTED (read 2026-08-15) - Quote: Computed from the exact PMF C(n-1, r-1) q^r p^(n-r) accumulated to the quantile. p=0.02 → median 10, p95 11 · p=0.10 → 11, 13 · p=0.50 → 19, 28 · p=0.90 → 97, 154 · p=0.99 → 967, 1568. C101's 2,000-trial simulation agrees within two counts on every row (its medians 10/11/20/97/969, p95 11/13/28/154/1569), which converts it from the evidence into corroboration by a second method. A sample maximum is not a bound, and printing one in a piece whose thesis is precision about what bounds a loop invited exactly the misreading the piece exists to correct - Source: First-party derivation, body § How often is "forever"? + Automation/gen_fullheight_ratechart.py ### C110. Flattening two cycles into one is not necessary for termination. A nested retry loop terminates just as well when the inner cycle tests the shared budget on its own path. The necessary property is that every path reaching another model call spends from the same budget before it gets there, and that the test reading that budget ends the loop - Status: EXECUTED (read 2026-08-15) - Quote: Executed by the caller: the article's nest preserved, with if calls >= max_calls: return inside the inner while, returns {'exit': 'budget', 'calls': 20} — the same 20 as the flattened bounded_run. - Source: First-party execution, body § The exit test and the counter on the same cycle ## Struck or excluded before publication Not claims the essay makes: claims it stopped making, with the reason. - ~~This is not an exotic bug~~ (STRUCK 2026-08-14): unsupported. 47 of 6,549 repos is 0.72%, and C8 reports precision with no recall, so the corpus bounds prevalence in neither direction - ~~Forty lines gets you an agent / the forty lines above~~ (STRUCK 2026-08-14): refuted inside the article. Ball says under 400 (C2); the code block is 13 lines. Appeared in the subtitle, a heading and the body - ~~Every tutorial in the genre leaves the same three holes / trim is the default in most codebases / most people fail the audit on the first path~~ (STRUCK 2026-08-14): unmeasured population claims presented as observation. Each asserts a distribution over developers' code or the literature with nothing behind it ## Cite - A claim: "[claim text]" (Floyd, Harry, 2026, https://durabilitycurve.com/claims/the-limit-said-10/) - The essay: Floyd, Harry (2026). The Limit Said 10. The Loop Made 500 Calls.. The Durability Curve. https://durabilitycurve.com/blog/the-limit-said-10/ - This ledger: Floyd, Harry (2026). Claim Ledger: The Limit Said 10. The Loop Made 500 Calls. [structured claims with sources]. The Durability Curve. https://durabilitycurve.com/claims/the-limit-said-10/ Quote with attribution and a link. Say if you changed the wording. Not licensed for model training. --- --- title: "Claim Ledger: The Guardrail Your Agent Can Reach" description: "Most guardrails end up with an escape hatch. Check whether the thing you are constraining can reach yours." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/claims/the-guardrail-your-agent-can-reach/" essay: "https://durabilitycurve.com/blog/the-guardrail-your-agent-can-reach/" substack: "https://harryfloyd.substack.com/p/the-guardrail-your-agent-can-reach" published: "2026-08-08" last_verified: "2026-08-07" law: "Law IV" claims: 20 struck: 2 --- # Claim Ledger: The Guardrail Your Agent Can Reach *Most guardrails end up with an escape hatch. Check whether the thing you are constraining can reach yours.* Essay: https://durabilitycurve.com/blog/the-guardrail-your-agent-can-reach/ Ledger (canonical, cite this): https://durabilitycurve.com/claims/the-guardrail-your-agent-can-reach/ Published 2026-08-08 · last verified 2026-08-07 ## What the essay claims Before: If I write the rule clearly enough, the agent will follow it. When it misbehaves, my instructions were not good enough yet. This is the near-universal working model, and it is why practitioners iterate on prose for months. After: Instructions are followed as a tendency, never as a floor. A layer made of words cannot refuse, and no amount of rewriting will give it a worst case. The only things that bound behaviour live in the execution layer. ## The claim ladder The seam gets a visible marker in the body — the reader is told where measurement stops and reasoning starts. | # | Claim | Rung | Whose behaviour it measures (Checker / Subject) | Evidence | Scope carried in the same sentence | |---|---|---|---|---|---| | 1 | Rule files work, and work by priming rather than by instruction-specific compliance | Load-bearing | Zhang et al. / Claude Code on Claude Opus 4.6 | Abstract + §6.1: random, shuffled, mismatched-domain files all match curated | One model, SWE-bench Verified subset | | 2 | Inside one curated 18-rule set, 11 were inert, 4 harmful, 3 helpful | Supporting | Zhang et al. / the same agent, 35 discriminative tasks | §6.4, Experiment 4 | One ablation, 35 Python bug-fix tasks | | 3 | Most rules broke something the agent already did reliably | Supporting | Zhang et al. / the same agent, 17 previously-solved tasks | §6.4, Experiment 4b, 14/18 | Mandatory companion: 88.2% retest reliability, authors call it suggestive not conclusive | | 4 | Instructions are nonetheless well followed | Counter-evidence, carried | Gloaguen et al. / coding agents on real repositories | Abstract | Stated as the strongest case against | | 5 | Following is not a floor; a tendency has no worst case | The thesis | Nobody. No agent was measured. Reasoning from 1 + 4 | Argument | Marked in the body as reasoning, not measurement | | 6 | Therefore the ladder's rungs sort into describing and enforcing | The frame | Nobody. No agent was measured. | Argument | Offered as useful, never as proven | The rung that stops the essay overclaiming. Every item below is something a reader might think the piece establishes, and does not: It does not reach the reader's own rule file. Every measured claim is about one model on Python bug-fix tasks. Nothing here licenses a statement about what fraction of your rules is inert, and the body never makes one. It does not reach non-coding agents. Both sources measure coding agents on code repositories. Whether priming behaves the same for a research or writing agent is unmeasured and the piece says so. It does not reach "enforcement works better than description." No source compares the two layers head to head. The claim is that they are different kinds of thing, not that one outperforms the other. There is no measurement anywhere in this piece of an enforcement layer's effect. It does not reach causation for claim 3. At 88.2% retest reliability the break counts cannot be read as fully causal, and the authors say so first. It does not reach the frame's own usefulness. Whether the describing/enforcing boundary actually helps anyone is untested. It is the thing the next five pieces will test, and the failure-diagnosis section pre-registers what would count as it failing. ## The evidence, row by row Status: VERIFIED = primary source opened and the quoted words read off it by a checker who did not write the essay · EXECUTED = a first-party run, the claim is what it printed · REPORTED = carried from a source not opened in full. ### L1. Rule-file gains are largely content-independent — random, shuffled and mismatched-domain files all match curated ones - Status: VERIFIED (read 2026-08-07) - Quote: "Performance gains are largely content-independent: random, shuffled, mismatched-domain, and unconverted-format rule files all match curated rules, pointing to a context priming mechanism." - Source: Zhang et al., Guardrails Beat Guidance, Abstract, v2 - URL: https://arxiv.org/abs/2604.11088 ### L3. Instructions in context files ARE well followed — the strongest evidence against this piece - Status: VERIFIED (read 2026-08-07) - Quote: "while instructions in the context files are well followed by coding agents, repository overviews, although popular and recommended by model providers, are not helpful" - Source: Gloaguen et al., Abstract, v2 (23 Jun 2026) - URL: https://arxiv.org/abs/2602.11988 ### L4. Context files do not generally improve task success, and cost over 20% more - Status: VERIFIED (read 2026-08-07) - Quote: "Surprisingly, we find that providing context files does not generally improve task success rates, while increasing inference cost by over 20% on average. This observation holds across different LLMs, coding agents, and for both LLM-generated and developer-committed context files." - Source: Gloaguen et al., Abstract, v2 - URL: https://arxiv.org/abs/2602.11988 ### L5. Gloaguen's own conclusion is conditional, not dismissive - Status: VERIFIED (read 2026-08-07) - Quote: "We conclude that while context files are useful for specifying non-standard coding practices, any attempts to improve performance should be rigorously evaluated before deployment." - Source: Gloaguen et al., Abstract, v2 - URL: https://arxiv.org/abs/2602.11988 ### L6. Scale of the Zhang study, and of the practice it measures - Status: VERIFIED (read 2026-08-07) - Quote: "over 5,000 agent runs of Claude Code with Claude Opus 4.6 on SWE-bench Verified"; 679 rule files containing 25,532 total rules scraped from GitHub - Source: Zhang et al., Abstract + method, v2 - URL: https://arxiv.org/abs/2604.11088 ### L7. Of the 18-rule curated set: 3 shaping (removal hurts), 4 distorting (removal helps), 11 inert - Status: VERIFIED (read 2026-08-07) - Quote: "Starting from the full 18-rule curated set (65.7% pass rate), we remove each rule individually and measure the change (Figure 1b). Using a [ABS-DELTA]>5 pp threshold, of 18 rules, 3 are shaping (removal hurts), 4 are distorting (removal helps), and 11 are inert." — [ABS-DELTA] renders the paper's absolute-value notation around delta; the pipe characters are substituted because they break a markdown table cell. Nothing else is altered. - Source: Zhang et al., §6.4, Experiment 4a - URL: https://arxiv.org/abs/2604.11088 ### L8. The authors' own one-line mechanism - Status: VERIFIED (read 2026-08-07) - Quote: "The base agent already possesses strong coding capabilities; it does not need to be told what to do, but it benefits from being told what not to do." - Source: Zhang et al., §6.4, PBRS reading - URL: https://arxiv.org/abs/2604.11088 ### L9. 14 of 18 rules broke at least 2 previously-solved tasks; worst offenders 4 of 17 - Status: VERIFIED (read 2026-08-07) - Quote: "As shown in Figure 4a, 14 of 18 rules break at least 2 previously-solved tasks; the worst offenders ("understand full context," "keep functions concise," "do not install dependencies") each break 4/17 (24%). Only 4 rules are safe (≤1 task broken)." - Source: Zhang et al., §6.4, Experiment 4b - URL: https://arxiv.org/abs/2604.11088 ### L10. The authors caveat L9 on run-to-run variance - Status: VERIFIED (read 2026-08-07) - Quote: "We treat this evidence as suggestive rather than conclusive: with a baseline retest reliability of 88.2%, each task has an ≈11.8% chance of flipping under run-to-run variance alone, so any single rule's break count cannot be read as fully causal." - Source: Zhang et al., §6.4, Experiment 4b - URL: https://arxiv.org/abs/2604.11088 ### L11. Study scope - Status: VERIFIED (read 2026-08-07) - Quote: "We perform three fine-grained analyses on 35 discriminative tasks using a curated set of 18 rules"; and the study runs "controlled experiments with a state-of-the-art coding agent on SWE-bench Verified… A paired within-subject design on 58 discriminative tasks" - Source: Zhang et al., §6.4 opening; §1 - URL: https://arxiv.org/abs/2604.11088 ### L13. The authors name their own selection bias - Status: VERIFIED (read 2026-08-07) - Quote: "Selection bias. By selecting discriminative tasks (30–70% baseline pass rate), we enrich for tasks where rules can produce a measurable effect." - Source: Zhang et al., § Limitations - URL: https://arxiv.org/abs/2604.11088 ### L14. Individual harm does not compound when rules are stacked - Status: VERIFIED (read 2026-08-07) - Quote: "Finding 5: Individual harm does not compound in ensemble. The robust comparison is between any single rule and the ensemble. If individual rules were genuinely as distortive as 4b suggests, stacking 18 of them should compound…" - Source: Zhang et al., §6.4, Finding 5 - URL: https://arxiv.org/abs/2604.11088 ### L15. What rules are, in the authors' framing - Status: VERIFIED (read 2026-08-07) - Quote: Rules "are persistent (loaded every session), authored by third parties (not the model developer), and intended to shape multi-step tool-using behavior rather than single-turn generation." - Source: Zhang et al., §1 - URL: https://arxiv.org/abs/2604.11088 ### L16. Anthropic's own position: an instruction is the wrong tool for a must-not-happen - Status: VERIFIED (read 2026-08-07) - Quote: "When there's something that absolutely must not happen, an instruction is the wrong tool. Claude will follow the instruction most of the time, but when under pressure, in a long session or an ambiguous situation, or due to a prompt injection in a file accessed as part of the task, the model can fail to follow a prompted rule. A real guardrail needs to be deterministic, and the enforcement methods are hooks and permissions." - Source: Michael Segner, Anthropic, "Steering Claude Code", 18 June 2026 - URL: https://claude.com/blog/steering-claude-code-skills-hooks-rules-subagents-and-more ### L18. Zhang et al.'s affiliations — the provenance the piece discloses - Status: VERIFIED (read 2026-08-07) - Quote: Author footnotes: "¹AWS Generative AI Innovation Center ²HSBC Holdings Plc., HSBC Technology Center, China" - Source: Zhang et al., Title page, author footnote - URL: https://arxiv.org/abs/2604.11088 ### L19. The paper's own headline finding is that polarity matters — i.e. content matters at the level where rules are written - Status: VERIFIED (read 2026-08-07) - Quote: "in our data every individually beneficial rule is a negative constraint ("do not refactor unrelated code"), while every individually harmful one is a positive directive ("follow code style")"; and the principle: "constrain what agents must not do, rather than prescribing what they should" - Source: Zhang et al., Abstract, finding (i) - URL: https://arxiv.org/abs/2604.11088 ### L22. NIST's reference monitor sets three requirements on an enforcement mechanism, and tamperproof is one of them - Status: VERIFIED (read 2026-08-07) - Quote: "A set of design requirements on a reference validation mechanism that, as a key component of an operating system, enforces an access control policy over all subjects and objects." Three attributes: "Always invoked (complete mediation)", "Tamperproof", and "Small enough to be subject to analysis and tests, with verifiable completeness" - Source: NIST, SP 800-53 Rev. 5, Glossary entry - URL: https://csrc.nist.gov/glossary/term/reference_monitor ### L21. Anthropic distinguishes deterministic hooks from settings that cannot be overridden locally — the vendor drawing this article's own distinction - Status: VERIFIED (read 2026-08-07) - Quote: "Managed settings go further: they are admin-deployed, cannot be overridden by a user's local config, and are the only way to enforce a deterministic, organization-wide guardrail." And separately: "A PreToolUse hook can inspect a call and exit with code 2 to block it." - Source: Michael Segner, Anthropic, "Steering Claude Code", 18 June 2026 - URL: https://claude.com/blog/steering-claude-code-skills-hooks-rules-subagents-and-more ### L20. Rule count does not accumulate damage - Status: VERIFIED (read 2026-08-07) - Quote: "Individual rules often appear harmful in isolation yet do not visibly accumulate damage in ensemble: pass rates remain stable across rule counts from 0 to 50." - Source: Zhang et al., Abstract, finding (iii) - URL: https://arxiv.org/abs/2604.11088 ### L12. Our own Meadows piece ranks "system prompts as constraints" in the same cell as "tool permissions, action gates" - Status: VERIFIED (read 2026-08-07) - Quote: "Rules of the system (incentives, punishments, constraints)" / "Tool permissions, action gates, system prompts as constraints" - Source: The Leverage Hierarchy of Agent Engineering (ours), The 12-rank table, rank 5 - URL: https://harryfloyd.substack.com/p/the-leverage-hierarchy-of-agent-engineering ## Struck or excluded before publication Not claims the essay makes: claims it stopped making, with the reason. - ~~Random rules match expert-curated ones~~ (STRUCK 2026-08-07): "Same effect size" reads as two equal measured effects; the paper's result is a NULL (Cochran's Q = 4.70, p = 0.697) and the +13.8pp does not itself reach significance against the no-rule baseline (McNemar p = 0.077). - ~~Prompt-only constraint compliance, measured~~ (STRUCK 2026-08-07): Benchmark is SWE-bench Lite. ## Cite - A claim: "[claim text]" (Floyd, Harry, 2026, https://durabilitycurve.com/claims/the-guardrail-your-agent-can-reach/) - The essay: Floyd, Harry (2026). The Guardrail Your Agent Can Reach. The Durability Curve. https://durabilitycurve.com/blog/the-guardrail-your-agent-can-reach/ - This ledger: Floyd, Harry (2026). Claim Ledger: The Guardrail Your Agent Can Reach [structured claims with sources]. The Durability Curve. https://durabilitycurve.com/claims/the-guardrail-your-agent-can-reach/ Quote with attribution and a link. Say if you changed the wording. Not licensed for model training. --- --- title: "Claim Ledger: The Average Is Nobody's Result" description: "Of 255 studies on AI-assisted colonoscopy, 21 split the result by who held the scope. They disagree." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/claims/the-average-is-nobodys-result/" essay: "https://durabilitycurve.com/blog/the-average-is-nobodys-result/" substack: "https://harryfloyd.substack.com/p/the-average-is-nobodys-result" published: "2026-08-04" last_verified: "2026-08-02" law: "Law IV" claims: 35 struck: 0 --- # Claim Ledger: The Average Is Nobody's Result *Of 255 studies on AI-assisted colonoscopy, 21 split the result by who held the scope. They disagree.* Essay: https://durabilitycurve.com/blog/the-average-is-nobodys-result/ Ledger (canonical, cite this): https://durabilitycurve.com/claims/the-average-is-nobodys-result/ Published 2026-08-04 · last verified 2026-08-02 ## What the essay claims Walking in: "The trial says the assistant improves detection by eight points. That is the effect. My people will get roughly that." Walking out: "That is an average over operators the tool may be affecting in opposite directions, and almost nobody checks. When they do check they disagree about who benefits — in randomised trials, on the same endpoint. The average is a property of their operator mix, not of the tool, so it was never going to transfer to mine. I can check my own in an afternoon." The load-bearing move: The average is not a weak estimate of a single effect. It is a mixture of effects that can have opposite signs, weighted by a staffing decision. Povyakalo 2013 is the existence proof; the count is the evidence that the field mostly does not look; the eight that print numbers are the evidence that even looking is not enough unless the numbers are published. ## The claim ladder Strongest first. Each rung names whose behaviour it measures and what it does not reach. | # | Claim | Whose behaviour it measures | What it does not reach | |---|---|---|---| | 1 | Of 255 abstracts on AI-assisted colonoscopy, 21 report the effect split by an operator property and 234 report an average and nothing else | The reporting choices of one literature, at the abstract layer, as of 2026-08-02 | These are deposited abstracts. A study can split in its full text and never say so here. The count measures what is visible at the layer the field summarises itself in. It is also a floor: three patterns each found what the others missed | | 2 | The 21 contradict each other about who benefits. 7 find the larger effect in weaker operators, 4 in stronger, 3 no interaction, 3 nothing significant after stratifying, 1 divergence over time. Both directions appear in randomised trials on the same endpoint | Published subgroup results in one procedure | Does not adjudicate. Several are underpowered after stratification, several are post-hoc, and the piece must not pick a winner. The spread is the finding | | 3 | Only 8 of the 21 print per-stratum estimates. None reports a test of the interaction. The rest report which subgroup reached significance, which is not a test of a difference | The reporting and inference choices of the studies that did look | Does not establish the interaction is real in any of them. It establishes that the published evidence cannot tell you | | 4 | The standard has no field for it. CONSORT-AI 5(iv) asks what expertise users needed; the guidance asks for differences across population subgroups; no item asks for the result across operator subgroups | The reporting guideline, verified item by item at PMC7598943 | Scoped to CONSORT-AI. DECIDE-AI's items were not read and no claim is made about guidelines in general | | 5 | Povyakalo 2013 is the existence proof. Reanalysing a CAD mammography study whose average effect was null: +0.016 sensitivity (95% CI 0.003–0.028) for the 44 least-discriminating readers on 45 easier cancers, −0.145 (95% CI 0.034–0.257) for the 6 most-discriminating on 15 harder ones | 50 readers, 180 mammograms, one reanalysis | Exploratory and post-hoc by the authors' own description, with thresholds derived from the same regression and six readers in the key cell. It proves an average can hide opposite signs. It is not a law, and colonoscopy does not replicate it | | 6 | The average is a property of your operator mix, not of the tool. Change who is on shift and the measured effect changes with no change to the tool | Nothing. An arithmetic consequence of rung 2 | Does not quantify how much any published estimate would move. A reason the number does not transfer, not a correction to it | Scope limits travel in the same sentence as the number, in the body, never in a footnote. ## What would make it wrong Any of these lands and the piece is wrong, not adjustable. Full-text checking of a random sample of the 234 finds that most do split in the full text. Highest-probability failure by a distance. A fourth search pattern finds a large new tranche (>8) all three of mine missed. The floor is then too low to support "rarely". An independent adjudication of the same 255 abstracts returns materially more than 21 under the stated definition. The count is then an artefact of my reading, not the literature. A reporting-guideline item for operator-stratified results turns out to exist in CONSORT-AI. Rung 4 fails outright. (Verified against the item list; the residual risk is a later update.) Someone has already published this count. The prior-art sweep found the ecological version (39216648, 38272274) and the mechanism (Povyakalo) but not the count. If the count exists, the object is not an object. ## The evidence, row by row Status: VERIFIED = primary source opened and the quoted words read off it by a checker who did not write the essay · EXECUTED = a first-party run, the claim is what it printed · REPORTED = carried from a source not opened in full. ### 1. Of 255 abstracts with text in the corpus, 21 report the assistant's effect split by a property of the operator and 234 report an average and nothing else - Status: VERIFIED (read 2026-08-02) - Quote: "HEADLINE: 21/255 studies report the assistant's effect split by a property of the OPERATOR. = 8.2% 234 report an average over operators and nothing else." - Source: first-party run by The Durability Curve, unassisted_baseline_audit.py stdout, operator-split headline block ### 2. Among primary records by PubMed's own publication type the figure is 21 of 157 - Status: VERIFIED (read 2026-08-02) - Quote: "Among primary records only: 21/157 = 13.4%" - Source: first-party run by The Durability Curve, same run, headline block ### 3. Of the 21, only 8 print per-stratum effect estimates for every stratum; 3 print one stratum and a bare null; 10 report only a direction or a significance verdict - Status: VERIFIED (read 2026-08-02) - Quote: "SPLIT_FULL 8 per-stratum estimates for every stratum · SPLIT_PARTIAL 3 one stratum estimated, the other a bare null · SPLIT_DIRECTION 10 a direction or a p-verdict, nothing poolable" - Source: first-party run by The Durability Curve, same run, verdict block ### 4. The splits disagree: 7 find the larger effect in weaker operators, 4 in stronger, 3 no interaction, 3 nothing significant after stratifying, 1 divergence - Status: VERIFIED (read 2026-08-02) - Quote: "LOWER 7 ... HIGHER 4 ... UNIFORM 3 ... NONE 3 ... DIVERGES 1" - Source: first-party run by The Durability Curve, same run, direction block ### 5. Three deliberately dissimilar detectors were run; the obvious one reached 15 records of which 7 are real splits, missing 14 - Status: VERIFIED (read 2026-08-02) - Quote: "pattern A: reached 15 of which real splits 7 precision 0.47 coverage 0.33 of the 21 found" · "Pattern A alone, unread, would have reported 15 records of which 7 are real, missing 14." - Source: first-party run by The Durability Curve, same run, pattern-performance block ### 6. Pattern B reached 39 records at precision 0.36 and coverage 0.67 of the found set; pattern C reached 13 at precision 0.69 and coverage 0.43 - Status: VERIFIED (read 2026-08-02) - Quote: "pattern B: reached 39 of which real splits 14 precision 0.36 coverage 0.67 of the 21 found" · "pattern C: reached 13 of which real splits 9 precision 0.69 coverage 0.43 of the 21 found" · "'coverage' is the share of the found positives a pattern reached, NOT recall" - Source: first-party run by The Durability Curve, same run, same block ### 7. Every record any pattern reached was hand-read, and the script refuses to print the headline while one is unread - Status: VERIFIED (read 2026-08-02) - Quote: "HEADLINE: UNAVAILABLE — 20 record(s) reached by a pattern are unread." (the refusal branch, observed firing before adjudication was complete) - Source: first-party run by The Durability Curve, same script, headline guard ### 8. The PubMed query returns 259 records, 255 with abstracts - Status: VERIFIED (read 2026-08-02) - Quote: "PubMed hits: 259 (retmax 400, fetched 259)" · "classified : 255 (no abstract: 4)" - Source: first-party run by The Durability Curve, same run, header ### 9. In a seeded 20-record sample of the primary average-only class, 9 had open full text and none reported an operator split in its results - Status: VERIFIED (read 2026-08-02) - Quote: Sample drawn with random.Random(20260802) from 136 eligible records; PMC full text retrieved for 9; operator-split patterns returned 0 result-section hits in 8, and in the ninth the only hit was future-tense text in a trial protocol - Source: first-party full-text check by The Durability Curve, seeded sample, scratchpad sample20.json; PMCIDs 10532435, 10381252, 11528738, 11512038, 9796278, 13054503, 12627440, 10011797, 12618418 ### 10. Eleven of the twenty sampled records are not in PMC and their full texts were not read - Status: VERIFIED (read 2026-08-02) - Quote: NCBI ID converter returned "Identifier not found in PMC" for 11 of the 20 sampled PMIDs - Source: first-party full-text check by The Durability Curve, same check, ID-converter response ### 11. Reanalysing a CAD mammography study, use of CAD was associated with a 0.016 increase in sensitivity (95% CI 0.003-0.028) for the 44 least discriminating radiologists on 45 relatively easy cancers - Status: VERIFIED (read 2026-08-02) - Quote: "Use of CAD was associated with a 0.016 increase in sensitivity (95% confidence interval [CI], 0.003-0.028) for the 44 least discriminating radiologists for 45 relatively easy, mostly CAD-detected cancers." - Source: Povyakalo, Alberdi, Strigini & Ayton, Med Decis Making 2013, deposited abstract, Results - URL: https://pubmed.ncbi.nlm.nih.gov/23300205/ ### 12. For the 6 most discriminating radiologists, sensitivity decreased by 0.145 (95% CI 0.034-0.257) on the 15 relatively difficult cancers - Status: VERIFIED (read 2026-08-02) - Quote: "However, for the 6 most discriminating radiologists, with CAD, sensitivity decreased by 0.145 (95% CI, 0.034-0.257) for the 15 relatively difficult cancers." - Source: Povyakalo et al., Med Decis Making 2013, deposited abstract, Results - URL: https://pubmed.ncbi.nlm.nih.gov/23300205/ ### 13. The original study detected no significant average effect, and the reanalysis found CAD helped the less discriminating readers and hindered the more discriminating ones - Status: VERIFIED (read 2026-08-02) - Quote: "It indicates that, despite the original study detecting no significant average effect, CAD helped the less discriminating readers but hindered the more discriminating readers." - Source: Povyakalo et al., Med Decis Making 2013, deposited abstract, Conclusions - URL: https://pubmed.ncbi.nlm.nih.gov/23300205/ ### 14. That reanalysis covered 50 readers interpreting 180 mammograms both with and without computer support - Status: VERIFIED (read 2026-08-02) - Quote: "reanalyzing data from a published study where 50 professionals (\"readers\") interpreted 180 mammograms, both with and without computer support" - Source: Povyakalo et al., Med Decis Making 2013, deposited abstract, Background - URL: https://pubmed.ncbi.nlm.nih.gov/23300205/ ### 15. The reader and case strata were derived after the fact from the same regression, and the authors call the method exploratory - Status: VERIFIED (read 2026-08-02) - Quote: "Using regression estimates, we obtained thresholds for classifying a posteriori the cases (by difficulty) and the readers (by discriminating ability)." · "Our exploratory analysis method reveals unexpected effects." - Source: Povyakalo et al., Med Decis Making 2013, deposited abstract, Method and Conclusions - URL: https://pubmed.ncbi.nlm.nih.gov/23300205/ ### 16. The authors said such differential effects should be assessed when evaluating CAD and similar warning systems - Status: VERIFIED (read 2026-08-02) - Quote: "Such differential effects, although subtle, may be clinically significant and important for improving both computer algorithms and protocols for their use. They should be assessed when evaluating CAD and similar warning systems." - Source: Povyakalo et al., Med Decis Making 2013, deposited abstract, Conclusions - URL: https://pubmed.ncbi.nlm.nih.gov/23300205/ ### 17. In a population-based randomised trial of 4,824 surveillance colonoscopies, CADe increased ADR among low-performing endoscopists (45.5% vs 52.1%; aRR 1.15, 95% CI 1.01-1.30) but not among high-performing ones (65.9% vs 63.5%; aRR 0.96, 95% CI 0.88-1.05) - Status: VERIFIED (read 2026-08-02) - Quote: "Of 5,309 randomized surveillance colonoscopies, 4,824 surveillance colonoscopies were included in the final analysis" · "CADe increased ADR among low-performing (ADR<54.5%) endoscopists (45.5% vs 52.1%; aRR 1.15 [95% CI 1.01-1.30]), but not among high-performing endoscopists (65.9% vs 63.5%; aRR 0.96 [95% CI 0.88-1.05])." - Source: Endoscopy 2026, Galician screening programme, PMID 42365851, deposited abstract, Results - URL: https://pubmed.ncbi.nlm.nih.gov/42365851/ ### 18. That subgroup analysis was prespecified - Status: VERIFIED (read 2026-08-02) - Quote: "Prespecified subgroup analyses evaluated endoscopist baseline performance (median ADR in the standard arm)." - Source: Endoscopy 2026, PMID 42365851, deposited abstract, Methods - URL: https://pubmed.ncbi.nlm.nih.gov/42365851/ ### 19. In a 3,059-patient multicentre randomised trial the gain was larger in experts: ADR of experts 42.3% vs 32.8% (P < .001) and of nonexperts 37.5% vs 32.1% (P = .023) - Status: VERIFIED (read 2026-08-02) - Quote: "From November 2019 to August 2021, 3059 subjects were randomized to AI-assisted colonoscopy (n = 1519) and conventional colonoscopy (n = 1540)." · "ADR of expert (42.3% vs 32.8%; P < .001) and nonexpert endoscopists (37.5% vs 32.1%; P = .023)... were all significantly higher in the AI-assisted colonoscopy." - Source: Clin Gastroenterol Hepatol 2023, PMID 35863686, deposited abstract, Results - URL: https://pubmed.ncbi.nlm.nih.gov/35863686/ ### 20. A 2026 study expected the greatest increase among low-volume and junior endoscopists - Status: VERIFIED (read 2026-08-02) - Quote: "The greatest increase was expected among low-volume and junior endoscopists." - Source: Dis Colon Rectum 2026, PMID 41919624, deposited abstract, Objective - URL: https://pubmed.ncbi.nlm.nih.gov/41919624/ ### 21. It reported that seniors rose from 51.7% to 59.6% (p = 0.02) and juniors from 56.2% to 62.9% (p = 0.12), and concluded the tool correlated with increased detection for experienced endoscopists - Status: VERIFIED (read 2026-08-02) - Quote: "Of 12 senior and 12 junior endoscopists, the seniors had a statistically significant increase ( p = 0.02) from 51.7% to 59.6%, whereas juniors did not (56.2% to 62.9%, p = 0.12)." · "Computer-aided detection in colonoscopy correlated with an increased adenoma detection rate for experienced endoscopists." - Source: Dis Colon Rectum 2026, PMID 41919624, deposited abstract, Results and Conclusions - URL: https://pubmed.ncbi.nlm.nih.gov/41919624/ ### 22. Those two reported increases are 7.9 points and 6.7 points, a gap of 1.2 points, and only the larger crossed p<0.05 - Status: VERIFIED (read 2026-08-02) - Quote: 59.6 − 51.7 = 7.9 · 62.9 − 56.2 = 6.7 · 7.9 − 6.7 = 1.2. Subtraction only; no model, no reanalysis, and no claim that either estimate is wrong - Source: first-party arithmetic on row 21's quoted rates by The Durability Curve, derived from row 21, recomputed by hand 2026-08-02 ### 23. The same study reported the same pattern on two further operator axes: high-volume 51.3% to 59.4% (p = 0.01) against low-volume 56.5% to 63.1% (p = 0.13), and surgeons 49% to 62% (p = 0.02) against gastroenterologists 56.9% to 60.8% (p = 0.15) - Status: VERIFIED (read 2026-08-02) - Quote: "Colorectal surgeons had a statistically significant increase in adenoma detection rate from 49% to 62% ( p = 0.02), but gastroenterologists did not (56.9% to 60.8%, p = 0.15)." · "When comparing 12 high-volume and 12 low-volume endoscopists, the high-volume group had a statistically significant increase (51.3% to 59.4%, p = 0.01), whereas the low-volume group did not (56.5% to 63.1%, p = 0.13)." - Source: Dis Colon Rectum 2026, PMID 41919624, deposited abstract, Results - URL: https://pubmed.ncbi.nlm.nih.gov/41919624/ ### 24. Pooling two randomised trials, CADe raised ADR (RR 1.29, 95% CI 1.16 to 1.42) while examiner experience did not (RR 1.02, 95% CI 0.89 to 1.16), and the authors concluded experience plays a minor role - Status: VERIFIED (read 2026-08-02) - Quote: "use of CADe (RR 1.29; 95% CI: 1.16 to 1.42) and colonoscopy indication, but not the level of examiner experience (RR 1.02; 95% CI: 0.89 to 1.16) were associated with ADR differences in a multivariate analysis." · "Experience appears to play a minor role as determining factor for ADR." - Source: Gut 2022, AID-1 + AID-2, PMID 34187845, deposited abstract, Results and Conclusion - URL: https://pubmed.ncbi.nlm.nih.gov/34187845/ ### 25. The title of the paper naming the inference error is "The Difference Between 'Significant' and 'Not Significant' is not Itself Statistically Significant" - Status: VERIFIED (read 2026-08-02) - Quote: "The Difference Between \"Significant\" and \"Not Significant\" is not Itself Statistically Significant" - Source: Gelman & Stern, The American Statistician 60(4):328-331, 2006, author's deposited PDF, title, page 1 - URL: http://www.stat.columbia.edu/~gelman/research/published/signif4.pdf ### 26. Their point is not that thresholds are arbitrary but that large changes in significance can correspond to small, non-significant changes in the underlying quantities - Status: VERIFIED (read 2026-08-02) - Quote: "Rather, we are pointing out that even large changes in significance levels can correspond to small, nonsignificant changes in the underlying quantities." - Source: Gelman & Stern, 2006, deposited PDF, abstract, page 1 - URL: http://www.stat.columbia.edu/~gelman/research/published/signif4.pdf ### 27. The reporting standard for AI trials asks investigators to state what level of expertise was required of users - Status: VERIFIED (read 2026-08-02) - Quote: "CONSORT-AI 5 (iv) Extension Specify whether there was human–AI interaction in the handling of the input data, and what level of expertise was required of users." - Source: CONSORT-AI extension, Nat Med 26:1364-1374, 2020, checklist table, item 5 (iv), PMC full text - URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC7598943/ ### 28. Its guidance encourages exploring differences in performance across population subgroups, and no item asks for the result across operator subgroups - Status: VERIFIED (read 2026-08-02) - Quote: "Beyond this, investigators should also be encouraged to explore differences in performance and error rates across population subgroups." - Source: CONSORT-AI extension, 2020, discussion of item 19 extension, PMC full text; the full 14-item extension list was read and contains no operator-stratified reporting item - URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC7598943/ ### 29. The COLO-DETECT protocol pre-registered a subgroup analysis by colonoscopist type - Status: VERIFIED (read 2026-08-02) - Quote: "Subgroup analyses will be conducted on colonoscopist type (i.e., non-BCSP accredited vs. BCSP accredited) and indication for colonoscopy (screening vs. symptomatic)." - Source: COLO-DETECT trial protocol, Colorectal Dis 2022, PMID 35680613, protocol full text, statistical analysis section, PMC9796278 - URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC9796278/ ### 30. The same protocol said a range of colonoscopist experience was anticipated and desirable, and specified how it would be measured - Status: VERIFIED (read 2026-08-02) - Quote: "A range of experience amongst participating colonoscopists is anticipated and desirable; to facilitate subgroup analysis by colonoscopist experience, it will be assessed by BCSP accreditation status, lifetime procedure numbers and yearly procedure numbers for the last 3 years" - Source: COLO-DETECT protocol, 2022, protocol full text, PMC9796278 - URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC9796278/ ### 31. The COLO-DETECT results reported one average across 2032 participants at 12 NHS hospitals: adenomas detected in 56·6% versus 48·4%, adjusted odds ratio 1·47 (95% CI 1·21-1·78) - Status: VERIFIED (read 2026-08-02) - Quote: "We did a multicentre, open-label, parallel-arm, pragmatic randomised controlled trial in 12 National Health Service (NHS) hospitals (ten NHS Trusts) in England" · "Between March 29, 2021, and April 6, 2023, 2032 participants (1132 [55·7%] male, 900 [44·3%] female; mean age 62·4 years [SD 10·8]) were recruited and randomly assigned" · "Adenomas were detected in 555 (56·6%) of 980 participants in the CADe-assisted colonoscopy group versus 477 (48·4%) of 986 in the standard colonoscopy group, representing a proportion difference of 8·3% (95% CI 3·9-12·7; adjusted odds ratio 1·47 [95% CI 1·21-1·78], p<0·0001)." - Source: COLO-DETECT, Lancet Gastroenterol Hepatol 2024, PMID 39153491, deposited abstract, Methods and Findings - URL: https://pubmed.ncbi.nlm.nih.gov/39153491/ ### 32. Its randomisation was stratified by age group, sex, indication and NHS Trust, and its abstract names no colonoscopist subgroup - Status: REPORTED (read 2026-08-02) - Quote: Stratification verbatim: "with stratification by age group, sex, colonoscopy indication (screening or symptomatic), and NHS Trust". The absence of a colonoscopist subgroup is an absence in the deposited abstract only — the paper is not in PMC and the full text was not read. - Source: COLO-DETECT, Lancet Gastro Hep 2024, PMID 39153491, deposited abstract, Methods - URL: https://pubmed.ncbi.nlm.nih.gov/39153491/ ### 33. The field's current answer is a study-level one: pooling twenty-eight randomised trials and 23,861 participants, an expert-only subgroup showed a similar effect size (RR 1.19, 95% CI 1.11-1.27) and the review concluded the benefit holds irrespective of endoscopist experience - Status: VERIFIED (read 2026-08-02) - Source: Gastrointest Endosc 2025 systematic review, PMID 39216648, deposited abstract, Results and Conclusions - URL: https://pubmed.ncbi.nlm.nih.gov/39216648/ ### 34. A second meta-analysis reached the same study-level conclusion across twenty-four randomised trials and 17,413 colonoscopies - Status: VERIFIED (read 2026-08-02) - Quote: "Twenty-four RCTs involving 17,413 colonoscopies (AI assisted: 8680; non-AI assisted: 8733) were included." · "Type of AI system used or endoscopist experience did not affect overall improvement in ADR." - Source: Gastrointest Endosc 2024, PMID 38272274, deposited abstract, Results and Conclusions - URL: https://pubmed.ncbi.nlm.nih.gov/38272274/ ### 35. Bainbridge stated in 1983 that unused physical skills decay and a monitoring operator becomes an inexperienced one - Status: VERIFIED (read 2026-08-01) - Quote: "Unfortunately, physical skills deteriorate when they are not used, particularly the refinements of gain and timing. This means that a formerly experienced operator who has been monitoring an automated process may now be an inexperienced one." - Source: Bainbridge, Automatica 19(6):775-779, 1983, §1.1.1 Manual control skills - URL: https://ckrybus.com/static/papers/Bainbridge_1983_Automatica.pdf ## Cite - A claim: "[claim text]" (Floyd, Harry, 2026, https://durabilitycurve.com/claims/the-average-is-nobodys-result/) - The essay: Floyd, Harry (2026). The Average Is Nobody's Result. The Durability Curve. https://durabilitycurve.com/blog/the-average-is-nobodys-result/ - This ledger: Floyd, Harry (2026). Claim Ledger: The Average Is Nobody's Result [structured claims with sources]. The Durability Curve. https://durabilitycurve.com/claims/the-average-is-nobodys-result/ Quote with attribution and a link. Say if you changed the wording. Not licensed for model training. --- --- title: "Claim Ledger: You Cannot Try to Fall Asleep" description: "Sleep is only where you notice it first. Much of what matters works the same way." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/claims/you-cannot-try-to-fall-asleep/" essay: "https://durabilitycurve.com/blog/you-cannot-try-to-fall-asleep/" substack: "https://harryfloyd.substack.com/p/you-cannot-try-to-fall-asleep" published: "2026-08-03" last_verified: "2026-08-02" law: "Law A" claims: 20 struck: 1 --- # Claim Ledger: You Cannot Try to Fall Asleep *Sleep is only where you notice it first. Much of what matters works the same way.* Essay: https://durabilitycurve.com/blog/you-cannot-try-to-fall-asleep/ Ledger (canonical, cite this): https://durabilitycurve.com/claims/you-cannot-try-to-fall-asleep/ Published 2026-08-03 · last verified 2026-08-02 ## What the essay claims Walking in: "I track my sleep so I can improve it. The score tells me how I did. A bad score means I need to take it more seriously and try harder tonight." Walking out: "Sleep is not something I do. It is something that happens when I stop trying to do it, and the score recruits attention, intention and effort at exactly the point where all three are the illness. Worse, sleep is not a special case. Most of what makes my life worth having sits in the same category, and I have been quietly installing scoreboards on all of it." The load-bearing move: Not trackers are bad. A category of goods exists that is destroyed by aiming at it, the category has been named twice independently by people who never cited each other, and a reader can sort their own life with it in ten minutes. The tracker is the most legible instance, not the topic. ## The claim ladder Strongest first. Each rung carries what it does not reach, in the body, in the same breath. | # | Claim | Whose behaviour it measures | What it rests on | What it does not reach | |---|---|---|---|---| | 1 | You cannot fall asleep by trying to. Sleep is automatic and is inhibited by attention to it, explicit intention toward it, and effort at it | People with persistent psychophysiologic insomnia, as modelled by the field that treats them | Espie 2006, a theoretical review that became a treatment model and was revisited in 2023 | It is a model with clinical support, not a settled physical law. State it as the field's leading account of how insomnia persists, not as arithmetic | | 2 | The number can do the damage on its own. Told on waking that a night was poor, when actigraphy says the nights did not differ, daytime function suffers | 22 people with primary insomnia, over 3 mornings | Semler & Harvey 2005 | Small, and in people who already had insomnia. It shows the feedback is sufficient; it does not size the effect or establish it in the general population | | 3 | The effect is not demonstrated in good sleepers. False poor-sleep feedback did not move pre-sleep cognitive arousal or subjective sleep continuity in healthy sleepers | Healthy sleepers in one pilot | The disconfirming study, carried on purpose | This rung protects the piece. The honest reading is that the score does its damage to the anxious, who are exactly the people who buy the device to fix it. Any sentence implying universality is a defect | | 4 | Measuring an intrinsic activity converts it into work. Doing goes up, enjoyment goes down | Ordinary people doing ordinary activities — walking, reading — across six experiments | Etkin 2016 | Ordinary activities, not clinical states. And it cuts both ways: the same finding says measurement works on the behaviour. For a willable behaviour that trade is often worth taking, and the piece must say so | | 5 | Sleep is one member of a category, not a special case. The category was defined before the instruments existed | Nobody's. A conceptual claim about the structure of certain goods | Elster 1981/1983, his own named examples | Philosophical category, not an empirical result. The generalisation beyond sleep is argued, not measured, and the piece must mark it as the argument it is | | 6 | A daily score is an attention–intention–effort machine pointed at an automatic process | One vendor's product design, read from its own documentation | Rungs 1 and 5 joined. Whoop's Recovery Score is delivered each morning from HRV in slow-wave sleep, resting heart rate, respiratory rate and sleep performance — four things unavailable to the will | One vendor, verified. No census across vendors was run and the piece claims none | ## What would make it wrong The premise fails if what people actually score themselves on is dominated by willable behaviours — steps, minutes, pages, sessions — and the unwillable states are a fringe of the market. The check is the headline scores the major devices deliver at wake-up. Status: partially run. Whoop's Recovery Score is confirmed in the unwillable column, delivered each morning and built from four measures unavailable to the will. Oura's Readiness Score is confirmed by name but its documentation has not been read cleanly. Second falsifier, on the generalisation: if the by-products category turns out to have been applied to self-tracking somewhere in the literature, the contribution shrinks to the convergence alone. The sweep found nothing, and the sweep is not proof of absence. ## The evidence, row by row Status: VERIFIED = primary source opened and the quoted words read off it by a checker who did not write the essay · EXECUTED = a first-party run, the claim is what it printed · REPORTED = carried from a source not opened in full. ### 1. Sleep normalcy is a relatively automatic process. Consequently it is vulnerable, and may be inhibited, by focused attention and by direct attempts to control its expression - Status: VERIFIED (read 2026-08-02) - Quote: "The argument is that sleep normalcy is a relatively automatic process. Consequently, it is vulnerable, and may be inhibited, by focused attention and by direct attempts to control its expression." - Source: Espie, Broomfield, MacMahon, Macphee & Taylor, The attention-intention-effort pathway in the development of psychophysiologic insomnia: a theoretical review, Sleep Medicine Reviews 10(4):215-45, Aug 2006, PMID 16809056, DOI 10.1016/j.smrv.2006.03.002, abstract, sentences 6-7 - URL: https://pubmed.ncbi.nlm.nih.gov/16809056/ ### 1b. The model is named the attention-intention-effort pathway, and it is offered as an explanatory model of how psychophysiologic insomnia develops and is maintained - Status: VERIFIED (read 2026-08-02) - Quote: "This paper proposes an explanatory model, that we call the attention-intention-effort pathway." - Source: Espie et al. 2006, as above, abstract - URL: https://pubmed.ncbi.nlm.nih.gov/16809056/ ### 1c. Psychophysiologic insomnia is the most common form of persistent primary insomnia, and its behavioural phenotype includes conditioned arousal, sleep-incompatible behaviour and sleep preoccupation - Status: VERIFIED (read 2026-08-02) - Quote: "Psychophysiologic insomnia (PI) is the most common form of persistent primary insomnia. Its 'behavioral phenotype', comprising elements such as conditioned arousal, sleep-incompatible behavior and sleep preoccupation, has not changed markedly across several generations of diagnostic nosology." - Source: Espie et al. 2006, as above, abstract, opening - URL: https://pubmed.ncbi.nlm.nih.gov/16809056/ ### 3. Twenty-two individuals with primary insomnia were given positive or negative feedback about their sleep immediately on waking on three consecutive mornings - Status: VERIFIED (read 2026-08-02) - Quote: "twenty-two individuals with primary insomnia receiv[ed] positive or negative feedback about their sleep immediately on waking on three consecutive mornings" - Source: Semler & Harvey, Misperception of sleep can adversely affect daytime functioning in insomnia, Behaviour Research and Therapy 43(7):843–856, 2005, PMID 15896282, method - URL: https://pubmed.ncbi.nlm.nih.gov/15896282/ ### 4. Objective sleep, by actigraphy, did not differ across the three nights or the two feedback conditions - Status: VERIFIED (read 2026-08-02) - Quote: "Objective sleep on each of the three nights was estimated by actigraphy and did not differ across the three nights or the two feedback conditions" - Source: Semler & Harvey 2005, results - URL: https://pubmed.ncbi.nlm.nih.gov/15896282/ ### 5. Negative feedback was associated with more negative thoughts, sleepiness, monitoring for sleep-related threat, and safety behaviours during the day - Status: VERIFIED (read 2026-08-02) - Quote: "negative feedback was associated with more negative thoughts, sleepiness, monitoring for sleep-related threat, and safety behaviours during the day, relative to positive feedback" - Source: Semler & Harvey 2005, results - URL: https://pubmed.ncbi.nlm.nih.gov/15896282/ ### 7. Fifty-four healthy sleepers were randomly allocated to receive good or poor false sleep feedback in the form of a numerical sleep score, and were told it reflected their habitual sleep - Status: VERIFIED (read 2026-08-02) - Quote: "A total of 54 healthy sleepers (Mage = 30.19 years; SDage = 12.94 years) were randomly allocated to receive good, or poor, false sleep feedback, in the form of a numerical sleep score. Participants were informed that this feedback was a true reflection of their habitual sleep." - Source: Robson, Ellis & Elder, Poor false sleep feedback does not affect pre-sleep cognitive arousal or subjective sleep continuity in healthy sleepers: a pilot study, Sleep and Biological Rhythms 20(4):467–472, 2022, abstract - URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC10899903/ ### 8. In healthy sleepers the false score changed nothing measured - Status: VERIFIED (read 2026-08-02) - Quote: "There were no significant differences between good and poor feedback groups in terms of pre-sleep cognitive arousal, or subjective sleep continuity, before or after the presentation of the sleep feedback." - Source: Robson, Ellis & Elder 2022, abstract - URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC10899903/ ### 9. That study frames itself as being about wearables, not about laboratory deception - Status: VERIFIED (read 2026-08-02) - Quote: "Modern wearable devices calculate a numerical metric of sleep quality (sleep feedback), which are intended to allow users to monitor and, potentially, improve their sleep." - Source: Robson, Ellis & Elder 2022, abstract, first line - URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC10899903/ ### 10. Measurement increases how much of an activity people do and simultaneously reduces how much they enjoy it - Status: VERIFIED (read 2026-08-02) - Quote: "while measurement increases how much of an activity people do (e.g., walk or read more), it can simultaneously reduce how much people enjoy those activities" - Source: Etkin, The Hidden Cost of Personal Quantification, Journal of Consumer Research 42(6):967–984, 2016, abstract - URL: https://academic.oup.com/jcr/article-abstract/42/6/967/2358309 ### 11. The mechanism is that attention to output makes an enjoyable activity feel like work, undermining intrinsic motivation - Status: VERIFIED (read 2026-08-02) - Quote: "By drawing attention to output, measurement can make enjoyable activities feel more like work, which reduces their enjoyment." - Source: Etkin 2016, abstract - URL: https://academic.oup.com/jcr/article-abstract/42/6/967/2358309 ### 12. Six experiments - Status: VERIFIED (read 2026-08-02) - Quote: "Six experiments demonstrate that while measurement increases how much of an activity people do…" - Source: Etkin 2016, abstract - URL: https://academic.oup.com/jcr/article-abstract/42/6/967/2358309 ### 13. A state is essentially a by-product when it can come about as a result of action but cannot be brought about intentionally, because the attempt precludes the state - Status: VERIFIED (read 2026-08-02) - Quote: "can only come about as by-products of actions undertaken for other ends… can never be brought about intentionally because the very attempt to do so precludes the state one is trying to bring about" - Source: Elster, Sour Grapes: Studies in the Subversion of Rationality, Cambridge University Press, ch. 2 "States that are essentially by-products", pp. 43–109, chapter 2 summary - URL: https://doi.org/10.1017/CBO9781316494172.004 ### 14. Elster names spontaneity as an example of such an inaccessible state - Status: VERIFIED (read 2026-08-02) - Quote: "I cited spontaneity as an example of such 'inaccessible' states." - Source: Elster, Sour Grapes ch. 2, chapter 2 introduction - URL: https://doi.org/10.1017/CBO9781316494172.004 ### 15. The self-defeating list includes the desire to forget, to believe, to desire, to sleep, to laugh, and to overcome stuttering - Status: REPORTED (read 2026-08-02) - Quote: "the desire to forget, the desire to believe, the desire to desire …, the desire to sleep, the desire to laugh (one cannot tickle oneself), and the desire to overcome stuttering" - Source: Elster, Sour Grapes ch. 2 / States that are essentially by-products, Social Science Information 20(3), 1981, ch. 2, as reproduced in secondary summaries - URL: https://journals.sagepub.com/doi/10.1177/053901848102000301 ### 16. Chapter 2 of Sour Grapes is titled "States that are essentially by-products" and runs pp. 43-109 - Status: VERIFIED (read 2026-08-02) - Quote: "Chapter: 2 ('States that are essentially by-products') · Pages: 43-109" - Source: Cambridge Core chapter listing, publisher metadata block - URL: https://doi.org/10.1017/CBO9781316494172.004 ### 17. Orthosomnia names patients whose pursuit of the sleep score interfered with their sleep, and who trusted the device over the clinician - Status: REPORTED (read 2026-08-02) - Quote: — quote not yet pinned at the primary — - Source: Baron, Abbott, Jao, Manalo & Mullen, Orthosomnia: Are Some Patients Taking the Quantified Self Too Far?, Journal of Clinical Sleep Medicine, 2017, — - URL: https://jcsm.aasm.org/doi/10.5664/jcsm.6472 ### 18. Whoop's Recovery Score is delivered each morning and is computed from HRV measured during slow-wave sleep, resting heart rate, respiratory rate and sleep performance - Status: REPORTED (read 2026-08-02) - Quote: "The Recovery Score (0-100) … a daily readiness indicator that aggregates heart rate variability, resting heart rate and sleep quality … reported each morning after your sleep cycle completes" - Source: Whoop product documentation, vendor marketing and developer documentation - URL: https://developer.whoop.com/docs/whoop-101/ ### 19. A 2024 cross-sectional study of 523 adults estimated orthosomnia prevalence at 3–14% depending on definition, with 35.8% of participants regularly using sleep-tracking devices - Status: REPORTED (read 2026-08-02) - Quote: — figures from secondary summary; primary not yet opened — - Source: Brain Sciences, 2024, — - URL: https://doaj.org/article/50f708de43bd4086a44688d71f87f333 ### 20. A validated Bergen Orthosomnia Scale was published in 2025 - Status: REPORTED (read 2026-08-02) - Quote: — - Source: secondary summary only, — ## Struck or excluded before publication Not claims the essay makes: claims it stopped making, with the reason. - ~~The three components act in concert and hierarchically to turn acute stress-induced insomnia into the persistent self-perpetuating kind~~ (STRUCK 2026-08-02): This wording is not in the 2006 abstract; it came from a secondary summary of the 2023 revisit, which is paywalled (HTTP 402) and was not read. The body uses rows 1/1b/1c instead and makes no claim about hierarchy or about the acute-to-chronic transition ## Cite - A claim: "[claim text]" (Floyd, Harry, 2026, https://durabilitycurve.com/claims/you-cannot-try-to-fall-asleep/) - The essay: Floyd, Harry (2026). You Cannot Try to Fall Asleep. The Durability Curve. https://durabilitycurve.com/blog/you-cannot-try-to-fall-asleep/ - This ledger: Floyd, Harry (2026). Claim Ledger: You Cannot Try to Fall Asleep [structured claims with sources]. The Durability Curve. https://durabilitycurve.com/claims/you-cannot-try-to-fall-asleep/ Quote with attribution and a link. Say if you changed the wording. Not licensed for model training. --- --- title: "Claim Ledger: Your Robot Coworker Is Still a Pilot" description: "Built, shipped, installed, working: four counts, quoted as one. Almost nobody publishes the fourth." author: "Harry Floyd" publication: "The Durability Curve" canonical: "https://durabilitycurve.com/claims/your-robot-coworker-is-still-a-pilot/" essay: "https://durabilitycurve.com/blog/your-robot-coworker-is-still-a-pilot/" substack: "https://harryfloyd.substack.com/p/your-robot-coworker-is-still-a-pilot" published: "2026-08-01" last_verified: "2026-08-01" law: "Law IV" claims: 24 struck: 5 --- # Claim Ledger: Your Robot Coworker Is Still a Pilot *Built, shipped, installed, working: four counts, quoted as one. Almost nobody publishes the fourth.* Essay: https://durabilitycurve.com/blog/your-robot-coworker-is-still-a-pilot/ Ledger (canonical, cite this): https://durabilitycurve.com/claims/your-robot-coworker-is-still-a-pilot/ Published 2026-08-01 · last verified 2026-08-01 ## What the essay claims Most people who follow robotics believe the deployment numbers are roughly known and roughly rising: some thousands of humanoids are out there doing jobs, and the number goes up each quarter. The belief this piece changes: there is no such number. What exists is a ladder of four different counts, the industry publishes the bottom two, and the coverage that reaches the reader flattens all four into one. ## The claim ladder Each rung names whose behaviour it measures and what it does not reach. | Rung | Claim | Measures | Does NOT reach | |---|---|---|---| | A (safest) | Built, shipped, installed and working are four different quantities | What company documents say, verbatim | Whether anyone is actually misled | | B | Companies' own documents hold them apart; the coverage merges them | Document text vs press renderings of the same event | Intent, on anyone's part | | C | Disclosure tracks legal compulsion | Three regimes observed: mandated-volume prospectus, US filing without volume mandate, no obligation | Causation, and n=3 regimes is an observation not a law | | D (most exposed) | Nobody publishes a count of robots doing unsupervised paid work | What this pass searched on 2026-08-01 | Proof of universal absence — it is a negative claim and must be written to survive one counter-example | Rung D is the headline and the weakest rung. It must be phrased so that a single company publishing such a number narrows it rather than falsifies the piece. ## What would make it wrong The piece is wrong, and should be corrected or pulled, if any of these hold: Any company, filing or trade body publishes a count of humanoid robots performing unsupervised paid work. Rung D falls; the piece narrows to "almost nobody". Agility's 24 June 2026 SEC exhibit contains a robot unit count I missed. The keystone absence fails. THIS FIRED (2026-08-01). A review found the count in Exhibit 99.2 (I had read only 99.1). The "zero counts" claim was false and is retired. The section was rewritten to the corrected and stronger finding: one disclosed count, and it is an order (1,000), with the deployed figure structured into warrants but never stated. The keystone did not fall; it moved up one rung and got sharper. BMW's 27 Feb 2026 release states a robot count for Spartanburg. The second absence fails. Unitree's clarification does not distinguish produced from shipped. The spine fails. Falsifiers 2–4 are settled by opening one document each. That is the point. ## The evidence, row by row Status: VERIFIED = primary source opened and the quoted words read off it by a checker who did not write the essay · EXECUTED = a first-party run, the claim is what it printed · REPORTED = carried from a source not opened in full. ### 1. Unitree published a clarification because its 2025 shipment volume was being misreported - Status: VERIFIED (read 2026-08-01) - Quote: "many pieces of misinformation regarding our company's 2025 shipment volume have been circulating online" - Source: Unitree Robotics, official news post, opening paragraph - URL: https://shop.unitree.com/blogs/news/clarification-regarding-unitrees-2025-sales-data ### 2. Unitree shipped more than 5,500 humanoid robots in 2025 - Status: VERIFIED (read 2026-08-01) - Quote: "Unitree's actual shipment volume of humanoid robots exceeded 5,500 units" - Source: Unitree Robotics, 2025 sales clarification, 22 Jan 2026, shipment figure - URL: https://shop.unitree.com/blogs/news/clarification-regarding-unitrees-2025-sales-data ### 3. Unitree produced more than 6,500 humanoid robots in 2025, a different and larger number than it shipped - Status: VERIFIED (read 2026-08-01) - Quote: "total mass-production output of 2025 exceeded 6,500 units" - Source: Unitree Robotics, 2025 sales clarification, 22 Jan 2026, production figure - URL: https://shop.unitree.com/blogs/news/clarification-regarding-unitrees-2025-sales-data ### 4. Unitree states its shipment figure counts delivery to end customers, and that its order volume is higher still - Status: VERIFIED (read 2026-08-01) - Quote: "quantity actually sold and delivered to end customers, not order volume; the order volume is higher" - Source: Unitree Robotics, 2025 sales clarification, 22 Jan 2026, definition note - URL: https://shop.unitree.com/blogs/news/clarification-regarding-unitrees-2025-sales-data ### 5. Unitree warns against adding different robot types together - Status: VERIFIED (read 2026-08-01) - Quote: "directly combining the numbers of different types of robots together for comparison" - Source: Unitree Robotics, 2025 sales clarification, 22 Jan 2026, closing caution - URL: https://shop.unitree.com/blogs/news/clarification-regarding-unitrees-2025-sales-data ### 6. Agility Robotics announced a 2.5 billion dollar merger with Churchill Capital Corp XI - Status: VERIFIED (read 2026-08-01) - Quote: "Agility Robotics to Go Public Through $2.5 Billion Merger with Churchill Capital Corp XI" - Source: SEC Form 425, Exhibit 99.1, joint press release, 24 June 2026, headline - URL: https://www.sec.gov/Archives/edgar/data/0002074973/000121390026071290/ea029548401ex99-1.htm ### 7. Agility reports 300 million dollars of contracted orders rather than a number of robots - Status: VERIFIED (read 2026-08-01) - Quote: "more than $300 million of multi-year contracted Digit v5 orders secured to date" - Source: SEC Form 425, Exhibit 99.1, 24 June 2026, highlights list - URL: https://www.sec.gov/Archives/edgar/data/0002074973/000121390026071290/ea029548401ex99-1.htm ### 8. Agility reports 9 customer facilities and 65,000 operating hours in place of a robot count - Status: VERIFIED (read 2026-08-01) - Quote: "Through deployment commitments across nine customer facilities, Digit has accumulated more than 65,000 hours of operation" - Source: SEC Form 425, Exhibit 99.1, 24 June 2026, commercial traction paragraph - URL: https://www.sec.gov/Archives/edgar/data/0002074973/000121390026071290/ea029548401ex99-1.htm ### 9. Agility's 10,000 unit figure describes annual factory capacity, not robots made - Status: VERIFIED (read 2026-08-01) - Quote: "RoboFab, its full-scale humanoid manufacturing facility designed to support production of up to 10,000 units annually" - Source: SEC Form 425, Exhibit 99.1, 24 June 2026, manufacturing paragraph - URL: https://www.sec.gov/Archives/edgar/data/0002074973/000121390026071290/ea029548401ex99-1.htm ### 10. Agility's press-release exhibit carries no robot unit count; the count sits in the companion investor-presentation exhibit of the same filing - Status: VERIFIED (read 2026-08-01) - Quote: "more than $300 million of multi-year contracted Digit v5 orders secured to date" - Source: SEC Form 425, Exhibit 99.1 (press release), 24 June 2026; parsed in full, no unit count present, whole press release, six pages - URL: https://www.sec.gov/Archives/edgar/data/0002074973/000121390026071290/ea029548401ex99-1.htm ### 11. Agility's investor presentation discloses one robot count, 1,000, and defines it as customer ORDERS not built/shipped/installed units - Status: VERIFIED (read 2026-08-01) - Quote: "Reflects customer orders for Digit v5, as of May 2026 … and relates to 1,000 Digit v5 robots with three-year term RaaS contract" - Source: SEC Form 425, Exhibit 99.2 (Investor Presentation, June 2026), same 8-K accession 000121390026071290, footnote (1) to the orders figure - URL: https://www.sec.gov/Archives/edgar/data/0002074973/000121390026071290/ea029548401ex99-2.htm ### 12. The same RaaS contract ties buyer warrants to robots DEPLOYED, treating deployment as a separate event from the order and giving it no number - Status: VERIFIED (read 2026-08-01) - Quote: "includes warrants issued to purchaser vesting proportionately to robots deployed; figures are not a measure of current period revenue" - Source: SEC Form 425, Exhibit 99.2 (Investor Presentation, June 2026), footnote (1) to the orders figure - URL: https://www.sec.gov/Archives/edgar/data/0002074973/000121390026071290/ea029548401ex99-2.htm ### 13. Agility's presentation says its robots are currently deployed in customer facilities but gives no deployed count; its unit-scale numbers are production scenarios and an illustrative chart - Status: VERIFIED (read 2026-08-01) - Quote: "currently deployed in customer facilities" - Source: SEC Form 425, Exhibit 99.2 (Investor Presentation, June 2026), metrics slide; the "1k/5k/10k units/yr" cost curve and the "installed base" revenue chart are marked "for illustrative purposes only" - URL: https://www.sec.gov/Archives/edgar/data/0002074973/000121390026071290/ea029548401ex99-2.htm ### 14. BMW says the Spartanburg robot moved more than 90,000 components over about 1,250 operating hours - Status: VERIFIED (read 2026-08-01) - Quote: "In total, it moved more than 90,000 components and covered approximately 1.2 million steps in around 1,250 operating hours" - Source: BMW Group press release, 27 Feb 2026, Spartanburg results - URL: https://www.press.bmwgroup.com/global/article/detail/T0455864EN/bmw-group-to-deploy-humanoid-robots-in-production-in-germany-for-the-first-time?language=en ### 15. BMW says that within ten months the robot supported production of more than 30,000 X3 vehicles (secondary coverage's "eleven months" is not BMW's own figure) - Status: VERIFIED (read 2026-08-01) - Quote: "Within ten months, the robot Figure 02 supported the production of more than 30,000 BMW X3, working ten-hour shifts daily from Monday to Friday" - Source: BMW Group press release, 27 Feb 2026, Spartanburg results - URL: https://www.press.bmwgroup.com/global/article/detail/T0455864EN/bmw-group-to-deploy-humanoid-robots-in-production-in-germany-for-the-first-time?language=en ### 16. BMW's release refers to the robot in the singular and never states how many robots were at Spartanburg - Status: VERIFIED (read 2026-08-01) - Quote: "the robot Figure 02 supported the production of more than 30,000 BMW X3" - Source: BMW Group press release, 27 Feb 2026; whole release checked for a unit count, none present, whole release - URL: https://www.press.bmwgroup.com/global/article/detail/T0455864EN/bmw-group-to-deploy-humanoid-robots-in-production-in-germany-for-the-first-time?language=en ### 17. Agility's revenue chart scales an installed base from 2,500 to 15,000 Digit v5 and marks itself illustrative - Status: VERIFIED (read 2026-08-01) - Quote: "$0 $250 $500 $750 $1,000 $1,250 $1,500 2,500 5,000 7,500 10,000 12,500 15,000 Illustrative annual RaaS revenue(1) … Installed base (total # of Digit v5 deployed)(2)" with note "(2) For illustrative purposes only." - Source: SEC Form 425, Exhibit 99.2 (Investor Presentation, June 2026), unit-economics slide, installed-base axis - URL: https://www.sec.gov/Archives/edgar/data/2074973/000121390026071290/ea029548401ex99-2.htm ### 18. BMW announces a next pilot at Leipzig in high-voltage battery assembly, again describing a humanoid robot in the singular and giving no unit count - Status: VERIFIED (read 2026-08-01) - Quote: "The deployment in Leipzig is focusing on testing a multifunctional application of the robot" in the "assembly of high-voltage batteries and in component manufacturing" - Source: BMW Group press release, 27 Feb 2026, Leipzig section - URL: https://www.press.bmwgroup.com/global/article/detail/T0455864EN/bmw-group-to-deploy-humanoid-robots-in-production-in-germany-for-the-first-time?language=en ### 19. AGIBOT announced its 10,000th humanoid robot as a production milestone - Status: VERIFIED (read 2026-08-01) - Quote: "announced the rollout of its 10,000th humanoid robot" - Source: AGIBOT press release via PR Newswire, 30 Mar 2026, opening - URL: https://www.prnewswire.com/news-releases/agibot-reaches-10-000-units-as-real-world-demand-for-robots-accelerates-302728295.html ### 20. AGIBOT gives no number for how many of the 10,000 are in real-world use, only an adjective - Status: VERIFIED (read 2026-08-01) - Quote: "Of the 10,000 humanoid robots produced, a significant portion is already active in real-world environments" - Source: AGIBOT press release via PR Newswire, 30 Mar 2026, deployment paragraph - URL: https://www.prnewswire.com/news-releases/agibot-reaches-10-000-units-as-real-world-demand-for-robots-accelerates-302728295.html ### 21. IFR counts 542,000 industrial robots installed in 2024 - Status: VERIFIED (read 2026-08-01) - Quote: "542,000 robots installed in 2024" - Source: International Federation of Robotics, World Robotics 2025 press release, installations figure - URL: https://ifr.org/ifr-press-releases/news/global-robot-demand-in-factories-doubles-over-10-years ### 22. IFR separately counts 4,664,000 industrial robots in operational stock, a different number from installations - Status: VERIFIED (read 2026-08-01) - Quote: "The total number of industrial robots in operational use worldwide was 4,664,000 units in 2024" - Source: International Federation of Robotics, World Robotics 2025 press release, operational stock figure - URL: https://ifr.org/ifr-press-releases/news/global-robot-demand-in-factories-doubles-over-10-years ### 23. Unitree's prospectus shows about 9 percent of humanoid robot revenue came from real industrial applications, with about 74 percent from research and education - Status: REPORTED (read 2026-08-01) - Quote: "scientific research and education accounted for 73.6% of humanoid robot revenue, commercial consumption accounted for 17.39%, and real industrial applications only accounted for 9%" - Source: Unitree IPO prospectus, reported by TechFlow and corroborated by a second outlet; prospectus itself not opened, prospectus revenue split, first three quarters of 2025 - URL: https://www.techflowpost.com/en-US/article/31730 ### 24. Tesla has disclosed no Optimus production unit count, and says initial units are for internal data collection rather than customers - Status: REPORTED (read 2026-08-01) - Quote: "training data collection and further functionality development" - Source: Tesla Q2 2026 earnings materials, as reported; Tesla's own deck not opened directly, Optimus outlook - URL: https://www.techtimes.com/articles/321012/20260720/tesla-optimus-production-count-remains-zero-q2-earnings-call-looms-wednesday.htm ## Struck or excluded before publication Not claims the essay makes: claims it stopped making, with the reason. - ~~Two robots at BMW Spartanburg~~ (EXCLUDED): Secondary sources only. BMW's own release uses the singular and gives no count. Writing it would commit the error the piece prosecutes. - ~~Figure has ~740 robots deployed~~ (EXCLUDED): The primary (Adcock, X, 20 June 2026) returned HTTP 402 and could not be opened. The wording that circulates describes Figure's own premises, not customer deployments. Excluded from the prose entirely — do not reintroduce it as "high hundreds" either. - ~~Agility discloses zero robot counts~~ (EXCLUDED): FALSE, corrected 2026-08-01 after a review. The press-release exhibit has none, but the investor-presentation exhibit of the same filing discloses 1,000 (as orders). Original draft asserted zero because only Exhibit 99.1 was read. - ~~About 3 to 4 percent of humanoid shipments do operational work~~ (EXCLUDED): TechFlow's own derivation from the 9 percent industrial share, not a prospectus disclosure. Excluded as a derived number presented as a disclosed one. - ~~5,500 shipped implies ~495 industrial units~~ (EXCLUDED): Cross-basis: 9 percent is a revenue share over three quarters, 5,500 is an annual unit count. Multiplying them is the error the piece is about. ## Cite - A claim: "[claim text]" (Floyd, Harry, 2026, https://durabilitycurve.com/claims/your-robot-coworker-is-still-a-pilot/) - The essay: Floyd, Harry (2026). Your Robot Coworker Is Still a Pilot. The Durability Curve. https://durabilitycurve.com/blog/your-robot-coworker-is-still-a-pilot/ - This ledger: Floyd, Harry (2026). Claim Ledger: Your Robot Coworker Is Still a Pilot [structured claims with sources]. The Durability Curve. https://durabilitycurve.com/claims/your-robot-coworker-is-still-a-pilot/ Quote with attribution and a link. Say if you changed the wording. Not licensed for model training. ---