Start here: Start with the problem you have.

The Durability Curve is Harry Floyd’s publication on how AI changes what work is worth.

  • Run one thing on your own work
  • Read the piece to start with
  • Check the evidence, claim by claim

What are you working on?

Deciding whether to trust an agent

5 pieces 53 min of reading

  1. 1RunFive sliders

    Why does my AI agent fail on long jobs?

    An agent that gets each step right 95% of the time finishes a 40-step job 13% of the time. At 98% it finishes 45%.

    Run the Marathon Calculator
  2. 2Read10 min

    How Reliable Is Your AI Agent?

    A month running an autonomous agent. Everyone who does comes back having built the same thing: a verifier.

    Read it: How Reliable Is Your AI Agent?
  3. 3Check20 of 20 verified

    The Guardrail Your Agent Can Reach

    A companion piece on the same problem: every claim, with its source, and how far it was checked.

    See the evidence: The Guardrail Your Agent Can Reach

For your teamThe Seven-Layer Agent AuditTwo-page worksheet · PDF · no sign-up

Download The Seven-Layer Agent Audit (PDF)

Why your benchmark is lying to you

5 pieces 47 min of reading

  1. 1RunChoose one metric, then answer its audit

    My metric is green. Should I believe it?

    Name the number you steer by. You get the specific way it can mislead you, and the check that catches it.

    Run the Metric Validity Audit
  2. 2Read11 min

    Ten Lines of Code Scored 100%. One Agent Broke Eight Benchmarks.

    Not one task was actually solved, and the same blind spot is sitting in your own dashboard.

    Read it: Ten Lines of Code Scored 100%. One Agent Broke Eight Benchmarks.
  3. 3Check57 of 68 verified

    Your AI Grader Is Only as Good as Its Answer Key

    A companion piece on the same problem: every claim, with its source, and how far it was checked.

    See the evidence: Your AI Grader Is Only as Good as Its Answer Key

Building an agent that gets better

5 pieces 46 min of reading

  1. 1RunYour tasks, then four questions on each

    Should this be one agent or a team?

    List your tasks and answer four questions on each. You get a ranked call on which ones one agent in a loop handles better than a team of agents.

    Run the Multi-Agent Decision
  2. 2Read10 min

    Same Model, Different Product: The Case for Harness Engineering

    Harness engineering, the code wrapped around an AI model, now drives more of the performance gap than the model you pick.

    Read it: Same Model, Different Product: The Case for Harness Engineering
  3. 3Check21 of 34 verified

    The Limit Said 10. The Loop Made 500 Calls.

    A companion piece on the same problem: every claim, with its source, and how far it was checked.

    See the evidence: The Limit Said 10. The Loop Made 500 Calls.

For your teamThe Skill Bill of MaterialsTwo-page worksheet · PDF · no sign-up

Download The Skill Bill of Materials (PDF)

Where the money actually lands

5 pieces 57 min of reading

  1. 1RunTwo questions · 60 seconds

    How long will my AI advantage last?

    Two questions give you a dated window: roughly when your AI advantage stops paying, and which layer to build on next.

    Run the Two-Rate Diagnostic
  2. 2Read10 min

    How Long Until Your AI Edge Stops Paying?

    You adopted AI everywhere and it still didn't pay. The scarce layer keeps the money, until its clock runs out.

    Read it: How Long Until Your AI Edge Stops Paying?
  3. 3Check11 of 17 verified

    You Can Win the Wrong Game for Years

    A companion piece on the same problem: every claim, with its source, and how far it was checked.

    See the evidence: You Can Win the Wrong Game for Years

What the machine is doing to you

5 pieces 57 min of reading

  1. 1Run9 min + a cold check later

    Use AI to Learn Without Getting Worse at It

    Answer-giving AI left learners worse off once it was gone. Three setups built to keep the thinking with you, and a way to check what stuck.

    Open the runbook
  2. 2Read15 min

    The Difficulty You're Escaping Was Making You

    AI can lift the effort out of almost anything you find hard. Some of that effort was the thing turning you into someone.

    Read it: The Difficulty You're Escaping Was Making You
  3. 3Check10 of 11 verified

    Confidently Wrong

    Every claim in this piece, with its source, and how far it was checked.

    See the evidence: Confidently Wrong

Know what you want? All 7 tools Every piece, by path The essay map

§01 · The lens Two minutes

What survives when the surface changes?

Every piece here asks it of whatever looks impressive. Here it is, with a test you can run on your own work.

The surface Last 30 days
  1. Anthropic Claude Sonnet 5.5
  2. OpenAI GPT-6 Sol / Luna
  3. Anthropic Claude Opus 5.5
  4. xAI Grok 4.7
  5. DeepSeek DeepSeek-V4.1-Flash
  6. OpenAI GPT-6 Astra
  7. Google Gemini 3.8 Flash
  8. Meta Muse Spark 1.3
  9. Anthropic Claude Fable 5.1

9 frontier model releases in 30 days 88 since Nov 2022

Each row links to the lab’s own announcement. From the log behind the AI release chart, data through .

The surfacehere called canopy

  • this month’s model
  • the demo
  • the benchmark score
  • the prompt that works this week

Repriced every few weeks.

What lastshere called substrate

  • the tests you trust
  • your data
  • the workflow it plugs into
  • the record of what broke
  • the rule you decide by
  • the trust you’ve earned

Still has a job after the repricing.

Where the value goes nexthere called the bottleneck

To the neighbouring layer that is now hardest to get: up to whoever checks and chooses, or down to power, chips and permission. Something built to last still loses if the value moves away from it.

The test, in three moves

  1. Commit

    Write down what you think will last, before anything happens.

  2. Name the change

    Be exact: a cheaper open model, a channel that stops favouring you, a platform that ships your feature.

  3. Count

    What still has a job is how durable it was.

Commit before you name the change. Do it the other way round and whatever survived will look like the lasting part.

A worked case, at my own expense

I wrote a script to run the seven questions automatically, then threw it away, because a script reads your file names rather than your setup and returns a confident verdict on any stack it does not recognise. The automation was canopy. The questions were substrate. I had built the wrong one first.

Harry Floyd, on building The Seven-Layer Agent Audit

Three questions to carry

  1. What is the part that lasts here?
  2. What happens to it when the ground moves?
  3. Is the value moving toward it, or away?

The long versionthe original Start Here essay, 6 min

§02 · The record 380 claims · 14 ledgers

How to check what you read here

  • Verified: 278, Executed: 58, Checked: 16, Reported: 28.

    278 of 380

    claims verified at the primary source, across 14 pieces

    Open the claim ledger
  • 5 laws

    each with what would prove it wrong; all amended after testing

    Read the five laws
  • 88 releases

    frontier models from 10 labs since Nov 2022; the newest 23 checked to the day

    Open the release log

Who writes it

I’m Harry Floyd. I run a research system that turns AI, product, and market noise into instruments for seeing what survives when the surface changes.

About the publication

The newsletter · Free

Subscribe to get the next structural lens in your inbox.

15 new pieces here in the 30 days to 29 Sep 2026. 76 of the 82 are free in full.

Signup completes on Substack. Unsubscribe from any email.

Already subscribed? Every piece is in the archive, or follow by RSS.