---
title: "You Own an Asset You've Never Scored"
description: "The file your agent reads first is capital. Score it out of sixteen, fix the thinnest mark, and run the one check that tells you whether the spec, not the model, was your limit."
author: "Harry Floyd"
publication: "The Durability Curve"
canonical: "https://durabilitycurve.com/blog/the-specification-asset/"
date: "2026-09-21"
series: "FLAGSHIP, STANDALONE"
format: "markdown mirror of the canonical HTML page; figures are named, not embedded"
---

# You Own an Asset You've Never Scored

*The file your agent reads first is capital. Score it out of sixteen, fix the thinnest mark, and run the one check that tells you whether the spec, not the model, was your limit.*

By Harry Floyd · 2026-09-21 · canonical: https://durabilitycurve.com/blog/the-specification-asset/

Open the file your agent reads first. The CLAUDE.md, the system prompt, the AGENTS.md, whatever your setup calls it. You have one, and the same model behaves differently with it than without it. Now answer the harder question. Is it any good? Not "does it work today", but is it gaining value or losing it, and how would you know before it costs you.

Most people cannot answer. They can name the model they run and the prompt they tuned last week. The file that decides whether the model is worth running at all, they treat as a README they update when they remember to. That is the mistake this piece is about, and it is an expensive one, because the file is not documentation for the work. The file is doing the work.

## The edge that does not commoditise

The visible contest in agent engineering is about models and prompts. Which model, which framework, which clever instruction. That contest is real, and it is also the copyable part. Model access and generic prompt techniques are available to anyone; what is hard to copy is the operational knowledge specific to your project, accumulated from your own failures. The edge the visible contest buys is temporary by construction, because everyone can buy the same edge.

The part that does not commoditise is the specification: the accumulated, written-down statement of what the agent is for, how it should behave, what counts as done, and what it must never do. That statement is where your judgment lives once it stops living only in your head. Give two operators the same model and the same task and the outcomes need not be close: one is running it through a specification that has absorbed a hundred hours of hard-won corrections, the other through six lines written once and never revisited.

## The real variable is whether the file learns

It is tempting to say documentation depreciates and specifications appreciate, but that is too clean, because good documentation can compound and a bad specification can rot. The label is not the variable. The variable is whether the file is passive description or executable operational memory.

A file is passive description when it states how things are and nothing depends on it; it decays the way a README decays, the system moving on while the words stay put. It becomes operational memory when failures and decisions flow back into it and future behaviour depends on what is written. Each time your agent hits an edge case, misreads an instruction, or produces confident nonsense, you encode the fix: the failure mode, the constraint that prevents it, the check that catches it next time. Do that repeatedly and the file carries knowledge you could not reconstruct from memory, and each correction hands the next run a constraint it would otherwise have to guess. Whether you call it a spec, a system prompt, a runbook or a policy file is beside the point. What matters is that behaviour depends on it and that what it has learned flows back in.

The condition is maintenance, and it is strict. In a fast-moving project, a file that has not been revisited in a couple of weeks may already be drifting, because the project moved and the file did not. Worse, the agent now optimises against a description of a system that no longer exists, and it does so confidently. This is Goodhart's law pointed at your own instructions: the file becomes the target, and a file describing a system that is gone steers the work toward the version that is gone. A stale specification is not neutral. It works against you.

There is an honest objection here, and the best evidence to date sharpens it rather than softens it. When researchers at ETH Zurich evaluated repository context files across several coding agents and models, adding one did not generally raise task success, and it pushed inference cost up by more than a fifth.[^1] Part of that is the model: a stronger one infers more and carries a short prompt further, and instructions piled on past the point where they help start to bury the signal the agent needs. But the study's finding underneath the headline is the one that matters here. The instructions in a context file were followed and did their job; the repository overviews, the passages that only narrate the system, were what carried no weight. That is the seam this whole piece runs along: a file that narrates decays, a file that instructs compounds. And the study's own conclusion is the note this piece closes on, that any change you make to improve the file has to be measured before you trust it. So the file matters, up to a point set by your model and your work, and the check at the end is how you find that point.

## What you are actually holding

The specification layer is capital: the newest form of the oldest kind of invisible capital there is, the internally built know-how a good operation used to carry only in its people and its habits, the kind that never sat on a balance sheet because its cost cannot be separated from the cost of running the business at all. Others have named the shape of it already: context as a moat, the spec as the new code, [the layer you own as where the edge lives](https://durabilitycurve.com/blog/skills-are-package-management-for-your-ai/), capital that compounds while it is used and decays when it goes stale. The genre says the layer matters. The claim here is the narrower, more useful one: the quality of your particular layer is readable, and reading it tells you whether yours is gaining value or losing it before it costs you.

What is new is not the category. It is that this capital used to be tacit, trapped in people, transferable only by hiring and apprenticeship. Now it is externalised into a file, executable directly by a machine, and operable by one person at a scale that used to need a team. That last claim is a hypothesis, not a measured law, and here is its falsifier: if teams with agents match a solo operator's output per person, then the specification is not what is doing the lifting. That test has not been run at scale. What follows is the mechanism, and how to measure it on your own work.

This also explains why the value is so easy to miss. A specification has high value in use and much lower value in exchange. It compounds for the person who maintains it and reads it every day, and most of what makes it work is the tacit context in that person's head, which does not transfer when the file does. A mature specification bundled with its tests, workflows and data can carry real value to a buyer, so this is a large gap, not a literal zero. But the gap is the point: the asset producing the output sits in a repository nobody outside can read, price, or put on a book. Two forces are pushing value into it at once. The specification persists while the output it generates turns over daily, and as writing code and producing content commoditised, the binding constraint moved up to specifying what to build and how to verify it, which is [where the advantage now sits](https://durabilitycurve.com/blog/leverage-hierarchy-of-agent-engineering/).

## So is yours appreciating or depreciating

You cannot answer that by feel, and feel is where most people stop. A specification you wrote is one you are proud of, which tells you nothing about whether it is aimed at a system that still exists. You have to read it against fixed marks, the way you would read a balance sheet rather than admire a photograph of the building.

Eight marks separate a specification that gains value from one that is quietly losing it. The eighth is the one people forget: a file that only ever accumulates eventually depreciates from its own weight, as old constraints outlive their reason, rules start to conflict, and the important instructions get diluted by the stale ones. Compounding is not the same as hoarding. Score your own as you read.

[Figure: The eight-mark rubric: for each mark, the cold pole a file slips to and the gold target it should reach.]

One way to hold the eight marks together: a file that forgets is bad, a file that learns is good, and a file that hoards has quietly started to forget again, because its best rules are buried under the ones that no longer apply. Read honestly, most specifications score low, and that is the point rather than a failing grade on you. The file was treated as documentation, so it decayed like documentation. The distance between where yours scores and where it could score is value you have not collected yet, and collecting it costs no model upgrade at all.

The eight marks are also a two-page field card, free to every reader, with the eight anchors laid out to tick against your own file. Take it below and score your own file against the eight marks above before you go any further.

https://harryfloyd.substack.com/p/the-specification-self-check

**What's in the rest of this piece:** how to turn the self-check into a score out of sixteen. A weak specification scored, then hardened line for line so you can copy the moves. The rule for what to fix first. And the one test that tells you whether any of it changed the work.

[^1]: Thibaud Gloaguen, Niels Mündler, Mark Müller, Veselin Raychev and Martin Vechev (ETH Zurich, SRI Lab), "Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?" (2026): across several coding agents and models, providing a repository context file "does not generally improve task success rates" and raises inference cost "by over 20% on average"; but the instructions in the file "are well followed," while repository overviews "are not helpful," and any attempt to improve performance "should be rigorously evaluated before deployment." https://arxiv.org/abs/2602.11988

  
Paid subscribers

  
The rest of this piece is for paid subscribers, on any tier.

  
[Read the rest on Substack](https://harryfloyd.substack.com/p/you-own-an-asset-youve-never-scored)
