You Built a Number You Will Not Trust
Some things you build get graded by the world. Some only ever hand you back your own guess wearing a decimal. Telling the two apart before you start is the cheapest hour you will spend.
How this is checked Law V last tested Sep 2026
- Rests on
- Law V, The Targeting Problem. It has been put to a real test and amended or narrowed under that pressure. Last tested 2 Sep 2026.
- What would prove it wrong
- A system where capability aimed at a target that was a poor proxy for the objective (not a sufficient statistic), with the misfire undetectable or un-re-aimable, nonetheless produced sustained improvement rather than waste or harm. This is the law's test; the essay has no ledger of its own yet.
New here Contents sit in the right margin. Press F to read with nothing else on screen. Save keeps an essay in this browser so it opens offline, and nothing leaves it. The bar at the foot of the screen holds the contents, text size and focus mode. Save keeps an essay in this browser so it opens offline, and nothing leaves it.
There is probably something you are part of the way through building. A spreadsheet that will score your sales leads so you know which to ring first. A scorecard that will rank the people you interviewed. A formula that will tell you which product line to drop. You have the columns and the weights, and it more or less works. And some part of you already knows you will not trust the number it gives you. When it puts the wrong lead at the top, you will override it, the way you did last time.
That thing was built well and it is not going to help you, and the reason is not that you got the weights wrong. You could tune them all weekend. The reason is that you never said what would make its answer right. You built a machine to rank the leads before you had settled what a good ranking even was, so when its number surprises you, you have no honest reason to trust the surprise over your own read, and you override it. Working that out before you start is worth more than any amount of tuning, because it changes what you build.
What “it works” does not tell you
When a thing you made runs, you feel finished. The columns add up, the report generates, the score comes out. But working only tells you the thing does what you built it to do. It says nothing about the two questions that can kill the thing before you waste a weekend on it: whether you could ever tell that its answer was wrong, and whether it beats what you would have done without it. You can ask both before you write a line, and most of the cost of a wasted build is the cost of not asking.
The one people skip: what would tell you it was wrong
Start with whether you could ever tell the answer was wrong, because that is the question that gets skipped, and it is easy to hear it as a question about whether the tool runs. Those are not the same: a tool can run perfectly and still hand you an answer you could never catch out. A total, a count, a due date, an amount owed: each has an obvious check, because there is an answer outside the tool to compare it with. Those you can build and know you built them well.
Past those, it comes down to one thing that is easy to miss: whether the outcome that grades the answer arrives on its own, or whether acting on the answer is what decides which outcome you ever see. A forecast can be the first kind, as long as the thing you are predicting arrives regardless of what you do with the prediction. Predict how many people will show up or how many inbound calls will arrive, and the real number turns up whatever you guessed. Footfall, call volume, next week’s demand: the world returns a verdict on these whether you like it or not, which makes them possible to test and improve even when the answers are hard. You always get to find out.
The lead score looks like one of those and is not. The outcome you get to see is the one the score chose for you. You ring the leads at the top, some of them buy, and the ranking looks confirmed, but the leads it sent to the bottom you never ring, so they never get the chance to prove it wrong. It can bury half your real buyers and still show you a good week, because the ones it buried are the ones you will never check. The number decides which evidence you ever see, and the mistakes that matter are the ones it never shows you. That is the dangerous kind of number: the one whose very use hides its own mistakes, so the more you lean on it the less you can see them.
A tool like that leaves you two choices, and neither is simply to trust the score. You can keep the decision for yourself and build the thing that lays the evidence out, the renewal date, the value, the last time you spoke, so the call takes seconds and the decision stays where it can actually be made. Or, if you want the score, you keep a corner of the world honest: at random, ring some of the leads it tells you to skip, so it cannot set its own exam and then pass it. What you cannot do is act on it everywhere and call the result proof.

So the question that matters is not whether the answer is a fact or a judgement. Plenty of recommendations can be tested against what actually happens, over enough cases. The question is whether you would ever find out this one was wrong, for the options it turns down as much as the ones it takes. When nothing would, you have automated the decision without earning any reason to trust it.
I have watched myself skip this question, so I know why it gets skipped. Asking is quick and uncomfortable, because the honest answer is sometimes that nothing would ever tell you, and that ends the project in the first ten minutes. Building is slow and feels like progress, and every hour in makes the thing harder to walk away from. So the quick question loses to the comfortable work, and you find out at the end.
What it cost me to learn this
I learned this by paying for it twice in one evening. I had built a small tool that read a page of written instructions, the kind you might write to hand a job over to someone, and sorted each line into two piles: the lines that told you to do something, and the lines that were only describing the setup. It ran. Then I tested it against a page it had never seen, marking the lines myself first, and on the ones it was willing to call, it agreed with me a little over half the time. The tempting lesson is that the tool was bad. The real lesson was worse. I had never pinned down what the right pile was, and I could have: written down the rule for what counts as an instruction, had someone else label the same page, checked whether we agreed. That would have given me something to hold the tool to. I never did it, so the tool’s score could tell me how often it matched me, but not whether it had learned a rule worth trusting.
The target was buildable. I had just built everything except the target.
So I rebuilt it as the other kind of tool, one that only counts, showing each section’s share of the page and letting me tick what to cut. That has an obvious check, and I built it properly, with a list of thirty-four ways it could go wrong, twenty-seven test cases worked by hand, seven real pages it passed cleanly. It was correct. It was also useless, because the thing it replaced was opening the file and reading it, and it beat that by nothing. The first question had sent me to a tool that could at least be right. The second question killed it anyway. That second failure is the easier one to miss, because a tool that is correct feels like a tool that is worth having. It can be measurable, tested, and right, and still lose to the thing you would have done without it. What you would do instead belongs on the page next to the answer you are after.
Later I wrote out what could have sunk each build. The thing that actually sank the first one was not among the bugs I had spent the night fixing. It had been sitting above them from the start: I never said what a right answer would be.
Run it on the thing you are building
So take the thing you are building now. Before you add another column, write down the answer it is meant to give, and put it to the two questions.
First, if that answer came out wrong, would anything ever tell you? If the outcome arrives independently of what you do with the answer, build it and keep score. If acting on the answer is what picks the evidence you see, randomly test some of the choices it would otherwise turn down, or leave the decision to yourself and build the tool that lays the evidence out.
Second, say in one plain sentence what you would do instead if this did not exist, and why what you are building beats it. If that sentence will not come, you have your answer, and it is cheaper now than after the weekend.
The tell that you skipped both: you are deep in the build and neither answer is written down anywhere.

What the questions do not do
Two honest limits. The first is that even a target you can check can be the wrong target. Rank the leads by who converts and the tool may learn to love the small easy accounts and starve the big awkward ones, and it will be hitting the number you set the entire time. You can define success exactly and still have defined the wrong success. These two questions do not protect you from that. What they buy you is the right to find out you were wrong, not a promise you aimed at the right thing.
The second is that they do not make you wise in advance. Both times, I killed the tool after I had built it, not before. Asking first lowers how often you pay, but it will not take the count to zero, because now and then the only way to learn what a thing is worth is to build it and look. The questions are there to catch the builds you start because building feels easier than asking, not to promise that you will never lose another evening.
There is one last thing the questions never reach, and it was always going to be yours. When you have a target you can genuinely test, software can hand you a better answer than you would reach alone, and you should let it. What it cannot do is settle what counts as a good answer, settle how much evidence is enough, or answer for the choice to act on it. A machine can be right. Being the one who is answerable for trusting it is the part that stays with you.