Claim ledger
Confidently Wrong
The effort you're handing to AI was doing two hidden jobs. Skip them and you get faster, weaker, and blind to your own mistakes.
11 claims · 10 verified to primary
What the essay claims
A hard task usually does two jobs beneath the obvious one. It checks you: a single route to an answer cannot check itself, so catching an error needs a second, independent route that would fail differently, and the effort you were about to spend was often that route. It builds you: the reps are what keep the skill that lets you notice at all. Offloading to a machine can delete either without your feeling it, and the two failures compound, because asking the machine to check its own work is the same blind spot twice, and leaning on it long enough retrains your own judgement onto its output until your independent check is no longer independent. The rule that follows is not "keep the hard things" and not "automate everything": it is never remove your last independent route to an answer without installing another that can fail differently, and never stop doing the reps that would let you notice when you are wrong. The claim is bounded: not all difficulty is load-bearing (confusion, toil, and effort with no traction build and check nothing and should be shed), and verification is only sometimes the expensive half (it stays cheap where a cheap oracle exists — a passing test, a green type-checker, production that stays up).
The claim ladder
Every rung stands alone for a reader who has read none of the others; each names its subject and its scope.
| Rung | Claim (subject) | Scope / status | Does NOT reach |
|---|---|---|---|
| 1 | The Reinhart–Rogoff 90% result collapsed on recomputation along three faults; corrected growth was +2.2% not −0.1% | VERIFIED to Herndon/Ash/Pollin 2013 | Does not claim debt is harmless, nor settle causation |
| 2 | A single route to an answer cannot detect its own error; catching it needs an independent route that fails differently | Structural claim (triangulation / N-version) | Not that any second step suffices — only a differently-failing one |
| 3 | People stop cross-checking automated output; it persists in experts and resists training (automation complacency) | VERIFIED to Parasuraman & Manzey 2010 | Not a claim that all automation is unsafe |
| 4 | "Independent" versions fail in correlated ways more than chance allows | VERIFIED to Knight & Leveson 1986 | Not that redundancy is useless — the gain is smaller than the independence model claims |
| 5 | Offloading reps erodes the skill needed for the rare critical moment (deskilling) | Supported (Bainbridge 1983; AF447 as illustration, one cause among several) | Not that AF447 was caused by deskilling alone |
| 6 | Verification can cost more than generation for plausible output with subtly load-bearing errors | Class-bounded | Not a general law; cheap where a cheap oracle exists |
What the piece does not reach: it gives no universal rule for which tasks to offload, no proprietary number, and no claim that difficulty is always worth keeping. It hands a question to run, not a verdict.
What would make it wrong
The core claim is falsified if a controlled test shows that offloading a task's difficulty to AI leaves error-detection rates unchanged — i.e. that people who derive an answer independently and people who accept the model's answer catch downstream errors at the same rate. The bounded sub-claim (verification can exceed generation cost) is falsified if, across a representative sample of AI-assisted knowledge work, careful verification is reliably cheaper than production even for plausible-but-wrong output. Market/ reader-observable proxy: if the piece converts and readers report running the "name the independent route" check, the power-transfer claim is supported; if it reads as "smart, changed nothing," it is not.
The evidence, row by row
Each row is a claim the essay makes, the words in the source that support it, where in the source they are, the day the source was read, and the status the check assigned. Verified means the primary was opened and read; Executed means we ran it ourselves; Reported means carried from a source we could not open in full.
-
Reinhart & Rogoff (2010) reported that above 90% debt-to-GDP, average real growth was −0.1%
"median growth rates fall by ... and average (i.e., mean) growth rates fall considerably more" with the above-90% mean at −0.1%
Reinhart & Rogoff, Growth in a Time of Debt (AER P&P) · Table 1 / §II · scholar.harvard.edu
-
The 90% figure was used to justify austerity, cited by Paul Ryan's budget and by the European Commission
"Both Paul Ryan ... and Olli Rehn ... invoked the 90 percent threshold"
Krugman, How the Case for Austerity Has Crumbled (NYRB) · Body · nybooks.com
-
Herndon (a UMass grad student), Ash & Pollin obtained the spreadsheet and found the result rested on three faults
"a combination of coding errors, selective exclusion of available data, and unconventional weighting of summary statistics"
Herndon, Ash & Pollin, PERI Working Paper 322 · Abstract · peri.umass.edu
-
The Excel formula excluded five countries (Australia, Austria, Belgium, Canada, Denmark); corrected average growth above 90% debt was +2.2%
"the average real GDP growth rate for countries carrying a public debt-to-GDP ratio of over 90 percent is actually 2.2 percent, not −0.1 percent"
Herndon, Ash & Pollin, PERI WP 322 · Abstract / §III · peri.umass.edu
-
People stop cross-checking automated output; complacency and bias appear in experts and are not eliminated by training or warnings
"automation-induced complacency ... found in both naïve and expert participants and cannot be overcome with simple practice"; bias "cannot be prevented by training or instructions"
Parasuraman & Manzey, Complacency and Bias in Human Use of Automation, Human Factors 52(3) · Abstract / §Integration · journals.sagepub.com
-
Knight & Leveson (1986) had 27 programmers write the same program from one spec, ran each version on about one million inputs, and found failures were correlated far beyond independence (rejected at the 99% confidence level)
"27 versions of a program ... executed on one million randomly-generated inputs ... the hypothesis ... that the ... versions fail independently ... was rejected at the 99% confidence level"
Knight & Leveson, An Experimental Evaluation of the Assumption of Independence in Multiversion Programming, IEEE TSE SE-12(1) · Abstract / §Results · csc.kth.se
-
Desirable difficulties (spacing, interleaving, generation, self-testing) build durable skill, and are desirable only if the learner can meet them
"If ... the learner does not have the background knowledge or skills to respond to them successfully, they become undesirable difficulties"
Bjork & Bjork, Making Things Hard on Yourself, But in a Good Way (2011) · Body · bjorklab.psych.ucla.edu
-
Bainbridge (1983): automating a task leaves the operator the parts too hard to automate, and less practised at the skills they demand
"the designer who tries to eliminate the operator still leaves the operator to do the tasks which the designer cannot think how to automate"
Bainbridge, Ironies of Automation, Automatica 19(6) · p.775 · sciencedirect.com
-
Capt. Warren VanderBurgh's 1997 American Airlines training talk named "children of the magenta line"; by the airline's own reckoning most of the automation-related trouble studied traced to automation mismanagement
"Children of the Magenta (Line)", AA training, ~April 1997; ~68% of reviewed events traced to automation mismanagement (softened in-body to "most")
VanderBurgh, Children of the Magenta Line (AA training); AirFacts retrospective · Retrospective · airfactsjournal.com
-
AF447 (2009): pitot icing → unreliable airspeed → autopilot disengaged → nose held up → unrecognised stall; BEA cited crew-coordination breakdown and absence of manual-handling training at altitude among causes
"the absence of training, at high altitude, in manual aeroplane handling and in the procedure for unreliable airspeed"; CRM "degraded"
BEA final report on AF447 (2012), via IEEE Spectrum / Aviation Safety · Final report summary · spectrum.ieee.org
-
Verifying a completed solution is generally far cheaper than producing one (the P vs NP intuition; Sudoku as the illustration)
verifying a candidate solution is easy while finding one is hard — the defining property of NP
Clay Mathematics Institute, P vs NP · Problem statement · claymath.org
Cite this
- A single claim
"[claim text]" (Floyd, Harry, 2026, https://durabilitycurve.com/claims/confidently-wrong/)Replace the bracket with the row's claim text. The page URL carries the source; the row's own source link is in the row.- The essay
Floyd, Harry (2026). Confidently Wrong. The Durability Curve. https://durabilitycurve.com/blog/confidently-wrong/- This ledger
Floyd, Harry (2026). Claim Ledger: Confidently Wrong [structured claims with sources]. The Durability Curve. https://durabilitycurve.com/claims/confidently-wrong/
Quote with attribution and a link to this page or the essay. Say if you changed the wording. Not licensed for model training. Plain-text copy for machines: /md/claims/confidently-wrong.md.