Derived principle A

Goodhart Corollary

Optimising an imperfect proxy degrades the construct it stands for — non-monotonically — unless the proxy is a sufficient statistic for it (which you usually cannot know in advance).

All horizons. Structural. 4 essays 39 sourced claims

The claim in full

Any metric used as an optimization target degrades the construct it stands for under sustained optimization pressure — UNLESS the proxy is a sufficient statistic for the construct, in which case optimizing the proxy improves it (a condition rarely total and usually not knowable in advance). AMENDED 2026-09-02: the degradation is NON-MONOTONIC — the construct typically tracks the proxy early and decouples only past an over-optimization point (an inverted-U), and its severity falls as proxy quality rises; the prior "erodes monotonically with optimizer capability" is WITHDRAWN (the trajectory is non-monotone, and the coefficient that actually scales is proxy quality, which reduces corruption). The gaming strategies generalise and compose.

What it is composed from

Law II — the proxy removes the load-bearing difficulty of real measurement. Law IV — optimising changes what the metric measures (Reflexivity). Law V — the proxy is the wrong abstraction level for the real construct. What makes it worth naming independently is the active adversarial feedback loop: the optimizer doesn't just miss the target, it degrades the signal it is chasing.

What would falsify it

A law that cannot fail is not a law. Each of these is the observation that would break this one, written before the evidence was looked for, so the frame can lose.

Below the threshold

Sustained high-pressure optimization on an imperfect proxy that is NOT a sufficient statistic for the construct — one where gaming the proxy without helping the construct is both possible and accessible to the optimizer — which nonetheless improved (or failed to degrade) the construct with no decoupling, even under sustained pressure. Supersedes the prior falsifier ("continued to improve the underlying construct with no degradation, even as the optimizer grew more capable"), which was met — contested — by grokking/double descent and ImageNet-accuracy→transfer (both sufficient-statistic or near-sufficient-statistic proxies) and could not distinguish a genuinely corruptible proxy from one that simply tracks its construct.

Above the threshold

A high-capability adaptive evaluand with strong incentive for conditional cooperation that, under sustained pressure, continued to satisfy the evaluation AND preserved aligned deployment behaviour, with the alignment verified by an instrument the evaluand could not model. Equivalent counter-finding: a more capable system that became easier to evaluate under the same protocol with no redesign.

The sharpest form to use

Above the strategic-evaluand threshold, stable evaluation metrics are not evidence of stable behaviour. Verify deployment independently of the evaluation surface.

Amendments

The statement above is not the original. Each row is a time the law was rewritten because it failed a test, with the reason it failed.

  1. 17 Apr 2026

    Capability-threshold extension, absorbing former Candidate Derived Principle C (The Evaluation Inversion). Derived A has two regimes separated by whether the optimizer can represent the evaluation boundary as an object distinct from the underlying objective. PRE-THRESHOLD: the signal degrades visibly; fix the metric. ABOVE-THRESHOLD (strategic-evaluand): the signal looks healthy while deployment behaviour diverges; the failure mode is false confidence, not visible noise, and improving the metric does not help because the evaluand can model your improvement and pass it too.

  2. 17 Apr 2026

    Evaluation-Instrument Triangle. Six orphan analyses reduced to the same compositional core (Law IV + Derived A) at three independently-named layers: contract, construct, trace. Implication: triangulation must be between-layer, not within-layer, because each layer shares the same Goodhart-vulnerable surface.

  3. 2 Sept 2026

    Non-monotonicity fix + sufficient-statistic scope clause, from a pre-registered cross-domain probe, 5 cases, all primary/strong-secondary grounded. TWO modifiers fell; the CORE held. (1) "erodes monotonically with optimizer capability" was FALSE of the optimization TRAJECTORY and named the WRONG scaling axis. The gold-vs-proxy curve is an inverted-U (Gao et al. 2023, ICML; best-of-n d(α−βd), RL d(α−β·log d): "gold reward initially increases … eventually peaks and declines"); grokking and double descent show the construct improving late. The coefficient that actually scales is REWARD-MODEL (proxy) size, and "larger reward models are more robust to over-optimization" — a better proxy corrupts LESS — so severity falls as proxy quality rises, the OPPOSITE of the withdrawn wording. The optimizer-STRENGTH direction stays a vault-internal claim (Synthesis — Reward Hacking) NOT re-grounded by this probe. (2) The universal "any metric degrades" is falsified, CONTESTED, by ImageNet-accuracy→transfer (Kornblith CVPR 2019, r=0.99/0.96; Recht ICML 2019, no adaptive overfitting) — a near-sufficient-statistic proxy whose optimization improved the construct. Contested by a decoupling subspace (regularization tricks that move transfer but not accuracy), so the exception's primitive is UNDER-DETERMINED (statistical sufficiency vs the optimizer lacking incentive to trade construct for proxy); this probe does not disentangle them, so the convergence with Law V's sufficient-statistic facet is looser than a clean match and stays UNCONFIRMED (Derived A composes II+IV+V, so it cannot independently confirm it — the tie-breaker is a clean Derived B probe). CORE (a non-sufficient-statistic imperfect proxy under sustained strong optimization eventually decouples/harms the construct) confirmed on Campbell's-Law test-score inflation, cardiac report cards (Dranove et al. JPE 2003), and reward over-optimization. Blocking checks passed: (a) contradicts no historical supporter and is corroborated by two (Reward Hacking's capability-monotonicity is the kept reading; the Evaluation Inversion had already flagged its "monotonicity evidence is directional, not proven"); (b) the non-monotonic/proxy-quality part is usable ex ante, the sufficient-statistic exception is not knowable in advance and is demoted to an explanatory limit, exactly as Law V facet (b). An operator-gated adversarial correctness review BEFORE applying caught and fixed two errors in an earlier draft of this amendment (the wrong scaling-axis citation and an over-tight Law V convergence). STANDING: Derived A's prior load-bearing rested on the 2026-04-17 ABSORPTION of Candidate C (an extension that made the principle bigger, never a narrowing); this is the FIRST falsifier to bite and narrow it, so its load-bearing is now EARNED UNDER FALSIFICATION rather than by absorption.

Essays that stress-test it

  1. PROOF & TRUST Everyone Got Safer. That's the Problem. Safety has two numbers: how often each system fails, and whether they fail together. The field has spent years driving the first one down while almost no dashboard reports the second. 19 sourced claims
  2. THE HUMAN LAYER You Cannot Try to Fall Asleep Sleep is only where you notice it first. Much of what matters works the same way. 20 sourced claims
  3. PROOF & TRUST Ten Lines of Code Scored 100%. One Agent Broke Eight Benchmarks. Not one task was actually solved, and the same blind spot is sitting in your own dashboard.
  4. PROOF & TRUST The Stable Liar Every metric you optimise quietly stops measuring what you meant. The dangerous ones never break. They keep reporting green while the thing underneath rots.

Read this law in the framework essay All writing The claim ledger