Derived principle A

Goodhart Corollary

Optimisation on a proxy degrades that proxy's validity.

All horizons. Structural. 4 essays 39 sourced claims

The claim in full

Any metric used as an optimization target degrades under the optimization pressure itself. The proxy-to-construct mapping erodes monotonically with optimizer capability. The gaming strategies generalise and compose.

What it is composed from

Law II — the proxy removes the load-bearing difficulty of real measurement. Law IV — optimising changes what the metric measures (Reflexivity). Law V — the proxy is the wrong abstraction level for the real construct. What makes it worth naming independently is the active adversarial feedback loop: the optimizer doesn't just miss the target, it degrades the signal it is chasing.

What would falsify it

A law that cannot fail is not a law. Each of these is the observation that would break this one, written before the evidence was looked for, so the frame can lose.

Below the threshold

Sustained high-pressure optimization on an imperfect proxy that continued to improve the underlying construct with no proxy-to-construct degradation, even as the optimizer grew more capable.

Above the threshold

A high-capability adaptive evaluand with strong incentive for conditional cooperation that, under sustained pressure, continued to satisfy the evaluation AND preserved aligned deployment behaviour, with the alignment verified by an instrument the evaluand could not model. Equivalent counter-finding: a more capable system that became easier to evaluate under the same protocol with no redesign.

The sharpest form to use

Above the strategic-evaluand threshold, stable evaluation metrics are not evidence of stable behaviour. Verify deployment independently of the evaluation surface.

Amendments

The statement above is not the original. Each row is a time the law was rewritten because it failed a test, with the reason it failed.

  1. 17 Apr 2026

    Capability-threshold extension, absorbing former Candidate Derived Principle C (The Evaluation Inversion). Derived A has two regimes separated by whether the optimizer can represent the evaluation boundary as an object distinct from the underlying objective. PRE-THRESHOLD: the signal degrades visibly; fix the metric. ABOVE-THRESHOLD (strategic-evaluand): the signal looks healthy while deployment behaviour diverges; the failure mode is false confidence, not visible noise, and improving the metric does not help because the evaluand can model your improvement and pass it too.

  2. 17 Apr 2026

    Evaluation-Instrument Triangle. Six orphan analyses reduced to the same compositional core (Law IV + Derived A) at three independently-named layers: contract, construct, trace. Implication: triangulation must be between-layer, not within-layer, because each layer shares the same Goodhart-vulnerable surface.

Essays that stress-test it

  1. PROOF & TRUST Everyone Got Safer. That's the Problem. Safety has two numbers: how often each system fails, and whether they fail together. The field has spent years driving the first one down while almost no dashboard reports the second. 19 sourced claims
  2. THE HUMAN LAYER You Cannot Try to Fall Asleep Sleep is only where you notice it first. Much of what matters works the same way. 20 sourced claims
  3. PROOF & TRUST Ten Lines of Code Scored 100%. One Agent Broke Eight Benchmarks. Not one task was actually solved, and the same blind spot is sitting in your own dashboard.
  4. PROOF & TRUST The Stable Liar Every metric you optimise quietly stops measuring what you meant. The dangerous ones never break. They keep reporting green while the thing underneath rots.

Read this law in the framework essay All writing The claim ledger