Claim ledger

Your AI Grader Is Only as Good as Its Answer Key

I went through 77 studies of AI graders. Without a verified answer key, they are least reliable on exactly the questions they can't answer themselves. Here is a twenty-item check for yours.

Read the essay On Substack Published

Claims
68
Verified at the primary source
57
Removed before publication
0
Last checked

68 claims: 57 verified, 11 executed.

Download this ledger: CSV · JSON · Markdown · CC BY 4.0

What the essay claims

On correctness, the variable that decides whether an AI judge can be trusted is not which model grades but whether it has a verified reference to grade against. Without one, judges agree with experts mostly on the questions they could answer themselves, and fall apart on the rest; with one, even small models grade close to how well humans agree with each other. When judges are wrong about correctness, they mostly pass flawed work (12 of 19 studies that measure direction; code review is the counter-case). On matters of taste and expert judgement, where no key exists, judges more often fall short of expert–expert agreement than match it. The popular fixes (panels, reasoning first, fine-tuned judges, calibration on a few human labels) help less than claimed. The test that screens it for any one setup is the reader's own: twenty items, with and without a key, split by whether the judge could solve them.

The claim ladder

RungClaimEvidence (ledger)Whose behaviourWhat it does NOT reach
R1With a verified key, judges grade hard correctness close to human level; a self-made or subtly wrong key failsL1–L5Judges on expert-graded finance/maths (GPT-4o + 4 open models); 2026 frontier judges on a proxy labelSubjective grading (no key exists); 2026 rows are corroboration, not human-labelled
R2Without a key, judges agree with experts mostly on questions they could answer themselvesL1, L5SameWhether the newest judges (Opus 5.5, GPT-6) still show it: untested in the literature
R3On taste and expert judgement, judges more often fall short of expert–expert agreementL7–L12Chat preference, science answers, clinical gradingMixed: HealthBench 5/7 themes at or above the average physician (vendor-authored); GRAND-ROUNDS 3/5
R4When wrong about correctness, judges mostly pass flawed workL6, L13 + direction sweep (12/19)Judges grading answers and proofsReverses in code review against a spec; not a law
R5Biases are judge-specificL14–L16Named 2025–26 judges on MT-Bench / JudgeBench / style pairsHuman baseline for markdown is n = 2 annotators
R6Self-preference is contestedL17–L18Judges grading their own vs others' answersUnresolved between the two papers; say so
R7Popular fixes help less than claimedL19–L21Panels, reasoning-first, human-label calibrationPanels can edge their best member (L19); some reasoning helps on hard pairs
R8The Judge Check can catch a generous grader on twenty items; it cannot clear one (about 100 per failure mode, L24a)Object Record (our run, n = 24)The check, on Haiku 4.5 and Sonnet 5 grading Haiku 4.5A demonstration, not evidence; one run, one task type

The evidence, row by row

Each row is a claim the essay makes, then the evidence recorded for it and the source it was checked against. For verified rows that is usually the exact words, where they sit and the day they were read. The label says how far it was checked.

What the labels mean
Verified
Checked against the primary source itself by a checker who did not draft the piece, usually with the exact words, where they sit and the date read.
Executed
A run or observation done for the piece: code, a query, a count, or a check made live in an app. The claim is what it returned. Run records are not published.
Checked
Checked against its source by a separate checker in the older ledger format, which recorded no quote, locator or date.
Reported
Not verified word for word against a primary source: carried from a secondary source, from one that could not be opened in full, or supported only in part. The essay words it accordingly.
Struck
Drafted, checked and removed before publication. Kept here, crossed out, with the reason.
Excluded
Considered and deliberately left out of the piece. Kept here with the reason.
Sources (21)
  1. Krumdick et al. 15 rows
  2. Chandak et al. 5 rows
  3. Zheng et al. 4 rows
  4. Kim et al. 1 row
  5. Calderon 1 row
  6. Chiang et al. 1 row
  7. Zhao et al. 3 rows
  8. Buckley et al. 4 rows
  9. Arora et al. 1 row
  10. Gonzalez et al. 4 rows
  11. Norman 4 rows
  12. Soumik 3 rows
  13. Roytburg et al. 1 row
  14. Zhang et al. 1 row
  15. Verga et al. 1 row
  16. Dorner 1 row
  17. Tan et al. 2 rows
  18. Eugene Yan 1 row
  19. Hamel Husain 2 rows
  20. Jin & Chen 2 rows
  21. First-party runs 11 rows
  1. L1 Verified Read

    Without a correct reference, judges agreed well with human experts only on questions they had been able to answer themselves

    "when not provided with a correct reference, judges show high agreement with human experts only on questions the judges were able to correctly answer themselves"

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, abstract · page-text offset 1.2% · arxiv.org

    Link
    Cite

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, abstract, page-text offset 1.2%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L1, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L1 (source read 2026-09-26).

    Report an error
  2. L1a Verified Read

    GPT-4o as judge, Cohen's kappa with experts. Pairwise: questions it answered correctly 0.78 with no key, 0.92 with a human key; questions it got wrong 0.30 with no key, 0.16 with its own answer as the key, 0.83 with a human key. Single grading: questions it answered correctly 0.46 with no key and 0.59 with a human key; questions it got wrong 0.16 and 0.81

    "GPT-4o Correct 0.78±0.02 0.86±0.02 0.92±0.02 0.46±0.07 0.52±0.06 0.59±0.06 Incorrect 0.30±0.09 0.16±0.08 0.83±0.06 0.16±0.14 0.13±0.14 0.81±0.13"

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Table 4 (GPT-4o row) · page-text offset 21.0% · arxiv.org

    Link
    Cite

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Table 4 (GPT-4o row), page-text offset 21.0%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L1a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L1a (source read 2026-09-26).

    Report an error
  3. L1b Verified Read

    Pairwise, questions each judge got wrong, no key vs human key: Llama 3.3 70B 0.39 vs 0.92; Phi 4 -0.08 vs 0.75; Qwen 2.5 7B 0.21 vs 0.63; Yi 1.5 34B 0.13 vs 0.56 (own answer as key: 0.35, -0.08, 0.14, 0.02)

    "Llama 3.3 70b Correct 0.79±0.03 0.86±0.03 0.93±0.02 0.61±0.08 0.54±0.08 0.73±0.07 Incorrect 0.39±0.06 0.35±0.07 0.92±0.03 0.13±0.06 0.23±0.10 0.69±0.09 Phi 4 Correct 0.67±0.07 0.88±0.04 0.91±0.04 0.48±0.09 0.47±0.07 0.67±0.07 Incorrect -0.08±0.09 -0.08±0.10 0.75±0.06 0.03±0.05 0.02±0.12 0.56±0.12 Qwen 2.5 7B Correct 0.65±0.05 0.69±0.04 0.78±0.04 0.31±0.11 0.37±0.09 0.54±0.09 Incorrect 0.21±0.07 0.14±0.07 0.63±0.06 0.08±0.07 0.21±0.10 0.56±0.08 Yi 1.5 34B Correct 0.31±0.12 0.53±0.12 0.77±0.09 0.23±0.11 0.29±0.08 0.43±0.09 Incorrect 0.13±0.09 0.02±0.09 0.56±0.08"

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Table 4 (other judges) · page-text offset 21.2% · arxiv.org

    Link
    Cite

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Table 4 (other judges), page-text offset 21.2%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L1b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L1b (source read 2026-09-26).

    Report an error
  4. L1c Verified Read

    The paper notes that a relatively small model (Qwen 2.5 7B) with a human reference can grade better than a larger model (GPT-4o) without one, on the same set of responses

    "providing a human reference to a relatively small model (such as Qwen 2.5 7B ) can yield better judgments than using a larger model without human references (such asGPT-4o) on the same set of responses"

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6 · page-text offset 22.2% · arxiv.org

    Link
    Cite

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6, page-text offset 22.2%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L1c, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L1c (source read 2026-09-26).

    Report an error
  5. L2 Verified Read

    With its own answer as the key, GPT-4o's pairwise kappa was 0.86 on questions it had answered correctly and 0.16 on questions it could not answer

    "decreases from 0.86 to 0.16 in the pairwise case on questions GPT-4o could not answer"

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6 · page-text offset 23.0% · arxiv.org

    Link
    Cite

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6, page-text offset 23.0%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L2, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L2 (source read 2026-09-26).

    Report an error
  6. L3 Verified Read

    Averaged across all five graders on questions GPT-4o answered correctly (Table 3 does not say whether single and pairwise grading are pooled): kappa with a human key 0.69, a verified GPT-4o answer as key 0.61, a completely unrelated key 0.50, no key 0.46, a subtly wrong key 0.21

    "Human 0.69 ±0.06 0.74±0.06 0.57±0.08 GPT-4o (✓) 0.61 ±0.07 0.66±0.07 0.52±0.09 None 0.46 ±0.08 0.56±0.08 0.22±0.09 Random 0.50 ±0.08 0.47±0.08 0.09±0.1 Wrong 0.21 ±0.06"

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Table 3 · page-text offset 19.3% · arxiv.org

    Link
    Cite

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Table 3, page-text offset 19.3%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L3, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3 (source read 2026-09-26).

    Report an error
  7. L3c Verified Read

    Table 3's figures are across all five graders

    "Table 3 presents results for these reference types across all five judges"

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6 · page-text offset 26.3% · arxiv.org

    Link
    Cite

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6, page-text offset 26.3%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L3c, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3c (source read 2026-09-26).

    Report an error
  8. L3a Verified Read

    The paper says a slightly incorrect reference can in some cases be worse than no reference at all

    "in some cases be worse than"

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6 · page-text offset 25.4% · arxiv.org

    Link
    Cite

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6, page-text offset 25.4%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L3a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3a (source read 2026-09-26).

    Report an error
  9. L3d Verified Read

    The authors describe the Random key as a completely unrelated reference, and say a slightly incorrect reference can in some cases be worse than a completely unrelated one or none

    "providing a slightly incorrect reference can in some cases be worse than providing a completely unrelated reference or no reference at all"

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6 · page-text offset 25.4% · arxiv.org

    Link
    Cite

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6, page-text offset 25.4%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L3d, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3d (source read 2026-09-26).

    Report an error
  10. L3b Verified Read

    The authors conclude that verifying a model-generated answer, rather than writing and checking one by hand, was sufficient as a key in their setting

    "Verifying responses generated by a model rather than manually writing and checking answers via annotators is sufficient in this setting"

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §7 · page-text offset 28.6% · arxiv.org

    Link
    Cite

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §7, page-text offset 28.6%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L3b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3b (source read 2026-09-26).

    Report an error
  11. L4 Verified Read

    2026 judges, questions each got wrong, against a stand-in label: single grading no key vs human key, Opus 4.7 0.26 vs 0.62, GPT-5.4 0.32 vs 0.67, Gemini 3.1 0.44 vs 0.71; pairwise 0.33 vs 0.66, 0.46 vs 0.87, 0.67 vs 0.95 (own answer as key, pairwise: 0.33, 0.45, 0.62)

    "Opus 4.7 Correct 0.86±0.04 0.90±0.04 0.91±0.04 0.63±0.06 0.69±0.06 0.77±0.05 Incorrect 0.33±0.16 0.33±0.15 0.66±0.13 0.26±0.14 0.27±0.12 0.62±0.11 GPT-5.4 Correct 0.75±0.06 0.87±0.04 0.89±0.04 0.51±0.06 0.64±0.06 0.66±0.06 Incorrect 0.46±0.11 0.45±0.13 0.87±0.08 0.32±0.11 0.25±0.12 0.67±0.09 Gemini 3.1 Correct 0.78±0.06 0.80±0.06 0.83±0.05 0.55±0.06 0.62±0.06 0.69±0.06 Incorrect 0.67±0.13 0.62±0.14 0.95±0.07 0.44±0.13 0.35±0.14 0.71±0.10"

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, Table 12 · page-text offset 90.1% · arxiv.org

    Link
    Cite

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, Table 12, page-text offset 90.1%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L4, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4 (source read 2026-09-26).

    Report an error
  12. L4a Verified Read

    The authors treat the 2026-judge rows, scored against a stand-in label, as corroboration

    "we regard these results as corroboration rather than equivalent to our main table"

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J · page-text offset 88.7% · arxiv.org

    Link
    Cite

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, page-text offset 88.7%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L4a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4a (source read 2026-09-26).

    Report an error
  13. L4b Verified Read

    The stand-in label is GPT-4o grading with the human-verified references

    "We used the GPT-4o Judge with the human-verified references as a proxy correctness label"

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J · page-text offset 88.4% · arxiv.org

    Link
    Cite

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, page-text offset 88.4%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L4b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4b (source read 2026-09-26).

    Report an error
  14. L4c Verified Read

    The authors report that higher reasoning effort can narrow the no-reference gap, for example GPT-5.4 at high effort

    "higher reasoning effort can narrow the no-reference gap (e.g.,GPT-5.4 at high effort)"

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J · page-text offset 88.9% · arxiv.org

    Link
    Cite

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, page-text offset 88.9%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L4c, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4c (source read 2026-09-26).

    Report an error
  15. L4d Verified Read

    GPT-5.4 at high reasoning effort, on questions it got wrong, against the stand-in label: pairwise 0.68 with no key and 0.84 with the human key; single grading 0.38 and 0.67

    "GPT-5.4 low Correct 0.85±0.05 0.90±0.04 0.54±0.06 0.57±0.05 Incorrect 0.55±0.13 0.95±0.06 0.34±0.11 0.56±0.09 default Correct 0.75±0.06 0.89±0.04 0.51±0.06 0.66±0.06 Incorrect 0.46±0.11 0.87±0.08 0.32±0.11 0.67±0.09 high Correct 0.80±0.06 0.87±0.05 0.68±0.08 0.67±0.06 Incorrect 0.68±0.14 0.84±0.11 0.38±0.19 0.67±0.11"

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, Table 13 (GPT-5.4 rows) · page-text offset 91.4% · arxiv.org

    Link
    Cite

    Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, Table 13 (GPT-5.4 rows), page-text offset 91.4%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L4d, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4d (source read 2026-09-26).

    Report an error
  16. L5 Verified Read

    On MATH, DeepSeek V3 as a judge without a reference reached Scott's pi 0.72 against ground truth; as a matcher with the reference, 0.98

    "DeepSeek v3 model achieves only modest agreement π = 0.72, while as a matcher, it achieves π = 0.98"

    Chandak et al., Answer Matching Outperforms Multiple Choice (preprint), arXiv 2507.02856v1, §3.1 · page-text offset 22.6% · arxiv.org

    Link
    Cite

    Chandak et al., Answer Matching Outperforms Multiple Choice (preprint), arXiv 2507.02856v1, §3.1, page-text offset 22.6%, https://arxiv.org/pdf/2507.02856v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L5, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L5 (source read 2026-09-26).

    Report an error
  17. L5a Verified Read

    A 1.7-billion-parameter Qwen3 model matching answers to the reference reached pi 0.97 against ground truth on MATH

    "answer matching, even with the 1.7 billion parameter Qwen3 model (non-thinking mode), achieves near-perfect alignment with the ground-truth (π = 0.97)"

    Chandak et al., Answer Matching Outperforms Multiple Choice (preprint), arXiv 2507.02856v1, §3.1 · page-text offset 22.4% · arxiv.org

    Link
    Cite

    Chandak et al., Answer Matching Outperforms Multiple Choice (preprint), arXiv 2507.02856v1, §3.1, page-text offset 22.4%, https://arxiv.org/pdf/2507.02856v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L5a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L5a (source read 2026-09-26).

    Report an error
  18. L5b Verified Read

    Answer matching with recent models, even small ones, reached near-perfect agreement with human grading, in the range of inter-annotator agreement

    "matching using recent models-even small ones-achieves near-perfect agreement, in the range of inter-annotator agreement"

    Chandak et al., Answer Matching (preprint), arXiv 2507.02856v1, abstract · page-text offset 1.3% · arxiv.org

    Link
    Cite

    Chandak et al., Answer Matching (preprint), arXiv 2507.02856v1, abstract, page-text offset 1.3%, https://arxiv.org/pdf/2507.02856v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L5b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L5b (source read 2026-09-26).

    Report an error
  19. L5c Verified Read

    On MATH, the 'true grades' come from MATH-Verify, a rule-based checker that compares each answer with the reference

    "MATH-Verify library (Kydlicek et al., 2025) implements rule-based ground-truth evaluations of generative responses"

    Chandak et al., Answer Matching (preprint), arXiv 2507.02856v1, §3.1 · page-text offset 22.0% · arxiv.org

    Link
    Cite

    Chandak et al., Answer Matching (preprint), arXiv 2507.02856v1, §3.1, page-text offset 22.0%, https://arxiv.org/pdf/2507.02856v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L5c, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L5c (source read 2026-09-26).

    Report an error
  20. L6 Verified Read

    For frontier judges (DeepSeek V3, o4-mini), 80%+ of grading errors were false passes: the judge marked correct what human annotation marked incorrect

    "errors disproportionately (80%+) arise from false positives"

    Chandak et al., arXiv 2507.02856v1, §3.2 · page-text offset 28.5% · arxiv.org

    Link
    Cite

    Chandak et al., arXiv 2507.02856v1, §3.2, page-text offset 28.5%, https://arxiv.org/pdf/2507.02856v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L6, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L6 (source read 2026-09-26).

    Report an error
  21. L7 Verified Read

    On MT-Bench, without ties, GPT-4 agreed with expert labellers 85% of the time against 81% between humans

    "The agreement under setup S2 (w/o tie) between GPT-4 and humans reaches 85%, which is even higher than the agreement among humans (81%)."

    Zheng et al., Judging LLM-as-a-Judge (NeurIPS 2023 D&B), arXiv 2306.05685v4, §4.2 + Table 5 · page-text offset 30.7% · arxiv.org

    Link
    Cite

    Zheng et al., Judging LLM-as-a-Judge (NeurIPS 2023 D&B), arXiv 2306.05685v4, §4.2 + Table 5, page-text offset 30.7%, https://arxiv.org/pdf/2306.05685v4. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L7, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L7 (source read 2026-09-26).

    Report an error
  22. L7a Verified Read

    First turn, with ties counted (setup S1): GPT-4 grading pairs agreed with humans 66%, GPT-4 grading single answers 60%, humans with each other 63%; without ties (S2) both GPT-4 modes 85% against 81%

    "G4-Pair 70% 1138 66% 1343 97% 662 85% 859 G4-Single - 60% 1280 - 85% 739 Human - 63% 721 - 81% 479 (a) First Turn"

    Zheng et al. (NeurIPS 2023 D&B), arXiv 2306.05685v4, Table 5(a) · page-text offset 34.1% · arxiv.org

    Link
    Cite

    Zheng et al. (NeurIPS 2023 D&B), arXiv 2306.05685v4, Table 5(a), page-text offset 34.1%, https://arxiv.org/pdf/2306.05685v4. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L7a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L7a (source read 2026-09-26).

    Report an error
  23. L7b Verified Read

    In 2023 the MT-Bench authors proposed a reference-guided grader that first answers the question itself and uses its own answer as the reference; on their maths questions it cut the failure rate from 70% to 15%

    "we first generate LLM judge's answer independently, and then display it as a reference answer in the judge prompt. In Table 4, we see a significant improvement in failure rate (from 70% to 15%) over the default prompt"

    Zheng et al. (NeurIPS 2023 D&B), arXiv 2306.05685v4, §3.4 · page-text offset 27.1% · arxiv.org

    Link
    Cite

    Zheng et al. (NeurIPS 2023 D&B), arXiv 2306.05685v4, §3.4, page-text offset 27.1%, https://arxiv.org/pdf/2306.05685v4. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L7b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L7b (source read 2026-09-26).

    Report an error
  24. L7c Verified Read

    The MT-Bench authors saw GPT-4 misjudge an answer to a maths problem it could solve when asked separately: it was misled by the answers it was shown

    "although GPT-4 can solve the problem (when asked separately), it was misled by the provided answers, ultimately resulting in incorrect judgment"

    Zheng et al. (NeurIPS 2023 D&B), arXiv 2306.05685v4, §3.3 · page-text offset 24.1% · arxiv.org

    Link
    Cite

    Zheng et al. (NeurIPS 2023 D&B), arXiv 2306.05685v4, §3.3, page-text offset 24.1%, https://arxiv.org/pdf/2306.05685v4. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L7c, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L7c (source read 2026-09-26).

    Report an error
  25. L26 Verified Read

    In 2023 the Prometheus authors found that removing the reference answer hurt their grader most, and said a reference relieves the grader of solving the question itself

    "relieves the need for the evaluator LM to internally solve the instruction and only focus on assessing the response"

    Kim et al., Prometheus (ICLR 2024, per the arXiv comments field), arXiv 2310.08491 · page-text offset 45.7% · arxiv.org

    Link
    Cite

    Kim et al., Prometheus (ICLR 2024, per the arXiv comments field), arXiv 2310.08491, page-text offset 45.7%, https://arxiv.org/pdf/2310.08491. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L26, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L26 (source read 2026-09-26).

    Report an error
  26. L8 Verified Read

    Re-tested on expert labels with a leave-one-annotator-out test, no LLM grader passed on MT-Bench, one of two datasets where none did

    "in two datasets (MT-Bench, and SummEval), none of the LLMs pass the test"

    Calderon, Reichart, Dror, The Alternative Annotator Test (ACL 2025), arXiv 2501.10970v4, §5 · page-text offset 21.8% · arxiv.org

    Link
    Cite

    Calderon, Reichart, Dror, The Alternative Annotator Test (ACL 2025), arXiv 2501.10970v4, §5, page-text offset 21.8%, https://arxiv.org/pdf/2501.10970v4. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L8, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L8 (source read 2026-09-26).

    Report an error
  27. L9 Verified Read

    In Chatbot Arena's expert relabel, GPT-4 agreed with the two experts 81.0% and 78.5% (Llama-2-13b battles) and 76.3% and 79.3% (GPT-3.5-Turbo battles); the experts agreed with each other 89.8% and 79.4%

    "Llama-2-13b Expert 1 Expert 2 GPT-4 Crowd 72.8% 77.8% 75.6% Expert 1 - 89.8% 81.0% Expert 2 - - 78.5% GPT-3.5-Turbo Expert 1 Expert 2 GPT-4 Crowd 73.8% 83.1% 75.6% Expert 1 - 79.4% 76.3% Expert 2 - - 79.3%"

    Chiang et al., Chatbot Arena (ICML 2024), arXiv 2403.04132v1, Table 3 · page-text offset 35.9% · arxiv.org

    Link
    Cite

    Chiang et al., Chatbot Arena (ICML 2024), arXiv 2403.04132v1, Table 3, page-text offset 35.9%, https://arxiv.org/pdf/2403.04132v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L9, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L9 (source read 2026-09-26).

    Report an error
  28. L10 Verified Read

    On literature-grounded science answers, the best judge, o3, reached 65.1% accuracy

    "Even the best-performing model, o3, achieves only 65.1% accuracy."

    Zhao et al., SciArena (NeurIPS 2025 D&B), arXiv 2507.01001v2, §6.2 + Table 1 · page-text offset 32.6% · arxiv.org

    Link
    Cite

    Zhao et al., SciArena (NeurIPS 2025 D&B), arXiv 2507.01001v2, §6.2 + Table 1, page-text offset 32.6%, https://arxiv.org/pdf/2507.01001v2. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L10, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L10 (source read 2026-09-26).

    Report an error
  29. L10a Verified Read

    Random guessing on the same task scores 50.0

    "Random Guess 50.0 o3 65.1"

    Zhao et al., SciArena (NeurIPS 2025 D&B), arXiv 2507.01001v2, Table 3 · page-text offset 32.3% · arxiv.org

    Link
    Cite

    Zhao et al., SciArena (NeurIPS 2025 D&B), arXiv 2507.01001v2, Table 3, page-text offset 32.3%, https://arxiv.org/pdf/2507.01001v2. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L10a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L10a (source read 2026-09-26).

    Report an error
  30. L10b Verified Read

    Expert annotators' average inter-annotator agreement: accuracy 0.82, kappa 0.76

    "Average 0.82 0.76 0.94 0.91"

    Zhao et al., SciArena, arXiv 2507.01001v2, Table 1 · page-text offset 21.9% · arxiv.org

    Link
    Cite

    Zhao et al., SciArena, arXiv 2507.01001v2, Table 1, page-text offset 21.9%, https://arxiv.org/pdf/2507.01001v2. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L10b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L10b (source read 2026-09-26).

    Report an error
  31. L11 Verified Read

    The best frontier judge reached physician agreement on 3 of 5 clinical grading tasks (NEJM Healer, BIDMC ER, Landmark) and fell short on two; no single model matched on all five

    "No single base model matched physician agreement across all five tasks."

    Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results · page-text offset 15.1% · numbers present in text: 82%, 77%, 91%, 88%, 68%, 67%, 87%, 92%, 84%, 95% · arxiv.org

    Link
    Cite

    Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results, page-text offset 15.1% · numbers present in text: 82%, 77%, 91%, 88%, 68%, 67%, 87%, 92%, 84%, 95%, https://arxiv.org/pdf/2609.12822v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L11, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L11 (source read 2026-09-26).

    Report an error
  32. L11a Verified Read

    Best judge vs physician agreement per task: NEJM Healer 82% vs 77%, BIDMC ER 91% vs 88%, Landmark 68% vs 67%; short on NEJM CPCs 87% vs 92% and Grey Matters Management 84% vs 95%

    "reached physician inter-rater agreement on the NEJM Healer (Claude, 82% vs. 77% for physicians), BIDMC ER (Claude, 91% vs. 88%), and Landmark Diagnostic Cases (Gemini, 68% vs. 67%), but fell short on the NEJM CPCs (GPT-5, 87% vs. 92%) and the Grey Matters Management Cases (Claude, 84% vs. 95%)"

    Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results · page-text offset 14.5% · arxiv.org

    Link
    Cite

    Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results, page-text offset 14.5%, https://arxiv.org/pdf/2609.12822v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L11a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L11a (source read 2026-09-26).

    Report an error
  33. L11b Verified Read

    In the clinical study, some tasks score a diagnosis against the known answer (the Bond score) and others score diagnostic reasoning against a rubric

    "A score of 5 indicates that the correct diagnosis is included in the differential"

    Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results, Methods, Task Rubrics · page-text offset 48.9% · arxiv.org

    Link
    Cite

    Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results, Methods, Task Rubrics, page-text offset 48.9%, https://arxiv.org/pdf/2609.12822v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L11b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L11b (source read 2026-09-26).

    Report an error
  34. L11c Verified Read

    The authors' own grader, PrecepTron, fine-tuned on a small number of physician-scored cases, matched or exceeded physician agreement on three of five tasks, so no single model, base or fine-tuned, matched on all five (the abstract's 'physician-level consistent scoring across tasks' is looser than the Results and must not be cited)

    "PrecepTron approaches physician inter-rater agreement, matching or exceeding it on three of five tasks (NEJM CPCs, NEJM Healer, and BIDMC ER)"

    Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results (PrecepTron section) · Results · arxiv.org

    Link
    Cite

    Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results (PrecepTron section), Results, https://arxiv.org/pdf/2609.12822v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L11c, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L11c (source read 2026-09-27).

    Report an error
  35. L12 Verified Read

    The GPT-4.1 grader using physician-written criteria scored above the average physician in five of seven themes (vendor-authored)

    "GPT-4.1 as a grader exceeds the random baseline for all themes as shown in Table 5. It exceeds the average physician score in five out of seven themes"

    Arora et al. (OpenAI), HealthBench, arXiv 2505.08775v1, §8.1 + Table 5 · page-text offset 42.6% · arxiv.org

    Link
    Cite

    Arora et al. (OpenAI), HealthBench, arXiv 2505.08775v1, §8.1 + Table 5, page-text offset 42.6%, https://arxiv.org/pdf/2505.08775v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L12, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L12 (source read 2026-09-26).

    Report an error
  36. L13 Verified Read

    Using the expert rubric, five of seven graders passed between 38.0% and 45.7% of the proofs experts failed and wrongly failed between 5.0% and 12.3% of the proofs experts passed; the other two passed 63.6% and 74.8% of the failed proofs (false fails 3.8% and 2.5%)

    "82.7% 43.0% 5.2% 82.0% 45.7% 5.0% 81.8% 38.0% 8.8% 80.9% 41.4% 8.7% 78.9% 39.6% 12.3% 77.0% 63.6% 3.8% 74.2% 74.8% 2.5%"

    Gonzalez et al., QEDBench (ICML 2026), arXiv 2602.20629v3, Fig. 5 · page-text offset 5.4% · arxiv.org

    Link
    Cite

    Gonzalez et al., QEDBench (ICML 2026), arXiv 2602.20629v3, Fig. 5, page-text offset 5.4%, https://arxiv.org/pdf/2602.20629v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L13, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L13 (source read 2026-09-26).

    Report an error
  37. L13a Verified Read

    The paper defines leniency as the false-positive rate and harshness as the false-negative rate against the expert rubric pass mark

    "We decompose errors into Leniency Rate(False Positives) and Harshness Rate(False"

    Gonzalez et al., QEDBench, Fig. 5 caption · page-text offset 5.5% · arxiv.org

    Link
    Cite

    Gonzalez et al., QEDBench, Fig. 5 caption, page-text offset 5.5%, https://arxiv.org/pdf/2602.20629v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L13a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L13a (source read 2026-09-26).

    Report an error
  38. L13b Verified Read

    The paper reports significant positive bias for certain frontier graders it names (Claude Opus 4.5, DeepSeek-V3, Qwen 2.5 Max, Llama 4 Maverick); it does not say this of all graders

    "frontier evaluators like Claude Opus 4.5, DeepSeek-V3, Qwen 2.5 Max, and Llama 4 Maverick exhibit significant positive bias"

    Gonzalez et al., QEDBench, abstract · page-text offset 0.5% · arxiv.org

    Link
    Cite

    Gonzalez et al., QEDBench, abstract, page-text offset 0.5%, https://arxiv.org/pdf/2602.20629v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L13b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L13b (source read 2026-09-26).

    Report an error
  39. L13c Verified Read

    QEDBench's leniency and harshness figures are for graders using the expert rubric

    "Judge Reliability Metrics (Expert Rubric Only)"

    Gonzalez et al., QEDBench (ICML 2026), arXiv 2602.20629v3, Fig. 5 label · page-text offset 5.4% · arxiv.org

    Link
    Cite

    Gonzalez et al., QEDBench (ICML 2026), arXiv 2602.20629v3, Fig. 5 label, page-text offset 5.4%, https://arxiv.org/pdf/2602.20629v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L13c, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L13c (source read 2026-09-26).

    Report an error
  40. L14 Verified Read

    Order-swap flip rate (MT-Bench / JudgeBench): Gemini 3.1 Pro 0.035 / 0.020; Claude Opus 4.6 0.038 / 0.022

    "Gemini 3.1 Pro 0.977 0.989 0.035 0.978 0.989 0.020 Claude Sonnet 4 0.960 0.976 0.074 0.941 0.966 0.146 GPT-4o-mini 0.959 0.975 0.127 0.896 0.939 0.380 Claude Opus 4.6 0.958 0.979 0.038 0.974 0.986 0.022"

    Norman, Rivera, Hughes, Reliability without Validity (preprint), arXiv 2606.19544v1, App. Table 6 · page-text offset 78.0% · arxiv.org

    Link
    Cite

    Norman, Rivera, Hughes, Reliability without Validity (preprint), arXiv 2606.19544v1, App. Table 6, page-text offset 78.0%, https://arxiv.org/pdf/2606.19544v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L14, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L14 (source read 2026-09-26).

    Report an error
  41. L14a Verified Read

    Order-swap flip rate (MT-Bench / JudgeBench): GPT-5.4 0.114 / 0.105

    "GPT-5.4 0.932 0.965 0.114 0.933 0.966 0.105"

    Norman, Rivera, Hughes, Reliability without Validity (preprint), arXiv 2606.19544v1, App. Table 6 · page-text offset 78.8% · arxiv.org

    Link
    Cite

    Norman, Rivera, Hughes, Reliability without Validity (preprint), arXiv 2606.19544v1, App. Table 6, page-text offset 78.8%, https://arxiv.org/pdf/2606.19544v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L14a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L14a (source read 2026-09-26).

    Report an error
  42. L15 Verified Read

    2026 JudgeBench, chance-corrected agreement (Cohen's kappa): Gemini 3.1 Pro 0.841, Claude Opus 4.6 0.875, Claude Sonnet 4.6 0.782, GPT-5.4 0.606; exact match 0.964, 0.956, 0.920, 0.812 (one order, ties excluded)

    "Gemini 3.1 Pro 0.849 0.511 33.8 0.964 0.841 12.3 0.956 0.898 5.9 Claude Opus 4.6 0.848 0.489 35.9 0.956 0.875 8.1 0.943 0.879 6.4 DeepSeek V3.2 0.845 0.486 35.9 0.791 0.545 24.5 0.921 0.826 9.5 Claude Sonnet 4.6 0.851 0.484 36.7 0.920 0.782 13.8 0.942 0.871 7.1 Llama 3.3 70B 0.841 0.465 37.6 0.664 0.283 38.1 0.892 0.769 12.3 Kimi K2.5 0.846 0.461 38.5 0.864 0.720 14.5 0.937 0.873 6.4 GPT-5.4 0.836 0.457 38.0 0.812 0.606"

    Norman, Rivera, Hughes, Reliability without Validity (preprint), arXiv 2606.19544v1, §4.1 Table 2 · page-text offset 24.2% · arxiv.org

    Link
    Cite

    Norman, Rivera, Hughes, Reliability without Validity (preprint), arXiv 2606.19544v1, §4.1 Table 2, page-text offset 24.2%, https://arxiv.org/pdf/2606.19544v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L15, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L15 (source read 2026-09-26).

    Report an error
  43. L15b Verified Read

    The same authors warn that raw agreement does not correct for chance and overstates how well a grader discriminates

    "This family of metrics does not correct for chance"

    Norman, Rivera, Hughes (preprint), arXiv 2606.19544v1, §2 · page-text offset 4.0% · arxiv.org

    Link
    Cite

    Norman, Rivera, Hughes (preprint), arXiv 2606.19544v1, §2, page-text offset 4.0%, https://arxiv.org/pdf/2606.19544v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L15b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L15b (source read 2026-09-26).

    Report an error
  44. L16 Verified Read

    Human annotators preferred the markdown side 57% of the time; four of five judges preferred it 73%–97% on the same pairs

    "human annotators prefer the markdown side only 57% of the time, while four of the five judges prefer it 73%-97% on the same pairs"

    Soumik, bias mitigation in LLM-as-a-Judge (TMLR 06/2026), arXiv 2604.23178v2 · page-text offset 44.2% · arxiv.org

    Link
    Cite

    Soumik, bias mitigation in LLM-as-a-Judge (TMLR 06/2026), arXiv 2604.23178v2, page-text offset 44.2%, https://arxiv.org/pdf/2604.23178v2. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L16, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L16 (source read 2026-09-26).

    Report an error
  45. L16a Verified Read

    The human comparison was two independent annotators on a 30-pair subsample: they preferred markdown 57% of the time on average, while four of the five judges preferred it 73%-97%

    "two independent annotators on a 30-pair subsample preferred markdown only 57% of the time on average, while four of the five judges preferred markdown 73%-97%"

    Soumik, bias mitigation in LLM-as-a-Judge (TMLR 06/2026), arXiv 2604.23178v2, §1 / App. G · page-text offset 28.5% · arxiv.org

    Link
    Cite

    Soumik, bias mitigation in LLM-as-a-Judge (TMLR 06/2026), arXiv 2604.23178v2, §1 / App. G, page-text offset 28.5%, https://arxiv.org/pdf/2604.23178v2. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L16a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L16a (source read 2026-09-26).

    Report an error
  46. L16b Verified Read

    The style test compared markdown with plain prose

    "STYLE (markdown vs. plain prose)"

    Soumik, bias mitigation in LLM-as-a-Judge (TMLR 06/2026), arXiv 2604.23178v2, §3 · page-text offset 22.0% · arxiv.org

    Link
    Cite

    Soumik, bias mitigation in LLM-as-a-Judge (TMLR 06/2026), arXiv 2604.23178v2, §3, page-text offset 22.0%, https://arxiv.org/pdf/2604.23178v2. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L16b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L16b (source read 2026-09-26).

    Report an error
  47. L17 Verified Read

    One study finds most measured self-preference is judges being unsure on hard items (its own summary: 89.6%, substantially reducing but not eliminating the evidence)

    "evaluator uncertainty accounts for an average of 89.6%"

    Roytburg et al., Are LLM Evaluators Really Narcissists? (ICML 2026), arXiv 2601.22548v4, §1 · page-text offset 9.2% · numbers present in text: 89.6% · arxiv.org

    Link
    Cite

    Roytburg et al., Are LLM Evaluators Really Narcissists? (ICML 2026), arXiv 2601.22548v4, §1, page-text offset 9.2% · numbers present in text: 89.6%, https://arxiv.org/pdf/2601.22548v4. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L17, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L17 (source read 2026-09-26).

    Report an error
  48. L18 Verified Read

    GPT-5 as judge: false-positive rate on its own wrong GSM8K answers 82.0 [72.4, 90.7], 67.3 points above its rate on others' wrong answers

    "GPT-5 82.0[72.4, 90.7]+67.3"

    Zhang et al. (Microsoft/MIT), Can We Trust LLM Judges (preprint), arXiv 2609.12002v1, App. B Table 10 (GSM8K) · page-text offset 87.4% · arxiv.org

    Link
    Cite

    Zhang et al. (Microsoft/MIT), Can We Trust LLM Judges (preprint), arXiv 2609.12002v1, App. B Table 10 (GSM8K), page-text offset 87.4%, https://arxiv.org/pdf/2609.12002v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L18, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L18 (source read 2026-09-26).

    Report an error
  49. L19 Verified Read

    Across three QA sets (NQ, TQA, HPQA), the panel (PoLL) beat GPT-4 alone on all three (0.763 vs 0.627; 0.906 vs 0.841; 0.867 vs 0.830), edged its best member on two (0.763 vs 0.749; 0.906 vs 0.902) and fell below Haiku on the third (0.867 vs 0.873)

    "GPT-4 0.627 0.841 0.830 CMD-R 0.734 0.902 0.815 Haiku 0.749 0.894 0.873 GPT-3.5 0.726 0.859 0.833 PoLL 0.763 0.906 0.867"

    Verga et al. (Cohere), Replacing Judges with Juries, arXiv 2404.18796v2, Table 1 · page-text offset 18.5% · arxiv.org

    Link
    Cite

    Verga et al. (Cohere), Replacing Judges with Juries, arXiv 2404.18796v2, Table 1, page-text offset 18.5%, https://arxiv.org/pdf/2404.18796v2. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L19, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L19 (source read 2026-09-26).

    Report an error
  50. L21 Verified Read

    When the judge is no better than the model it grades, correcting it with human labels can at most halve the labels you need

    "no debiasing method can decrease the required amount of ground truth labels by more than half"

    Dorner, Nastl, Hardt (ICLR 2025 Oral), arXiv 2410.13341v4, abstract · page-text offset 1.4% · arxiv.org

    Link
    Cite

    Dorner, Nastl, Hardt (ICLR 2025 Oral), arXiv 2410.13341v4, abstract, page-text offset 1.4%, https://arxiv.org/pdf/2410.13341v4. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L21, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L21 (source read 2026-09-26).

    Report an error
  51. L22 Verified Read

    In the original JudgeBench (2024), many strong graders, GPT-4o among them, scored only slightly better than random on hard correctness pairs

    "performing just slightly better than random guessing"

    Tan et al., JudgeBench (ICLR 2025), arXiv 2410.12784v2, abstract · page-text offset 1.8% · arxiv.org

    Link
    Cite

    Tan et al., JudgeBench (ICLR 2025), arXiv 2410.12784v2, abstract, page-text offset 1.8%, https://arxiv.org/pdf/2410.12784v2. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L22, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L22 (source read 2026-09-26).

    Report an error
  52. L22a Verified Read

    JudgeBench found a grader's ability to verify answers highly correlated with its ability to solve the problem itself

    "the ability of the judge to verify the solution pairs is highly correlated with its ability to solve the problem itself"

    Tan et al., JudgeBench (ICLR 2025), arXiv 2410.12784v2, §4.4 · page-text offset 41.3% · arxiv.org

    Link
    Cite

    Tan et al., JudgeBench (ICLR 2025), arXiv 2410.12784v2, §4.4, page-text offset 41.3%, https://arxiv.org/pdf/2410.12784v2. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L22a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L22a (source read 2026-09-26).

    Report an error
  53. L23 Verified Read

    PRIOR ART: Eugene Yan's August 2024 review of LLM-evaluators drew on about two dozen papers and already covers reference-based evaluation and position bias

    "Drawing from two dozen papers"

    Eugene Yan, Evaluating the Effectiveness of LLM-Evaluators (aka LLM-as-Judge), Aug 2024 · page-text offset 1.5% · eugeneyan.com

    Link
    Cite

    Eugene Yan, Evaluating the Effectiveness of LLM-Evaluators (aka LLM-as-Judge), Aug 2024, page-text offset 1.5%, https://eugeneyan.com/writing/llm-evaluators/. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L23, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L23 (source read 2026-09-26).

    Report an error
  54. L24 Verified Read

    PRIOR ART: Hamel Husain's guide (Oct 2024) has a principal domain expert make pass/fail judgments and tracks the judge's agreement with them; the warning that agreement misleads on imbalanced data was added in a 2025 revision

    "I also tracked agreement rates over time to ensure we were converging on a good prompt"

    Hamel Husain, Using LLM-as-a-Judge For Evaluation: A Complete Guide (Oct 2024) · page-text offset 60.7% · hamel.dev

    Link
    Cite

    Hamel Husain, Using LLM-as-a-Judge For Evaluation: A Complete Guide (Oct 2024), page-text offset 60.7%, https://hamel.dev/blog/posts/llm-judge/. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L24, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L24 (source read 2026-09-26).

    Report an error
  55. L24a Verified Read

    Husain's guide, as revised in September 2026, advises about 100 examples per failure mode to validate an automated grader and says below 60 the confidence intervals are often too wide; the October 2024 original did not contain this advice

    "Aim for about 100 examples per failure mode, with enough Pass and Fail examples to measure both classes. Below 60 examples, the confidence intervals are often too wide"

    Hamel Husain, Using LLM-as-a-Judge For Evaluation (Oct 2024) · page-text offset 50.8% · hamel.dev

    Link
    Cite

    Hamel Husain, Using LLM-as-a-Judge For Evaluation (Oct 2024), page-text offset 50.8%, https://hamel.dev/blog/posts/llm-judge/. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L24a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L24a (source read 2026-09-26).

    Report an error
  56. L25 Verified Read

    In code review against a written requirement, GPT-4o wrongly failed correct code (false-negative rate) 26.2% (HumanEval) and 35.9% (MBPP) when judging directly, rising to 73.2% and 87.9% when also asked to explain and repair

    "GPT-4o achieves a relatively low FNR in HumanEval (26.2%) and MBPP (35.9%), but once explana- tions and repairs are required, the FNR increases sharply to 73.2% in HumanEval and 87.9% in MBPP"

    Jin & Chen, Are LLMs Reliable Code Reviewers? Systematic Overcorrection (Automated Software Engineering 33, art. 90, 26 Jun 2026; read as arXiv 2603.00539v1, and the published text prints the same sentence), §5.2 · page-text offset 38.7% · numbers present in text: 26.2, 35.9, 73.2, 87.9 · arxiv.org

    Link
    Cite

    Jin & Chen, Are LLMs Reliable Code Reviewers? Systematic Overcorrection (Automated Software Engineering 33, art. 90, 26 Jun 2026; read as arXiv 2603.00539v1, and the published text prints the same sentence), §5.2, page-text offset 38.7% · numbers present in text: 26.2, 35.9, 73.2, 87.9, https://arxiv.org/pdf/2603.00539v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L25, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L25 (source read 2026-09-26).

    Report an error
  57. L25a Verified Read

    In that paper, FNR is rejecting correct code (a false fail) and FPR is accepting buggy code (a false pass)

    "FNR reflects over-correction (rejecting correct implementations), while FPR reflects unsafe acceptance"

    Jin & Chen, Are LLMs Reliable Code Reviewers? Systematic Overcorrection (Automated Software Engineering 33, art. 90, 2026; read as arXiv 2603.00539v1), Table 2 caption · page-text offset 37.2% · arxiv.org

    Link
    Cite

    Jin & Chen, Are LLMs Reliable Code Reviewers? Systematic Overcorrection (Automated Software Engineering 33, art. 90, 2026; read as arXiv 2603.00539v1), Table 2 caption, page-text offset 37.2%, https://arxiv.org/pdf/2603.00539v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L25a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L25a (source read 2026-09-26).

    Report an error
  58. P1 Executed Read

    The prior registered before any paper was read: graders would hold on easy, preference-style chat grading and fail on correctness of hard items they cannot solve themselves and in specialised domains

    Result: judges HOLD on easy, preference-style chat grading and FAIL on (a) correctness of hard items the judge cannot itself solve and (b) specialised domains

    first-party: pre-registration, 19:57 BST 2026-09-26

    A run or observation done for this piece. The run record is not published.

    Link
    Cite

    A run or observation by The Durability Curve (first-party: pre-registration, 19:57 BST 2026-09-26). Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row P1, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-P1 (source read 2026-09-26).

    Report an error
  59. J1 Executed Read

    As run, scored to the precision each question asked for (one pass, n = 24 arithmetic-style questions with one right answer each, listed in Line/judge-check/items.py, 2026-09-26 21:02 BST): Claude Haiku 4.5, answering without tools, got 4 of 24 wrong. Haiku 4.5 grading those answers without a key passed 3 of its 4 wrong answers; given the verified key, it passed 0 of 4

    Result: HEADLINE haiku passed 3/4 of its own wrong answers without a key, 0/4 with one

    first-party: Judge Check run

    A run or observation done for this piece. The run record is not published.

    Link
    Cite

    A run or observation by The Durability Curve (first-party: Judge Check run). Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row J1, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J1 (source read 2026-09-26).

    Report an error
  60. J2 Executed Read

    Claude Sonnet 5 grading the same answers without a key agreed with the key on 23 of 24: no false passes, one false fail (a correct final answer reached by an unjustified leap, which it failed although told to grade the final answer); with the key, 24 of 24

    Result: no key agree 23/24 · false passes 0 · false fails 1 · unparsed 0 / with key agree 24/24

    first-party: Judge Check run

    A run or observation done for this piece. The run record is not published.

    Link
    Cite

    A run or observation by The Durability Curve (first-party: Judge Check run). Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row J2, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J2 (source read 2026-09-26).

    Report an error
  61. J3 Executed Read

    Re-dressing the same answers in markdown changed 2 of 24 of Haiku's verdicts (both from pass to fail on wrong answers) and none of Sonnet's

    Result: markdown re-dress changed 2/24 verdicts / markdown re-dress changed 0/24 verdicts

    first-party: Judge Check run

    A run or observation done for this piece. The run record is not published.

    Link
    Cite

    A run or observation by The Durability Curve (first-party: Judge Check run). Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row J3, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J3 (source read 2026-09-26).

    Report an error
  62. J4 Executed Read

    As run (the precision each question asked for): Haiku as judge could not solve 2 of the 24 items itself and, without a key, passed the wrong answer on both; on H6 (digit sum of 50!) it passed a wrong answer although it solved the item correctly when asked. Under the registered rule it could not solve 1 item (H11) and passed the wrong answer on it; H6 holds under both rules

    Result: solved it itself: 22/24 · no-key agree on solved 21/22 · on unsolved 0/2

    first-party: Judge Check run

    A run or observation done for this piece. The run record is not published.

    Link
    Cite

    A run or observation by The Durability Curve (first-party: Judge Check run). Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row J4, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J4 (source read 2026-09-26).

    Report an error
  63. J5 Executed Read

    The run cost $2.83. As run, predictions P1, P4 and P5 passed, P3 passed for Haiku, P2 and Sonnet's P3 were unscorable. Under the pre-registered scoring rule (score.py --registered), P1 failed for both graders, P5 failed, Sonnet's P3 failed, P4 passed, Haiku's P3 passed

    Result: cost $2.8326 · Predictions (pre-registered)

    first-party: Judge Check run

    A run or observation done for this piece. The run record is not published.

    Link
    Cite

    A run or observation by The Durability Curve (first-party: Judge Check run). Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row J5, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J5 (source read 2026-09-26).

    Report an error
  64. T1 Executed Read

    Of 19 studies we found that report which way AI graders err against human or ground-truth labels, 12 lean toward passing flawed work (4 of them weakly, adversarially or in one direction only), 3 toward failing sound work (code review against a spec; older-model essay grading; a frontier panel on clinical diagnoses) and 4 are mixed

    Result: Tally (Part 2, excluding 7b): Supports 12 (1, 2, 3, 4, 5, 6, 9, 13, 14, 16, 18, 19; of these, 13, 16, 18 and 19 are weak, adversarial or one-directional) · Contradicts 3 (7, 8, 11) · Mixed 4 (10, 12, 15, 17)

    first-party: error-direction sweep, 2026-09-26

    A run or observation done for this piece. The run record is not published.

    Link
    Cite

    A run or observation by The Durability Curve (first-party: error-direction sweep, 2026-09-26). Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row T1, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-T1 (source read 2026-09-26).

    Report an error
  65. T2 Executed Read

    The evidence base is 77 distinct studies: 63 in the four sweeps and 14 more in the error-direction check

    Result: PAPERS TOTAL 77 distinct studies

    first-party: paper count

    A run or observation done for this piece. The run record is not published.

    Link
    Cite

    A run or observation by The Durability Curve (first-party: paper count). Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row T2, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-T2 (source read 2026-09-26).

    Report an error
  66. J6 Executed Read

    Under the pre-registered scoring rule (0.001 relative tolerance), Haiku's answers were wrong on 2 of 24; Haiku grading them passed both wrong answers without a key and neither with one; both graders, given the key, failed the two answers that missed the precision the question asked for

    Result: HEADLINE haiku passed 2/2 of its own wrong answers without a key, 0/2 with one

    first-party: Judge Check run, rescored

    A run or observation done for this piece. The run record is not published.

    Link
    Cite

    A run or observation by The Durability Curve (first-party: Judge Check run, rescored). Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row J6, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J6 (source read 2026-09-26).

    Report an error
  67. W1 Executed Read

    Our earlier piece A Green Score Is Not Evidence reported a groundedness metric that scored its best with the evidence removed

    Result: A groundedness metric scored its best with the evidence removed.

    first-party: published article (subtitle) · harryfloyd.substack.com

    A run or observation done for this piece. The run record is not published.

    Link
    Cite

    A run or observation by The Durability Curve (first-party: published article (subtitle)), https://harryfloyd.substack.com/p/a-green-score-is-not-evidence. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row W1, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-W1 (source read 2026-09-26).

    Report an error
  68. W2 Executed Read

    Our earlier piece Your AI Looks Best Where You Can Check It Least argued that the work that is hardest to check is where AI output looks best

    Result: Your AI Looks Best Where You Can Check It Least

    first-party: published article (title) · harryfloyd.substack.com

    A run or observation done for this piece. The run record is not published.

    Link
    Cite

    A run or observation by The Durability Curve (first-party: published article (title)), https://harryfloyd.substack.com/p/your-ai-looks-best-where-you-check-least. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row W2, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-W2 (source read 2026-09-26).

    Report an error

Cite and reuse

A single claim
Use the row's Cite button. It names the original source first, then this row, which carries its own link.
This ledger
Floyd, Harry (2026). Claim Ledger: Your AI Grader Is Only as Good as Its Answer Key. The Durability Curve. https://durabilitycurve.com/claims/your-ai-grader-answer-key/
The essay
Floyd, Harry (2026). Your AI Grader Is Only as Good as Its Answer Key. The Durability Curve. https://durabilitycurve.com/blog/your-ai-grader-answer-key/

The ledger data is licensed CC BY 4.0: reuse it, quote it or train on it, with credit to The Durability Curve and a link. Read the licence. Machine copies: JSON, CSV, Markdown.