Claim ledger
Your AI Grader Is Only as Good as Its Answer Key
I went through 77 studies of AI graders. Without a verified answer key, they are least reliable on exactly the questions they can't answer themselves. Here is a twenty-item check for yours.
What the essay claims
On correctness, the variable that decides whether an AI judge can be trusted is not which model grades but whether it has a verified reference to grade against. Without one, judges agree with experts mostly on the questions they could answer themselves, and fall apart on the rest; with one, even small models grade close to how well humans agree with each other. When judges are wrong about correctness, they mostly pass flawed work (12 of 19 studies that measure direction; code review is the counter-case). On matters of taste and expert judgement, where no key exists, judges more often fall short of expert–expert agreement than match it. The popular fixes (panels, reasoning first, fine-tuned judges, calibration on a few human labels) help less than claimed. The test that screens it for any one setup is the reader's own: twenty items, with and without a key, split by whether the judge could solve them.
The claim ladder
| Rung | Claim | Evidence (ledger) | Whose behaviour | What it does NOT reach |
|---|---|---|---|---|
| R1 | With a verified key, judges grade hard correctness close to human level; a self-made or subtly wrong key fails | L1–L5 | Judges on expert-graded finance/maths (GPT-4o + 4 open models); 2026 frontier judges on a proxy label | Subjective grading (no key exists); 2026 rows are corroboration, not human-labelled |
| R2 | Without a key, judges agree with experts mostly on questions they could answer themselves | L1, L5 | Same | Whether the newest judges (Opus 5.5, GPT-6) still show it: untested in the literature |
| R3 | On taste and expert judgement, judges more often fall short of expert–expert agreement | L7–L12 | Chat preference, science answers, clinical grading | Mixed: HealthBench 5/7 themes at or above the average physician (vendor-authored); GRAND-ROUNDS 3/5 |
| R4 | When wrong about correctness, judges mostly pass flawed work | L6, L13 + direction sweep (12/19) | Judges grading answers and proofs | Reverses in code review against a spec; not a law |
| R5 | Biases are judge-specific | L14–L16 | Named 2025–26 judges on MT-Bench / JudgeBench / style pairs | Human baseline for markdown is n = 2 annotators |
| R6 | Self-preference is contested | L17–L18 | Judges grading their own vs others' answers | Unresolved between the two papers; say so |
| R7 | Popular fixes help less than claimed | L19–L21 | Panels, reasoning-first, human-label calibration | Panels can edge their best member (L19); some reasoning helps on hard pairs |
| R8 | The Judge Check can catch a generous grader on twenty items; it cannot clear one (about 100 per failure mode, L24a) | Object Record (our run, n = 24) | The check, on Haiku 4.5 and Sonnet 5 grading Haiku 4.5 | A demonstration, not evidence; one run, one task type |
The evidence, row by row
Each row is a claim the essay makes, then the evidence recorded for it and the source it was checked against. For verified rows that is usually the exact words, where they sit and the day they were read. The label says how far it was checked.
What the labels mean
- Verified
- Checked against the primary source itself by a checker who did not draft the piece, usually with the exact words, where they sit and the date read.
- Executed
- A run or observation done for the piece: code, a query, a count, or a check made live in an app. The claim is what it returned. Run records are not published.
- Checked
- Checked against its source by a separate checker in the older ledger format, which recorded no quote, locator or date.
- Reported
- Not verified word for word against a primary source: carried from a secondary source, from one that could not be opened in full, or supported only in part. The essay words it accordingly.
- Struck
- Drafted, checked and removed before publication. Kept here, crossed out, with the reason.
- Excluded
- Considered and deliberately left out of the piece. Kept here with the reason.
Sources (21)
- Krumdick et al. 15 rows
- Chandak et al. 5 rows
- Zheng et al. 4 rows
- Kim et al. 1 row
- Calderon 1 row
- Chiang et al. 1 row
- Zhao et al. 3 rows
- Buckley et al. 4 rows
- Arora et al. 1 row
- Gonzalez et al. 4 rows
- Norman 4 rows
- Soumik 3 rows
- Roytburg et al. 1 row
- Zhang et al. 1 row
- Verga et al. 1 row
- Dorner 1 row
- Tan et al. 2 rows
- Eugene Yan 1 row
- Hamel Husain 2 rows
- Jin & Chen 2 rows
- First-party runs 11 rows
-
Without a correct reference, judges agreed well with human experts only on questions they had been able to answer themselves
"when not provided with a correct reference, judges show high agreement with human experts only on questions the judges were able to correctly answer themselves"
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, abstract · page-text offset 1.2% · arxiv.org
LinkReport an errorCite
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, abstract, page-text offset 1.2%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L1, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L1 (source read 2026-09-26). -
GPT-4o as judge, Cohen's kappa with experts. Pairwise: questions it answered correctly 0.78 with no key, 0.92 with a human key; questions it got wrong 0.30 with no key, 0.16 with its own answer as the key, 0.83 with a human key. Single grading: questions it answered correctly 0.46 with no key and 0.59 with a human key; questions it got wrong 0.16 and 0.81
"GPT-4o Correct 0.78±0.02 0.86±0.02 0.92±0.02 0.46±0.07 0.52±0.06 0.59±0.06 Incorrect 0.30±0.09 0.16±0.08 0.83±0.06 0.16±0.14 0.13±0.14 0.81±0.13"
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Table 4 (GPT-4o row) · page-text offset 21.0% · arxiv.org
LinkReport an errorCite
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Table 4 (GPT-4o row), page-text offset 21.0%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L1a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L1a (source read 2026-09-26). -
Pairwise, questions each judge got wrong, no key vs human key: Llama 3.3 70B 0.39 vs 0.92; Phi 4 -0.08 vs 0.75; Qwen 2.5 7B 0.21 vs 0.63; Yi 1.5 34B 0.13 vs 0.56 (own answer as key: 0.35, -0.08, 0.14, 0.02)
"Llama 3.3 70b Correct 0.79±0.03 0.86±0.03 0.93±0.02 0.61±0.08 0.54±0.08 0.73±0.07 Incorrect 0.39±0.06 0.35±0.07 0.92±0.03 0.13±0.06 0.23±0.10 0.69±0.09 Phi 4 Correct 0.67±0.07 0.88±0.04 0.91±0.04 0.48±0.09 0.47±0.07 0.67±0.07 Incorrect -0.08±0.09 -0.08±0.10 0.75±0.06 0.03±0.05 0.02±0.12 0.56±0.12 Qwen 2.5 7B Correct 0.65±0.05 0.69±0.04 0.78±0.04 0.31±0.11 0.37±0.09 0.54±0.09 Incorrect 0.21±0.07 0.14±0.07 0.63±0.06 0.08±0.07 0.21±0.10 0.56±0.08 Yi 1.5 34B Correct 0.31±0.12 0.53±0.12 0.77±0.09 0.23±0.11 0.29±0.08 0.43±0.09 Incorrect 0.13±0.09 0.02±0.09 0.56±0.08"
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Table 4 (other judges) · page-text offset 21.2% · arxiv.org
LinkReport an errorCite
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Table 4 (other judges), page-text offset 21.2%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L1b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L1b (source read 2026-09-26). -
The paper notes that a relatively small model (Qwen 2.5 7B) with a human reference can grade better than a larger model (GPT-4o) without one, on the same set of responses
"providing a human reference to a relatively small model (such as Qwen 2.5 7B ) can yield better judgments than using a larger model without human references (such asGPT-4o) on the same set of responses"
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6 · page-text offset 22.2% · arxiv.org
LinkReport an errorCite
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6, page-text offset 22.2%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L1c, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L1c (source read 2026-09-26). -
With its own answer as the key, GPT-4o's pairwise kappa was 0.86 on questions it had answered correctly and 0.16 on questions it could not answer
"decreases from 0.86 to 0.16 in the pairwise case on questions GPT-4o could not answer"
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6 · page-text offset 23.0% · arxiv.org
LinkReport an errorCite
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6, page-text offset 23.0%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L2, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L2 (source read 2026-09-26). -
Averaged across all five graders on questions GPT-4o answered correctly (Table 3 does not say whether single and pairwise grading are pooled): kappa with a human key 0.69, a verified GPT-4o answer as key 0.61, a completely unrelated key 0.50, no key 0.46, a subtly wrong key 0.21
"Human 0.69 ±0.06 0.74±0.06 0.57±0.08 GPT-4o (✓) 0.61 ±0.07 0.66±0.07 0.52±0.09 None 0.46 ±0.08 0.56±0.08 0.22±0.09 Random 0.50 ±0.08 0.47±0.08 0.09±0.1 Wrong 0.21 ±0.06"
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Table 3 · page-text offset 19.3% · arxiv.org
LinkReport an errorCite
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Table 3, page-text offset 19.3%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L3, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3 (source read 2026-09-26). -
Table 3's figures are across all five graders
"Table 3 presents results for these reference types across all five judges"
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6 · page-text offset 26.3% · arxiv.org
LinkReport an errorCite
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6, page-text offset 26.3%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L3c, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3c (source read 2026-09-26). -
The paper says a slightly incorrect reference can in some cases be worse than no reference at all
"in some cases be worse than"
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6 · page-text offset 25.4% · arxiv.org
LinkReport an errorCite
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6, page-text offset 25.4%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L3a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3a (source read 2026-09-26). -
The authors describe the Random key as a completely unrelated reference, and say a slightly incorrect reference can in some cases be worse than a completely unrelated one or none
"providing a slightly incorrect reference can in some cases be worse than providing a completely unrelated reference or no reference at all"
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6 · page-text offset 25.4% · arxiv.org
LinkReport an errorCite
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6, page-text offset 25.4%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L3d, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3d (source read 2026-09-26). -
The authors conclude that verifying a model-generated answer, rather than writing and checking one by hand, was sufficient as a key in their setting
"Verifying responses generated by a model rather than manually writing and checking answers via annotators is sufficient in this setting"
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §7 · page-text offset 28.6% · arxiv.org
LinkReport an errorCite
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §7, page-text offset 28.6%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L3b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3b (source read 2026-09-26). -
2026 judges, questions each got wrong, against a stand-in label: single grading no key vs human key, Opus 4.7 0.26 vs 0.62, GPT-5.4 0.32 vs 0.67, Gemini 3.1 0.44 vs 0.71; pairwise 0.33 vs 0.66, 0.46 vs 0.87, 0.67 vs 0.95 (own answer as key, pairwise: 0.33, 0.45, 0.62)
"Opus 4.7 Correct 0.86±0.04 0.90±0.04 0.91±0.04 0.63±0.06 0.69±0.06 0.77±0.05 Incorrect 0.33±0.16 0.33±0.15 0.66±0.13 0.26±0.14 0.27±0.12 0.62±0.11 GPT-5.4 Correct 0.75±0.06 0.87±0.04 0.89±0.04 0.51±0.06 0.64±0.06 0.66±0.06 Incorrect 0.46±0.11 0.45±0.13 0.87±0.08 0.32±0.11 0.25±0.12 0.67±0.09 Gemini 3.1 Correct 0.78±0.06 0.80±0.06 0.83±0.05 0.55±0.06 0.62±0.06 0.69±0.06 Incorrect 0.67±0.13 0.62±0.14 0.95±0.07 0.44±0.13 0.35±0.14 0.71±0.10"
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, Table 12 · page-text offset 90.1% · arxiv.org
LinkReport an errorCite
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, Table 12, page-text offset 90.1%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L4, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4 (source read 2026-09-26). -
The authors treat the 2026-judge rows, scored against a stand-in label, as corroboration
"we regard these results as corroboration rather than equivalent to our main table"
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J · page-text offset 88.7% · arxiv.org
LinkReport an errorCite
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, page-text offset 88.7%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L4a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4a (source read 2026-09-26). -
The stand-in label is GPT-4o grading with the human-verified references
"We used the GPT-4o Judge with the human-verified references as a proxy correctness label"
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J · page-text offset 88.4% · arxiv.org
LinkReport an errorCite
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, page-text offset 88.4%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L4b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4b (source read 2026-09-26). -
The authors report that higher reasoning effort can narrow the no-reference gap, for example GPT-5.4 at high effort
"higher reasoning effort can narrow the no-reference gap (e.g.,GPT-5.4 at high effort)"
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J · page-text offset 88.9% · arxiv.org
LinkReport an errorCite
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, page-text offset 88.9%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L4c, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4c (source read 2026-09-26). -
GPT-5.4 at high reasoning effort, on questions it got wrong, against the stand-in label: pairwise 0.68 with no key and 0.84 with the human key; single grading 0.38 and 0.67
"GPT-5.4 low Correct 0.85±0.05 0.90±0.04 0.54±0.06 0.57±0.05 Incorrect 0.55±0.13 0.95±0.06 0.34±0.11 0.56±0.09 default Correct 0.75±0.06 0.89±0.04 0.51±0.06 0.66±0.06 Incorrect 0.46±0.11 0.87±0.08 0.32±0.11 0.67±0.09 high Correct 0.80±0.06 0.87±0.05 0.68±0.08 0.67±0.06 Incorrect 0.68±0.14 0.84±0.11 0.38±0.19 0.67±0.11"
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, Table 13 (GPT-5.4 rows) · page-text offset 91.4% · arxiv.org
LinkReport an errorCite
Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, Table 13 (GPT-5.4 rows), page-text offset 91.4%, https://arxiv.org/pdf/2503.05061v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L4d, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4d (source read 2026-09-26). -
On MATH, DeepSeek V3 as a judge without a reference reached Scott's pi 0.72 against ground truth; as a matcher with the reference, 0.98
"DeepSeek v3 model achieves only modest agreement π = 0.72, while as a matcher, it achieves π = 0.98"
Chandak et al., Answer Matching Outperforms Multiple Choice (preprint), arXiv 2507.02856v1, §3.1 · page-text offset 22.6% · arxiv.org
LinkReport an errorCite
Chandak et al., Answer Matching Outperforms Multiple Choice (preprint), arXiv 2507.02856v1, §3.1, page-text offset 22.6%, https://arxiv.org/pdf/2507.02856v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L5, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L5 (source read 2026-09-26). -
A 1.7-billion-parameter Qwen3 model matching answers to the reference reached pi 0.97 against ground truth on MATH
"answer matching, even with the 1.7 billion parameter Qwen3 model (non-thinking mode), achieves near-perfect alignment with the ground-truth (π = 0.97)"
Chandak et al., Answer Matching Outperforms Multiple Choice (preprint), arXiv 2507.02856v1, §3.1 · page-text offset 22.4% · arxiv.org
LinkReport an errorCite
Chandak et al., Answer Matching Outperforms Multiple Choice (preprint), arXiv 2507.02856v1, §3.1, page-text offset 22.4%, https://arxiv.org/pdf/2507.02856v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L5a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L5a (source read 2026-09-26). -
Answer matching with recent models, even small ones, reached near-perfect agreement with human grading, in the range of inter-annotator agreement
"matching using recent models-even small ones-achieves near-perfect agreement, in the range of inter-annotator agreement"
Chandak et al., Answer Matching (preprint), arXiv 2507.02856v1, abstract · page-text offset 1.3% · arxiv.org
LinkReport an errorCite
Chandak et al., Answer Matching (preprint), arXiv 2507.02856v1, abstract, page-text offset 1.3%, https://arxiv.org/pdf/2507.02856v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L5b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L5b (source read 2026-09-26). -
On MATH, the 'true grades' come from MATH-Verify, a rule-based checker that compares each answer with the reference
"MATH-Verify library (Kydlicek et al., 2025) implements rule-based ground-truth evaluations of generative responses"
Chandak et al., Answer Matching (preprint), arXiv 2507.02856v1, §3.1 · page-text offset 22.0% · arxiv.org
LinkReport an errorCite
Chandak et al., Answer Matching (preprint), arXiv 2507.02856v1, §3.1, page-text offset 22.0%, https://arxiv.org/pdf/2507.02856v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L5c, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L5c (source read 2026-09-26). -
For frontier judges (DeepSeek V3, o4-mini), 80%+ of grading errors were false passes: the judge marked correct what human annotation marked incorrect
"errors disproportionately (80%+) arise from false positives"
Chandak et al., arXiv 2507.02856v1, §3.2 · page-text offset 28.5% · arxiv.org
LinkReport an errorCite
Chandak et al., arXiv 2507.02856v1, §3.2, page-text offset 28.5%, https://arxiv.org/pdf/2507.02856v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L6, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L6 (source read 2026-09-26). -
On MT-Bench, without ties, GPT-4 agreed with expert labellers 85% of the time against 81% between humans
"The agreement under setup S2 (w/o tie) between GPT-4 and humans reaches 85%, which is even higher than the agreement among humans (81%)."
Zheng et al., Judging LLM-as-a-Judge (NeurIPS 2023 D&B), arXiv 2306.05685v4, §4.2 + Table 5 · page-text offset 30.7% · arxiv.org
LinkReport an errorCite
Zheng et al., Judging LLM-as-a-Judge (NeurIPS 2023 D&B), arXiv 2306.05685v4, §4.2 + Table 5, page-text offset 30.7%, https://arxiv.org/pdf/2306.05685v4. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L7, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L7 (source read 2026-09-26). -
First turn, with ties counted (setup S1): GPT-4 grading pairs agreed with humans 66%, GPT-4 grading single answers 60%, humans with each other 63%; without ties (S2) both GPT-4 modes 85% against 81%
"G4-Pair 70% 1138 66% 1343 97% 662 85% 859 G4-Single - 60% 1280 - 85% 739 Human - 63% 721 - 81% 479 (a) First Turn"
Zheng et al. (NeurIPS 2023 D&B), arXiv 2306.05685v4, Table 5(a) · page-text offset 34.1% · arxiv.org
LinkReport an errorCite
Zheng et al. (NeurIPS 2023 D&B), arXiv 2306.05685v4, Table 5(a), page-text offset 34.1%, https://arxiv.org/pdf/2306.05685v4. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L7a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L7a (source read 2026-09-26). -
In 2023 the MT-Bench authors proposed a reference-guided grader that first answers the question itself and uses its own answer as the reference; on their maths questions it cut the failure rate from 70% to 15%
"we first generate LLM judge's answer independently, and then display it as a reference answer in the judge prompt. In Table 4, we see a significant improvement in failure rate (from 70% to 15%) over the default prompt"
Zheng et al. (NeurIPS 2023 D&B), arXiv 2306.05685v4, §3.4 · page-text offset 27.1% · arxiv.org
LinkReport an errorCite
Zheng et al. (NeurIPS 2023 D&B), arXiv 2306.05685v4, §3.4, page-text offset 27.1%, https://arxiv.org/pdf/2306.05685v4. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L7b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L7b (source read 2026-09-26). -
The MT-Bench authors saw GPT-4 misjudge an answer to a maths problem it could solve when asked separately: it was misled by the answers it was shown
"although GPT-4 can solve the problem (when asked separately), it was misled by the provided answers, ultimately resulting in incorrect judgment"
Zheng et al. (NeurIPS 2023 D&B), arXiv 2306.05685v4, §3.3 · page-text offset 24.1% · arxiv.org
LinkReport an errorCite
Zheng et al. (NeurIPS 2023 D&B), arXiv 2306.05685v4, §3.3, page-text offset 24.1%, https://arxiv.org/pdf/2306.05685v4. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L7c, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L7c (source read 2026-09-26). -
In 2023 the Prometheus authors found that removing the reference answer hurt their grader most, and said a reference relieves the grader of solving the question itself
"relieves the need for the evaluator LM to internally solve the instruction and only focus on assessing the response"
Kim et al., Prometheus (ICLR 2024, per the arXiv comments field), arXiv 2310.08491 · page-text offset 45.7% · arxiv.org
LinkReport an errorCite
Kim et al., Prometheus (ICLR 2024, per the arXiv comments field), arXiv 2310.08491, page-text offset 45.7%, https://arxiv.org/pdf/2310.08491. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L26, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L26 (source read 2026-09-26). -
Re-tested on expert labels with a leave-one-annotator-out test, no LLM grader passed on MT-Bench, one of two datasets where none did
"in two datasets (MT-Bench, and SummEval), none of the LLMs pass the test"
Calderon, Reichart, Dror, The Alternative Annotator Test (ACL 2025), arXiv 2501.10970v4, §5 · page-text offset 21.8% · arxiv.org
LinkReport an errorCite
Calderon, Reichart, Dror, The Alternative Annotator Test (ACL 2025), arXiv 2501.10970v4, §5, page-text offset 21.8%, https://arxiv.org/pdf/2501.10970v4. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L8, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L8 (source read 2026-09-26). -
In Chatbot Arena's expert relabel, GPT-4 agreed with the two experts 81.0% and 78.5% (Llama-2-13b battles) and 76.3% and 79.3% (GPT-3.5-Turbo battles); the experts agreed with each other 89.8% and 79.4%
"Llama-2-13b Expert 1 Expert 2 GPT-4 Crowd 72.8% 77.8% 75.6% Expert 1 - 89.8% 81.0% Expert 2 - - 78.5% GPT-3.5-Turbo Expert 1 Expert 2 GPT-4 Crowd 73.8% 83.1% 75.6% Expert 1 - 79.4% 76.3% Expert 2 - - 79.3%"
Chiang et al., Chatbot Arena (ICML 2024), arXiv 2403.04132v1, Table 3 · page-text offset 35.9% · arxiv.org
LinkReport an errorCite
Chiang et al., Chatbot Arena (ICML 2024), arXiv 2403.04132v1, Table 3, page-text offset 35.9%, https://arxiv.org/pdf/2403.04132v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L9, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L9 (source read 2026-09-26). -
On literature-grounded science answers, the best judge, o3, reached 65.1% accuracy
"Even the best-performing model, o3, achieves only 65.1% accuracy."
Zhao et al., SciArena (NeurIPS 2025 D&B), arXiv 2507.01001v2, §6.2 + Table 1 · page-text offset 32.6% · arxiv.org
LinkReport an errorCite
Zhao et al., SciArena (NeurIPS 2025 D&B), arXiv 2507.01001v2, §6.2 + Table 1, page-text offset 32.6%, https://arxiv.org/pdf/2507.01001v2. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L10, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L10 (source read 2026-09-26). -
Random guessing on the same task scores 50.0
"Random Guess 50.0 o3 65.1"
Zhao et al., SciArena (NeurIPS 2025 D&B), arXiv 2507.01001v2, Table 3 · page-text offset 32.3% · arxiv.org
LinkReport an errorCite
Zhao et al., SciArena (NeurIPS 2025 D&B), arXiv 2507.01001v2, Table 3, page-text offset 32.3%, https://arxiv.org/pdf/2507.01001v2. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L10a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L10a (source read 2026-09-26). -
Expert annotators' average inter-annotator agreement: accuracy 0.82, kappa 0.76
"Average 0.82 0.76 0.94 0.91"
Zhao et al., SciArena, arXiv 2507.01001v2, Table 1 · page-text offset 21.9% · arxiv.org
LinkReport an errorCite
Zhao et al., SciArena, arXiv 2507.01001v2, Table 1, page-text offset 21.9%, https://arxiv.org/pdf/2507.01001v2. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L10b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L10b (source read 2026-09-26). -
The best frontier judge reached physician agreement on 3 of 5 clinical grading tasks (NEJM Healer, BIDMC ER, Landmark) and fell short on two; no single model matched on all five
"No single base model matched physician agreement across all five tasks."
Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results · page-text offset 15.1% · numbers present in text: 82%, 77%, 91%, 88%, 68%, 67%, 87%, 92%, 84%, 95% · arxiv.org
LinkReport an errorCite
Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results, page-text offset 15.1% · numbers present in text: 82%, 77%, 91%, 88%, 68%, 67%, 87%, 92%, 84%, 95%, https://arxiv.org/pdf/2609.12822v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L11, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L11 (source read 2026-09-26). -
Best judge vs physician agreement per task: NEJM Healer 82% vs 77%, BIDMC ER 91% vs 88%, Landmark 68% vs 67%; short on NEJM CPCs 87% vs 92% and Grey Matters Management 84% vs 95%
"reached physician inter-rater agreement on the NEJM Healer (Claude, 82% vs. 77% for physicians), BIDMC ER (Claude, 91% vs. 88%), and Landmark Diagnostic Cases (Gemini, 68% vs. 67%), but fell short on the NEJM CPCs (GPT-5, 87% vs. 92%) and the Grey Matters Management Cases (Claude, 84% vs. 95%)"
Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results · page-text offset 14.5% · arxiv.org
LinkReport an errorCite
Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results, page-text offset 14.5%, https://arxiv.org/pdf/2609.12822v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L11a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L11a (source read 2026-09-26). -
In the clinical study, some tasks score a diagnosis against the known answer (the Bond score) and others score diagnostic reasoning against a rubric
"A score of 5 indicates that the correct diagnosis is included in the differential"
Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results, Methods, Task Rubrics · page-text offset 48.9% · arxiv.org
LinkReport an errorCite
Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results, Methods, Task Rubrics, page-text offset 48.9%, https://arxiv.org/pdf/2609.12822v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L11b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L11b (source read 2026-09-26). -
The authors' own grader, PrecepTron, fine-tuned on a small number of physician-scored cases, matched or exceeded physician agreement on three of five tasks, so no single model, base or fine-tuned, matched on all five (the abstract's 'physician-level consistent scoring across tasks' is looser than the Results and must not be cited)
"PrecepTron approaches physician inter-rater agreement, matching or exceeding it on three of five tasks (NEJM CPCs, NEJM Healer, and BIDMC ER)"
Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results (PrecepTron section) · Results · arxiv.org
LinkReport an errorCite
Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results (PrecepTron section), Results, https://arxiv.org/pdf/2609.12822v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L11c, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L11c (source read 2026-09-27). -
The GPT-4.1 grader using physician-written criteria scored above the average physician in five of seven themes (vendor-authored)
"GPT-4.1 as a grader exceeds the random baseline for all themes as shown in Table 5. It exceeds the average physician score in five out of seven themes"
Arora et al. (OpenAI), HealthBench, arXiv 2505.08775v1, §8.1 + Table 5 · page-text offset 42.6% · arxiv.org
LinkReport an errorCite
Arora et al. (OpenAI), HealthBench, arXiv 2505.08775v1, §8.1 + Table 5, page-text offset 42.6%, https://arxiv.org/pdf/2505.08775v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L12, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L12 (source read 2026-09-26). -
Using the expert rubric, five of seven graders passed between 38.0% and 45.7% of the proofs experts failed and wrongly failed between 5.0% and 12.3% of the proofs experts passed; the other two passed 63.6% and 74.8% of the failed proofs (false fails 3.8% and 2.5%)
"82.7% 43.0% 5.2% 82.0% 45.7% 5.0% 81.8% 38.0% 8.8% 80.9% 41.4% 8.7% 78.9% 39.6% 12.3% 77.0% 63.6% 3.8% 74.2% 74.8% 2.5%"
Gonzalez et al., QEDBench (ICML 2026), arXiv 2602.20629v3, Fig. 5 · page-text offset 5.4% · arxiv.org
LinkReport an errorCite
Gonzalez et al., QEDBench (ICML 2026), arXiv 2602.20629v3, Fig. 5, page-text offset 5.4%, https://arxiv.org/pdf/2602.20629v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L13, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L13 (source read 2026-09-26). -
The paper defines leniency as the false-positive rate and harshness as the false-negative rate against the expert rubric pass mark
"We decompose errors into Leniency Rate(False Positives) and Harshness Rate(False"
Gonzalez et al., QEDBench, Fig. 5 caption · page-text offset 5.5% · arxiv.org
LinkReport an errorCite
Gonzalez et al., QEDBench, Fig. 5 caption, page-text offset 5.5%, https://arxiv.org/pdf/2602.20629v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L13a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L13a (source read 2026-09-26). -
The paper reports significant positive bias for certain frontier graders it names (Claude Opus 4.5, DeepSeek-V3, Qwen 2.5 Max, Llama 4 Maverick); it does not say this of all graders
"frontier evaluators like Claude Opus 4.5, DeepSeek-V3, Qwen 2.5 Max, and Llama 4 Maverick exhibit significant positive bias"
Gonzalez et al., QEDBench, abstract · page-text offset 0.5% · arxiv.org
LinkReport an errorCite
Gonzalez et al., QEDBench, abstract, page-text offset 0.5%, https://arxiv.org/pdf/2602.20629v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L13b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L13b (source read 2026-09-26). -
QEDBench's leniency and harshness figures are for graders using the expert rubric
"Judge Reliability Metrics (Expert Rubric Only)"
Gonzalez et al., QEDBench (ICML 2026), arXiv 2602.20629v3, Fig. 5 label · page-text offset 5.4% · arxiv.org
LinkReport an errorCite
Gonzalez et al., QEDBench (ICML 2026), arXiv 2602.20629v3, Fig. 5 label, page-text offset 5.4%, https://arxiv.org/pdf/2602.20629v3. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L13c, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L13c (source read 2026-09-26). -
Order-swap flip rate (MT-Bench / JudgeBench): Gemini 3.1 Pro 0.035 / 0.020; Claude Opus 4.6 0.038 / 0.022
"Gemini 3.1 Pro 0.977 0.989 0.035 0.978 0.989 0.020 Claude Sonnet 4 0.960 0.976 0.074 0.941 0.966 0.146 GPT-4o-mini 0.959 0.975 0.127 0.896 0.939 0.380 Claude Opus 4.6 0.958 0.979 0.038 0.974 0.986 0.022"
Norman, Rivera, Hughes, Reliability without Validity (preprint), arXiv 2606.19544v1, App. Table 6 · page-text offset 78.0% · arxiv.org
LinkReport an errorCite
Norman, Rivera, Hughes, Reliability without Validity (preprint), arXiv 2606.19544v1, App. Table 6, page-text offset 78.0%, https://arxiv.org/pdf/2606.19544v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L14, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L14 (source read 2026-09-26). -
Order-swap flip rate (MT-Bench / JudgeBench): GPT-5.4 0.114 / 0.105
"GPT-5.4 0.932 0.965 0.114 0.933 0.966 0.105"
Norman, Rivera, Hughes, Reliability without Validity (preprint), arXiv 2606.19544v1, App. Table 6 · page-text offset 78.8% · arxiv.org
LinkReport an errorCite
Norman, Rivera, Hughes, Reliability without Validity (preprint), arXiv 2606.19544v1, App. Table 6, page-text offset 78.8%, https://arxiv.org/pdf/2606.19544v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L14a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L14a (source read 2026-09-26). -
2026 JudgeBench, chance-corrected agreement (Cohen's kappa): Gemini 3.1 Pro 0.841, Claude Opus 4.6 0.875, Claude Sonnet 4.6 0.782, GPT-5.4 0.606; exact match 0.964, 0.956, 0.920, 0.812 (one order, ties excluded)
"Gemini 3.1 Pro 0.849 0.511 33.8 0.964 0.841 12.3 0.956 0.898 5.9 Claude Opus 4.6 0.848 0.489 35.9 0.956 0.875 8.1 0.943 0.879 6.4 DeepSeek V3.2 0.845 0.486 35.9 0.791 0.545 24.5 0.921 0.826 9.5 Claude Sonnet 4.6 0.851 0.484 36.7 0.920 0.782 13.8 0.942 0.871 7.1 Llama 3.3 70B 0.841 0.465 37.6 0.664 0.283 38.1 0.892 0.769 12.3 Kimi K2.5 0.846 0.461 38.5 0.864 0.720 14.5 0.937 0.873 6.4 GPT-5.4 0.836 0.457 38.0 0.812 0.606"
Norman, Rivera, Hughes, Reliability without Validity (preprint), arXiv 2606.19544v1, §4.1 Table 2 · page-text offset 24.2% · arxiv.org
LinkReport an errorCite
Norman, Rivera, Hughes, Reliability without Validity (preprint), arXiv 2606.19544v1, §4.1 Table 2, page-text offset 24.2%, https://arxiv.org/pdf/2606.19544v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L15, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L15 (source read 2026-09-26). -
The same authors warn that raw agreement does not correct for chance and overstates how well a grader discriminates
"This family of metrics does not correct for chance"
Norman, Rivera, Hughes (preprint), arXiv 2606.19544v1, §2 · page-text offset 4.0% · arxiv.org
LinkReport an errorCite
Norman, Rivera, Hughes (preprint), arXiv 2606.19544v1, §2, page-text offset 4.0%, https://arxiv.org/pdf/2606.19544v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L15b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L15b (source read 2026-09-26). -
Human annotators preferred the markdown side 57% of the time; four of five judges preferred it 73%–97% on the same pairs
"human annotators prefer the markdown side only 57% of the time, while four of the five judges prefer it 73%-97% on the same pairs"
Soumik, bias mitigation in LLM-as-a-Judge (TMLR 06/2026), arXiv 2604.23178v2 · page-text offset 44.2% · arxiv.org
LinkReport an errorCite
Soumik, bias mitigation in LLM-as-a-Judge (TMLR 06/2026), arXiv 2604.23178v2, page-text offset 44.2%, https://arxiv.org/pdf/2604.23178v2. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L16, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L16 (source read 2026-09-26). -
The human comparison was two independent annotators on a 30-pair subsample: they preferred markdown 57% of the time on average, while four of the five judges preferred it 73%-97%
"two independent annotators on a 30-pair subsample preferred markdown only 57% of the time on average, while four of the five judges preferred markdown 73%-97%"
Soumik, bias mitigation in LLM-as-a-Judge (TMLR 06/2026), arXiv 2604.23178v2, §1 / App. G · page-text offset 28.5% · arxiv.org
LinkReport an errorCite
Soumik, bias mitigation in LLM-as-a-Judge (TMLR 06/2026), arXiv 2604.23178v2, §1 / App. G, page-text offset 28.5%, https://arxiv.org/pdf/2604.23178v2. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L16a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L16a (source read 2026-09-26). -
The style test compared markdown with plain prose
"STYLE (markdown vs. plain prose)"
Soumik, bias mitigation in LLM-as-a-Judge (TMLR 06/2026), arXiv 2604.23178v2, §3 · page-text offset 22.0% · arxiv.org
LinkReport an errorCite
Soumik, bias mitigation in LLM-as-a-Judge (TMLR 06/2026), arXiv 2604.23178v2, §3, page-text offset 22.0%, https://arxiv.org/pdf/2604.23178v2. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L16b, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L16b (source read 2026-09-26). -
One study finds most measured self-preference is judges being unsure on hard items (its own summary: 89.6%, substantially reducing but not eliminating the evidence)
"evaluator uncertainty accounts for an average of 89.6%"
Roytburg et al., Are LLM Evaluators Really Narcissists? (ICML 2026), arXiv 2601.22548v4, §1 · page-text offset 9.2% · numbers present in text: 89.6% · arxiv.org
LinkReport an errorCite
Roytburg et al., Are LLM Evaluators Really Narcissists? (ICML 2026), arXiv 2601.22548v4, §1, page-text offset 9.2% · numbers present in text: 89.6%, https://arxiv.org/pdf/2601.22548v4. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L17, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L17 (source read 2026-09-26). -
GPT-5 as judge: false-positive rate on its own wrong GSM8K answers 82.0 [72.4, 90.7], 67.3 points above its rate on others' wrong answers
"GPT-5 82.0[72.4, 90.7]+67.3"
Zhang et al. (Microsoft/MIT), Can We Trust LLM Judges (preprint), arXiv 2609.12002v1, App. B Table 10 (GSM8K) · page-text offset 87.4% · arxiv.org
LinkReport an errorCite
Zhang et al. (Microsoft/MIT), Can We Trust LLM Judges (preprint), arXiv 2609.12002v1, App. B Table 10 (GSM8K), page-text offset 87.4%, https://arxiv.org/pdf/2609.12002v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L18, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L18 (source read 2026-09-26). -
Across three QA sets (NQ, TQA, HPQA), the panel (PoLL) beat GPT-4 alone on all three (0.763 vs 0.627; 0.906 vs 0.841; 0.867 vs 0.830), edged its best member on two (0.763 vs 0.749; 0.906 vs 0.902) and fell below Haiku on the third (0.867 vs 0.873)
"GPT-4 0.627 0.841 0.830 CMD-R 0.734 0.902 0.815 Haiku 0.749 0.894 0.873 GPT-3.5 0.726 0.859 0.833 PoLL 0.763 0.906 0.867"
Verga et al. (Cohere), Replacing Judges with Juries, arXiv 2404.18796v2, Table 1 · page-text offset 18.5% · arxiv.org
LinkReport an errorCite
Verga et al. (Cohere), Replacing Judges with Juries, arXiv 2404.18796v2, Table 1, page-text offset 18.5%, https://arxiv.org/pdf/2404.18796v2. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L19, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L19 (source read 2026-09-26). -
When the judge is no better than the model it grades, correcting it with human labels can at most halve the labels you need
"no debiasing method can decrease the required amount of ground truth labels by more than half"
Dorner, Nastl, Hardt (ICLR 2025 Oral), arXiv 2410.13341v4, abstract · page-text offset 1.4% · arxiv.org
LinkReport an errorCite
Dorner, Nastl, Hardt (ICLR 2025 Oral), arXiv 2410.13341v4, abstract, page-text offset 1.4%, https://arxiv.org/pdf/2410.13341v4. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L21, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L21 (source read 2026-09-26). -
In the original JudgeBench (2024), many strong graders, GPT-4o among them, scored only slightly better than random on hard correctness pairs
"performing just slightly better than random guessing"
Tan et al., JudgeBench (ICLR 2025), arXiv 2410.12784v2, abstract · page-text offset 1.8% · arxiv.org
LinkReport an errorCite
Tan et al., JudgeBench (ICLR 2025), arXiv 2410.12784v2, abstract, page-text offset 1.8%, https://arxiv.org/pdf/2410.12784v2. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L22, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L22 (source read 2026-09-26). -
JudgeBench found a grader's ability to verify answers highly correlated with its ability to solve the problem itself
"the ability of the judge to verify the solution pairs is highly correlated with its ability to solve the problem itself"
Tan et al., JudgeBench (ICLR 2025), arXiv 2410.12784v2, §4.4 · page-text offset 41.3% · arxiv.org
LinkReport an errorCite
Tan et al., JudgeBench (ICLR 2025), arXiv 2410.12784v2, §4.4, page-text offset 41.3%, https://arxiv.org/pdf/2410.12784v2. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L22a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L22a (source read 2026-09-26). -
PRIOR ART: Eugene Yan's August 2024 review of LLM-evaluators drew on about two dozen papers and already covers reference-based evaluation and position bias
"Drawing from two dozen papers"
Eugene Yan, Evaluating the Effectiveness of LLM-Evaluators (aka LLM-as-Judge), Aug 2024 · page-text offset 1.5% · eugeneyan.com
LinkReport an errorCite
Eugene Yan, Evaluating the Effectiveness of LLM-Evaluators (aka LLM-as-Judge), Aug 2024, page-text offset 1.5%, https://eugeneyan.com/writing/llm-evaluators/. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L23, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L23 (source read 2026-09-26). -
PRIOR ART: Hamel Husain's guide (Oct 2024) has a principal domain expert make pass/fail judgments and tracks the judge's agreement with them; the warning that agreement misleads on imbalanced data was added in a 2025 revision
"I also tracked agreement rates over time to ensure we were converging on a good prompt"
Hamel Husain, Using LLM-as-a-Judge For Evaluation: A Complete Guide (Oct 2024) · page-text offset 60.7% · hamel.dev
LinkReport an errorCite
Hamel Husain, Using LLM-as-a-Judge For Evaluation: A Complete Guide (Oct 2024), page-text offset 60.7%, https://hamel.dev/blog/posts/llm-judge/. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L24, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L24 (source read 2026-09-26). -
Husain's guide, as revised in September 2026, advises about 100 examples per failure mode to validate an automated grader and says below 60 the confidence intervals are often too wide; the October 2024 original did not contain this advice
"Aim for about 100 examples per failure mode, with enough Pass and Fail examples to measure both classes. Below 60 examples, the confidence intervals are often too wide"
Hamel Husain, Using LLM-as-a-Judge For Evaluation (Oct 2024) · page-text offset 50.8% · hamel.dev
LinkReport an errorCite
Hamel Husain, Using LLM-as-a-Judge For Evaluation (Oct 2024), page-text offset 50.8%, https://hamel.dev/blog/posts/llm-judge/. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L24a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L24a (source read 2026-09-26). -
In code review against a written requirement, GPT-4o wrongly failed correct code (false-negative rate) 26.2% (HumanEval) and 35.9% (MBPP) when judging directly, rising to 73.2% and 87.9% when also asked to explain and repair
"GPT-4o achieves a relatively low FNR in HumanEval (26.2%) and MBPP (35.9%), but once explana- tions and repairs are required, the FNR increases sharply to 73.2% in HumanEval and 87.9% in MBPP"
Jin & Chen, Are LLMs Reliable Code Reviewers? Systematic Overcorrection (Automated Software Engineering 33, art. 90, 26 Jun 2026; read as arXiv 2603.00539v1, and the published text prints the same sentence), §5.2 · page-text offset 38.7% · numbers present in text: 26.2, 35.9, 73.2, 87.9 · arxiv.org
LinkReport an errorCite
Jin & Chen, Are LLMs Reliable Code Reviewers? Systematic Overcorrection (Automated Software Engineering 33, art. 90, 26 Jun 2026; read as arXiv 2603.00539v1, and the published text prints the same sentence), §5.2, page-text offset 38.7% · numbers present in text: 26.2, 35.9, 73.2, 87.9, https://arxiv.org/pdf/2603.00539v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L25, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L25 (source read 2026-09-26). -
In that paper, FNR is rejecting correct code (a false fail) and FPR is accepting buggy code (a false pass)
"FNR reflects over-correction (rejecting correct implementations), while FPR reflects unsafe acceptance"
Jin & Chen, Are LLMs Reliable Code Reviewers? Systematic Overcorrection (Automated Software Engineering 33, art. 90, 2026; read as arXiv 2603.00539v1), Table 2 caption · page-text offset 37.2% · arxiv.org
LinkReport an errorCite
Jin & Chen, Are LLMs Reliable Code Reviewers? Systematic Overcorrection (Automated Software Engineering 33, art. 90, 2026; read as arXiv 2603.00539v1), Table 2 caption, page-text offset 37.2%, https://arxiv.org/pdf/2603.00539v1. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row L25a, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L25a (source read 2026-09-26). -
The prior registered before any paper was read: graders would hold on easy, preference-style chat grading and fail on correctness of hard items they cannot solve themselves and in specialised domains
Result: judges HOLD on easy, preference-style chat grading and FAIL on (a) correctness of hard items the judge cannot itself solve and (b) specialised domains
first-party: pre-registration, 19:57 BST 2026-09-26
A run or observation done for this piece. The run record is not published.
LinkReport an errorCite
A run or observation by The Durability Curve (first-party: pre-registration, 19:57 BST 2026-09-26). Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row P1, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-P1 (source read 2026-09-26). -
As run, scored to the precision each question asked for (one pass, n = 24 arithmetic-style questions with one right answer each, listed in Line/judge-check/items.py, 2026-09-26 21:02 BST): Claude Haiku 4.5, answering without tools, got 4 of 24 wrong. Haiku 4.5 grading those answers without a key passed 3 of its 4 wrong answers; given the verified key, it passed 0 of 4
Result: HEADLINE haiku passed 3/4 of its own wrong answers without a key, 0/4 with one
first-party: Judge Check run
A run or observation done for this piece. The run record is not published.
LinkReport an errorCite
A run or observation by The Durability Curve (first-party: Judge Check run). Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row J1, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J1 (source read 2026-09-26). -
Claude Sonnet 5 grading the same answers without a key agreed with the key on 23 of 24: no false passes, one false fail (a correct final answer reached by an unjustified leap, which it failed although told to grade the final answer); with the key, 24 of 24
Result: no key agree 23/24 · false passes 0 · false fails 1 · unparsed 0 / with key agree 24/24
first-party: Judge Check run
A run or observation done for this piece. The run record is not published.
LinkReport an errorCite
A run or observation by The Durability Curve (first-party: Judge Check run). Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row J2, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J2 (source read 2026-09-26). -
Re-dressing the same answers in markdown changed 2 of 24 of Haiku's verdicts (both from pass to fail on wrong answers) and none of Sonnet's
Result: markdown re-dress changed 2/24 verdicts / markdown re-dress changed 0/24 verdicts
first-party: Judge Check run
A run or observation done for this piece. The run record is not published.
LinkReport an errorCite
A run or observation by The Durability Curve (first-party: Judge Check run). Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row J3, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J3 (source read 2026-09-26). -
As run (the precision each question asked for): Haiku as judge could not solve 2 of the 24 items itself and, without a key, passed the wrong answer on both; on H6 (digit sum of 50!) it passed a wrong answer although it solved the item correctly when asked. Under the registered rule it could not solve 1 item (H11) and passed the wrong answer on it; H6 holds under both rules
Result: solved it itself: 22/24 · no-key agree on solved 21/22 · on unsolved 0/2
first-party: Judge Check run
A run or observation done for this piece. The run record is not published.
LinkReport an errorCite
A run or observation by The Durability Curve (first-party: Judge Check run). Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row J4, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J4 (source read 2026-09-26). -
The run cost $2.83. As run, predictions P1, P4 and P5 passed, P3 passed for Haiku, P2 and Sonnet's P3 were unscorable. Under the pre-registered scoring rule (score.py --registered), P1 failed for both graders, P5 failed, Sonnet's P3 failed, P4 passed, Haiku's P3 passed
Result: cost $2.8326 · Predictions (pre-registered)
first-party: Judge Check run
A run or observation done for this piece. The run record is not published.
LinkReport an errorCite
A run or observation by The Durability Curve (first-party: Judge Check run). Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row J5, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J5 (source read 2026-09-26). -
Of 19 studies we found that report which way AI graders err against human or ground-truth labels, 12 lean toward passing flawed work (4 of them weakly, adversarially or in one direction only), 3 toward failing sound work (code review against a spec; older-model essay grading; a frontier panel on clinical diagnoses) and 4 are mixed
Result: Tally (Part 2, excluding 7b): Supports 12 (1, 2, 3, 4, 5, 6, 9, 13, 14, 16, 18, 19; of these, 13, 16, 18 and 19 are weak, adversarial or one-directional) · Contradicts 3 (7, 8, 11) · Mixed 4 (10, 12, 15, 17)
first-party: error-direction sweep, 2026-09-26
A run or observation done for this piece. The run record is not published.
LinkReport an errorCite
A run or observation by The Durability Curve (first-party: error-direction sweep, 2026-09-26). Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row T1, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-T1 (source read 2026-09-26). -
The evidence base is 77 distinct studies: 63 in the four sweeps and 14 more in the error-direction check
Result: PAPERS TOTAL 77 distinct studies
first-party: paper count
A run or observation done for this piece. The run record is not published.
LinkReport an errorCite
A run or observation by The Durability Curve (first-party: paper count). Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row T2, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-T2 (source read 2026-09-26). -
Under the pre-registered scoring rule (0.001 relative tolerance), Haiku's answers were wrong on 2 of 24; Haiku grading them passed both wrong answers without a key and neither with one; both graders, given the key, failed the two answers that missed the precision the question asked for
Result: HEADLINE haiku passed 2/2 of its own wrong answers without a key, 0/2 with one
first-party: Judge Check run, rescored
A run or observation done for this piece. The run record is not published.
LinkReport an errorCite
A run or observation by The Durability Curve (first-party: Judge Check run, rescored). Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row J6, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J6 (source read 2026-09-26). -
Our earlier piece A Green Score Is Not Evidence reported a groundedness metric that scored its best with the evidence removed
Result: A groundedness metric scored its best with the evidence removed.
first-party: published article (subtitle) · harryfloyd.substack.com
A run or observation done for this piece. The run record is not published.
LinkReport an errorCite
A run or observation by The Durability Curve (first-party: published article (subtitle)), https://harryfloyd.substack.com/p/a-green-score-is-not-evidence. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row W1, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-W1 (source read 2026-09-26). -
Our earlier piece Your AI Looks Best Where You Can Check It Least argued that the work that is hardest to check is where AI output looks best
Result: Your AI Looks Best Where You Can Check It Least
first-party: published article (title) · harryfloyd.substack.com
A run or observation done for this piece. The run record is not published.
LinkReport an errorCite
A run or observation by The Durability Curve (first-party: published article (title)), https://harryfloyd.substack.com/p/your-ai-looks-best-where-you-check-least. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row W2, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-W2 (source read 2026-09-26).
No rows match. Clear the search or pick another label.
Cite and reuse
- A single claim
- Use the row's Cite button. It names the original source first, then this row, which carries its own link.
- This ledger
Floyd, Harry (2026). Claim Ledger: Your AI Grader Is Only as Good as Its Answer Key. The Durability Curve. https://durabilitycurve.com/claims/your-ai-grader-answer-key/- The essay
Floyd, Harry (2026). Your AI Grader Is Only as Good as Its Answer Key. The Durability Curve. https://durabilitycurve.com/blog/your-ai-grader-answer-key/
The ledger data is licensed CC BY 4.0: reuse it, quote it or train on it, with credit to The Durability Curve and a link. Read the licence. Machine copies: JSON, CSV, Markdown.