---
title: "Claim Ledger: Your AI Grader Is Only as Good as Its Answer Key"
description: "I went through 77 studies of AI graders. Without a verified answer key, they are least reliable on exactly the questions they can't answer themselves. Here is a twenty-item check for yours."
author: "Harry Floyd"
publication: "The Durability Curve"
canonical: "https://durabilitycurve.com/claims/your-ai-grader-answer-key/"
essay: "https://durabilitycurve.com/blog/your-ai-grader-answer-key/"
substack: "https://harryfloyd.substack.com/p/your-ai-grader-answer-key"
published: "2026-09-28"
ledger_date: "2026-09-27"
claims: 68
removed_before_publication: 0
---

# Claim Ledger: Your AI Grader Is Only as Good as Its Answer Key

*I went through 77 studies of AI graders. Without a verified answer key, they are least reliable on exactly the questions they can't answer themselves. Here is a twenty-item check for yours.*

Essay: https://durabilitycurve.com/blog/your-ai-grader-answer-key/  
Ledger (canonical, cite this): https://durabilitycurve.com/claims/your-ai-grader-answer-key/  
Published 2026-09-28 · last verified 2026-09-27

## What the essay claims

On correctness, the variable that decides whether an AI judge can be trusted is not which model grades but whether it has a verified reference to grade against. Without one, judges agree with experts mostly on the questions they could answer themselves, and fall apart on the rest; with one, even small models grade close to how well humans agree with each other. When judges are wrong about correctness, they mostly pass flawed work (12 of 19 studies that measure direction; code review is the counter-case). On matters of taste and expert judgement, where no key exists, judges more often fall short of expert–expert agreement than match it. The popular fixes (panels, reasoning first, fine-tuned judges, calibration on a few human labels) help less than claimed. The test that screens it for any one setup is the reader's own: twenty items, with and without a key, split by whether the judge could solve them.

## The claim ladder

| Rung | Claim | Evidence (ledger) | Whose behaviour | What it does NOT reach |
|---|---|---|---|---|
| R1 | With a verified key, judges grade hard correctness close to human level; a self-made or subtly wrong key fails | L1–L5 | Judges on expert-graded finance/maths (GPT-4o + 4 open models); 2026 frontier judges on a proxy label | Subjective grading (no key exists); 2026 rows are corroboration, not human-labelled |
| R2 | Without a key, judges agree with experts mostly on questions they could answer themselves | L1, L5 | Same | Whether the newest judges (Opus 5.5, GPT-6) still show it: untested in the literature |
| R3 | On taste and expert judgement, judges more often fall short of expert–expert agreement | L7–L12 | Chat preference, science answers, clinical grading | Mixed: HealthBench 5/7 themes at or above the average physician (vendor-authored); GRAND-ROUNDS 3/5 |
| R4 | When wrong about correctness, judges mostly pass flawed work | L6, L13 + direction sweep (12/19) | Judges grading answers and proofs | Reverses in code review against a spec; not a law |
| R5 | Biases are judge-specific | L14–L16 | Named 2025–26 judges on MT-Bench / JudgeBench / style pairs | Human baseline for markdown is n = 2 annotators |
| R6 | Self-preference is contested | L17–L18 | Judges grading their own vs others' answers | Unresolved between the two papers; say so |
| R7 | Popular fixes help less than claimed | L19–L21 | Panels, reasoning-first, human-label calibration | Panels can edge their best member (L19); some reasoning helps on hard pairs |
| R8 | The Judge Check can catch a generous grader on twenty items; it cannot clear one (about 100 per failure mode, L24a) | Object Record (our run, n = 24) | The check, on Haiku 4.5 and Sonnet 5 grading Haiku 4.5 | A demonstration, not evidence; one run, one task type |

## The evidence, row by row

Labels: VERIFIED = checked against the primary source itself by a checker who did not draft the piece, usually with the exact words, where they sit and the date read · EXECUTED = a run or observation done for the piece (code, a query, a count, or a check made live in an app); the claim is what it returned; run records are not published · CHECKED = checked against its source by a separate checker in the older ledger format, which recorded no quote, locator or date · REPORTED = not verified word for word against a primary source (secondary, not openable in full, or supported only in part); the essay words it accordingly · STRUCK = drafted, checked and removed before publication · EXCLUDED = considered and deliberately left out.

### L1. Without a correct reference, judges agreed well with human experts only on questions they had been able to answer themselves

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L1
- Quote: "when not provided with a correct reference, judges show high agreement with human experts only on questions the judges were able to correctly answer themselves"
- Source: Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, abstract, page-text offset 1.2%
- URL: https://arxiv.org/pdf/2503.05061v3

### L1a. GPT-4o as judge, Cohen's kappa with experts. Pairwise: questions it answered correctly 0.78 with no key, 0.92 with a human key; questions it got wrong 0.30 with no key, 0.16 with its own answer as the key, 0.83 with a human key. Single grading: questions it answered correctly 0.46 with no key and 0.59 with a human key; questions it got wrong 0.16 and 0.81

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L1a
- Quote: "GPT-4o Correct 0.78±0.02 0.86±0.02 0.92±0.02 0.46±0.07 0.52±0.06 0.59±0.06 Incorrect 0.30±0.09 0.16±0.08 0.83±0.06 0.16±0.14 0.13±0.14 0.81±0.13"
- Source: Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Table 4 (GPT-4o row), page-text offset 21.0%
- URL: https://arxiv.org/pdf/2503.05061v3

### L1b. Pairwise, questions each judge got wrong, no key vs human key: Llama 3.3 70B 0.39 vs 0.92; Phi 4 -0.08 vs 0.75; Qwen 2.5 7B 0.21 vs 0.63; Yi 1.5 34B 0.13 vs 0.56 (own answer as key: 0.35, -0.08, 0.14, 0.02)

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L1b
- Quote: "Llama 3.3 70b Correct 0.79±0.03 0.86±0.03 0.93±0.02 0.61±0.08 0.54±0.08 0.73±0.07 Incorrect 0.39±0.06 0.35±0.07 0.92±0.03 0.13±0.06 0.23±0.10 0.69±0.09 Phi 4 Correct 0.67±0.07 0.88±0.04 0.91±0.04 0.48±0.09 0.47±0.07 0.67±0.07 Incorrect -0.08±0.09 -0.08±0.10 0.75±0.06 0.03±0.05 0.02±0.12 0.56±0.12 Qwen 2.5 7B Correct 0.65±0.05 0.69±0.04 0.78±0.04 0.31±0.11 0.37±0.09 0.54±0.09 Incorrect 0.21±0.07 0.14±0.07 0.63±0.06 0.08±0.07 0.21±0.10 0.56±0.08 Yi 1.5 34B Correct 0.31±0.12 0.53±0.12 0.77±0.09 0.23±0.11 0.29±0.08 0.43±0.09 Incorrect 0.13±0.09 0.02±0.09 0.56±0.08"
- Source: Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Table 4 (other judges), page-text offset 21.2%
- URL: https://arxiv.org/pdf/2503.05061v3

### L1c. The paper notes that a relatively small model (Qwen 2.5 7B) with a human reference can grade better than a larger model (GPT-4o) without one, on the same set of responses

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L1c
- Quote: "providing a human reference to a relatively small model (such as Qwen 2.5 7B ) can yield better judgments than using a larger model without human references (such asGPT-4o) on the same set of responses"
- Source: Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6, page-text offset 22.2%
- URL: https://arxiv.org/pdf/2503.05061v3

### L2. With its own answer as the key, GPT-4o's pairwise kappa was 0.86 on questions it had answered correctly and 0.16 on questions it could not answer

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L2
- Quote: "decreases from 0.86 to 0.16 in the pairwise case on questions GPT-4o could not answer"
- Source: Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6, page-text offset 23.0%
- URL: https://arxiv.org/pdf/2503.05061v3

### L3. Averaged across all five graders on questions GPT-4o answered correctly (Table 3 does not say whether single and pairwise grading are pooled): kappa with a human key 0.69, a verified GPT-4o answer as key 0.61, a completely unrelated key 0.50, no key 0.46, a subtly wrong key 0.21

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3
- Quote: "Human 0.69 ±0.06 0.74±0.06 0.57±0.08 GPT-4o (✓) 0.61 ±0.07 0.66±0.07 0.52±0.09 None 0.46 ±0.08 0.56±0.08 0.22±0.09 Random 0.50 ±0.08 0.47±0.08 0.09±0.1 Wrong 0.21 ±0.06"
- Source: Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Table 3, page-text offset 19.3%
- URL: https://arxiv.org/pdf/2503.05061v3

### L3c. Table 3's figures are across all five graders

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3c
- Quote: "Table 3 presents results for these reference types across all five judges"
- Source: Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6, page-text offset 26.3%
- URL: https://arxiv.org/pdf/2503.05061v3

### L3a. The paper says a slightly incorrect reference can in some cases be worse than no reference at all

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3a
- Quote: "in some cases be worse than"
- Source: Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6, page-text offset 25.4%
- URL: https://arxiv.org/pdf/2503.05061v3

### L3d. The authors describe the Random key as a completely unrelated reference, and say a slightly incorrect reference can in some cases be worse than a completely unrelated one or none

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3d
- Quote: "providing a slightly incorrect reference can in some cases be worse than providing a completely unrelated reference or no reference at all"
- Source: Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6, page-text offset 25.4%
- URL: https://arxiv.org/pdf/2503.05061v3

### L3b. The authors conclude that verifying a model-generated answer, rather than writing and checking one by hand, was sufficient as a key in their setting

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3b
- Quote: "Verifying responses generated by a model rather than manually writing and checking answers via annotators is sufficient in this setting"
- Source: Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §7, page-text offset 28.6%
- URL: https://arxiv.org/pdf/2503.05061v3

### L4. 2026 judges, questions each got wrong, against a stand-in label: single grading no key vs human key, Opus 4.7 0.26 vs 0.62, GPT-5.4 0.32 vs 0.67, Gemini 3.1 0.44 vs 0.71; pairwise 0.33 vs 0.66, 0.46 vs 0.87, 0.67 vs 0.95 (own answer as key, pairwise: 0.33, 0.45, 0.62)

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4
- Quote: "Opus 4.7 Correct 0.86±0.04 0.90±0.04 0.91±0.04 0.63±0.06 0.69±0.06 0.77±0.05 Incorrect 0.33±0.16 0.33±0.15 0.66±0.13 0.26±0.14 0.27±0.12 0.62±0.11 GPT-5.4 Correct 0.75±0.06 0.87±0.04 0.89±0.04 0.51±0.06 0.64±0.06 0.66±0.06 Incorrect 0.46±0.11 0.45±0.13 0.87±0.08 0.32±0.11 0.25±0.12 0.67±0.09 Gemini 3.1 Correct 0.78±0.06 0.80±0.06 0.83±0.05 0.55±0.06 0.62±0.06 0.69±0.06 Incorrect 0.67±0.13 0.62±0.14 0.95±0.07 0.44±0.13 0.35±0.14 0.71±0.10"
- Source: Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, Table 12, page-text offset 90.1%
- URL: https://arxiv.org/pdf/2503.05061v3

### L4a. The authors treat the 2026-judge rows, scored against a stand-in label, as corroboration

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4a
- Quote: "we regard these results as corroboration rather than equivalent to our main table"
- Source: Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, page-text offset 88.7%
- URL: https://arxiv.org/pdf/2503.05061v3

### L4b. The stand-in label is GPT-4o grading with the human-verified references

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4b
- Quote: "We used the GPT-4o Judge with the human-verified references as a proxy correctness label"
- Source: Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, page-text offset 88.4%
- URL: https://arxiv.org/pdf/2503.05061v3

### L4c. The authors report that higher reasoning effort can narrow the no-reference gap, for example GPT-5.4 at high effort

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4c
- Quote: "higher reasoning effort can narrow the no-reference gap (e.g.,GPT-5.4 at high effort)"
- Source: Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, page-text offset 88.9%
- URL: https://arxiv.org/pdf/2503.05061v3

### L4d. GPT-5.4 at high reasoning effort, on questions it got wrong, against the stand-in label: pairwise 0.68 with no key and 0.84 with the human key; single grading 0.38 and 0.67

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4d
- Quote: "GPT-5.4 low Correct 0.85±0.05 0.90±0.04 0.54±0.06 0.57±0.05 Incorrect 0.55±0.13 0.95±0.06 0.34±0.11 0.56±0.09 default Correct 0.75±0.06 0.89±0.04 0.51±0.06 0.66±0.06 Incorrect 0.46±0.11 0.87±0.08 0.32±0.11 0.67±0.09 high Correct 0.80±0.06 0.87±0.05 0.68±0.08 0.67±0.06 Incorrect 0.68±0.14 0.84±0.11 0.38±0.19 0.67±0.11"
- Source: Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, Table 13 (GPT-5.4 rows), page-text offset 91.4%
- URL: https://arxiv.org/pdf/2503.05061v3

### L5. On MATH, DeepSeek V3 as a judge without a reference reached Scott's pi 0.72 against ground truth; as a matcher with the reference, 0.98

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L5
- Quote: "DeepSeek v3 model achieves only modest agreement π = 0.72, while as a matcher, it achieves π = 0.98"
- Source: Chandak et al., Answer Matching Outperforms Multiple Choice (preprint), arXiv 2507.02856v1, §3.1, page-text offset 22.6%
- URL: https://arxiv.org/pdf/2507.02856v1

### L5a. A 1.7-billion-parameter Qwen3 model matching answers to the reference reached pi 0.97 against ground truth on MATH

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L5a
- Quote: "answer matching, even with the 1.7 billion parameter Qwen3 model (non-thinking mode), achieves near-perfect alignment with the ground-truth (π = 0.97)"
- Source: Chandak et al., Answer Matching Outperforms Multiple Choice (preprint), arXiv 2507.02856v1, §3.1, page-text offset 22.4%
- URL: https://arxiv.org/pdf/2507.02856v1

### L5b. Answer matching with recent models, even small ones, reached near-perfect agreement with human grading, in the range of inter-annotator agreement

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L5b
- Quote: "matching using recent models-even small ones-achieves near-perfect agreement, in the range of inter-annotator agreement"
- Source: Chandak et al., Answer Matching (preprint), arXiv 2507.02856v1, abstract, page-text offset 1.3%
- URL: https://arxiv.org/pdf/2507.02856v1

### L5c. On MATH, the 'true grades' come from MATH-Verify, a rule-based checker that compares each answer with the reference

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L5c
- Quote: "MATH-Verify library (Kydlicek et al., 2025) implements rule-based ground-truth evaluations of generative responses"
- Source: Chandak et al., Answer Matching (preprint), arXiv 2507.02856v1, §3.1, page-text offset 22.0%
- URL: https://arxiv.org/pdf/2507.02856v1

### L6. For frontier judges (DeepSeek V3, o4-mini), 80%+ of grading errors were false passes: the judge marked correct what human annotation marked incorrect

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L6
- Quote: "errors disproportionately (80%+) arise from false positives"
- Source: Chandak et al., arXiv 2507.02856v1, §3.2, page-text offset 28.5%
- URL: https://arxiv.org/pdf/2507.02856v1

### L7. On MT-Bench, without ties, GPT-4 agreed with expert labellers 85% of the time against 81% between humans

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L7
- Quote: "The agreement under setup S2 (w/o tie) between GPT-4 and humans reaches 85%, which is even higher than the agreement among humans (81%)."
- Source: Zheng et al., Judging LLM-as-a-Judge (NeurIPS 2023 D&B), arXiv 2306.05685v4, §4.2 + Table 5, page-text offset 30.7%
- URL: https://arxiv.org/pdf/2306.05685v4

### L7a. First turn, with ties counted (setup S1): GPT-4 grading pairs agreed with humans 66%, GPT-4 grading single answers 60%, humans with each other 63%; without ties (S2) both GPT-4 modes 85% against 81%

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L7a
- Quote: "G4-Pair 70% 1138 66% 1343 97% 662 85% 859 G4-Single - 60% 1280 - 85% 739 Human - 63% 721 - 81% 479 (a) First Turn"
- Source: Zheng et al. (NeurIPS 2023 D&B), arXiv 2306.05685v4, Table 5(a), page-text offset 34.1%
- URL: https://arxiv.org/pdf/2306.05685v4

### L7b. In 2023 the MT-Bench authors proposed a reference-guided grader that first answers the question itself and uses its own answer as the reference; on their maths questions it cut the failure rate from 70% to 15%

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L7b
- Quote: "we first generate LLM judge's answer independently, and then display it as a reference answer in the judge prompt. In Table 4, we see a significant improvement in failure rate (from 70% to 15%) over the default prompt"
- Source: Zheng et al. (NeurIPS 2023 D&B), arXiv 2306.05685v4, §3.4, page-text offset 27.1%
- URL: https://arxiv.org/pdf/2306.05685v4

### L7c. The MT-Bench authors saw GPT-4 misjudge an answer to a maths problem it could solve when asked separately: it was misled by the answers it was shown

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L7c
- Quote: "although GPT-4 can solve the problem (when asked separately), it was misled by the provided answers, ultimately resulting in incorrect judgment"
- Source: Zheng et al. (NeurIPS 2023 D&B), arXiv 2306.05685v4, §3.3, page-text offset 24.1%
- URL: https://arxiv.org/pdf/2306.05685v4

### L26. In 2023 the Prometheus authors found that removing the reference answer hurt their grader most, and said a reference relieves the grader of solving the question itself

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L26
- Quote: "relieves the need for the evaluator LM to internally solve the instruction and only focus on assessing the response"
- Source: Kim et al., Prometheus (ICLR 2024, per the arXiv comments field), arXiv 2310.08491, page-text offset 45.7%
- URL: https://arxiv.org/pdf/2310.08491

### L8. Re-tested on expert labels with a leave-one-annotator-out test, no LLM grader passed on MT-Bench, one of two datasets where none did

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L8
- Quote: "in two datasets (MT-Bench, and SummEval), none of the LLMs pass the test"
- Source: Calderon, Reichart, Dror, The Alternative Annotator Test (ACL 2025), arXiv 2501.10970v4, §5, page-text offset 21.8%
- URL: https://arxiv.org/pdf/2501.10970v4

### L9. In Chatbot Arena's expert relabel, GPT-4 agreed with the two experts 81.0% and 78.5% (Llama-2-13b battles) and 76.3% and 79.3% (GPT-3.5-Turbo battles); the experts agreed with each other 89.8% and 79.4%

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L9
- Quote: "Llama-2-13b Expert 1 Expert 2 GPT-4 Crowd 72.8% 77.8% 75.6% Expert 1 - 89.8% 81.0% Expert 2 - - 78.5% GPT-3.5-Turbo Expert 1 Expert 2 GPT-4 Crowd 73.8% 83.1% 75.6% Expert 1 - 79.4% 76.3% Expert 2 - - 79.3%"
- Source: Chiang et al., Chatbot Arena (ICML 2024), arXiv 2403.04132v1, Table 3, page-text offset 35.9%
- URL: https://arxiv.org/pdf/2403.04132v1

### L10. On literature-grounded science answers, the best judge, o3, reached 65.1% accuracy

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L10
- Quote: "Even the best-performing model, o3, achieves only 65.1% accuracy."
- Source: Zhao et al., SciArena (NeurIPS 2025 D&B), arXiv 2507.01001v2, §6.2 + Table 1, page-text offset 32.6%
- URL: https://arxiv.org/pdf/2507.01001v2

### L10a. Random guessing on the same task scores 50.0

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L10a
- Quote: "Random Guess 50.0 o3 65.1"
- Source: Zhao et al., SciArena (NeurIPS 2025 D&B), arXiv 2507.01001v2, Table 3, page-text offset 32.3%
- URL: https://arxiv.org/pdf/2507.01001v2

### L10b. Expert annotators' average inter-annotator agreement: accuracy 0.82, kappa 0.76

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L10b
- Quote: "Average 0.82 0.76 0.94 0.91"
- Source: Zhao et al., SciArena, arXiv 2507.01001v2, Table 1, page-text offset 21.9%
- URL: https://arxiv.org/pdf/2507.01001v2

### L11. The best frontier judge reached physician agreement on 3 of 5 clinical grading tasks (NEJM Healer, BIDMC ER, Landmark) and fell short on two; no single model matched on all five

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L11
- Quote: "No single base model matched physician agreement across all five tasks."
- Source: Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results, page-text offset 15.1% · numbers present in text: 82%, 77%, 91%, 88%, 68%, 67%, 87%, 92%, 84%, 95%
- URL: https://arxiv.org/pdf/2609.12822v1

### L11a. Best judge vs physician agreement per task: NEJM Healer 82% vs 77%, BIDMC ER 91% vs 88%, Landmark 68% vs 67%; short on NEJM CPCs 87% vs 92% and Grey Matters Management 84% vs 95%

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L11a
- Quote: "reached physician inter-rater agreement on the NEJM Healer (Claude, 82% vs. 77% for physicians), BIDMC ER (Claude, 91% vs. 88%), and Landmark Diagnostic Cases (Gemini, 68% vs. 67%), but fell short on the NEJM CPCs (GPT-5, 87% vs. 92%) and the Grey Matters Management Cases (Claude, 84% vs. 95%)"
- Source: Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results, page-text offset 14.5%
- URL: https://arxiv.org/pdf/2609.12822v1

### L11b. In the clinical study, some tasks score a diagnosis against the known answer (the Bond score) and others score diagnostic reasoning against a rubric

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L11b
- Quote: "A score of 5 indicates that the correct diagnosis is included in the differential"
- Source: Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results, Methods, Task Rubrics, page-text offset 48.9%
- URL: https://arxiv.org/pdf/2609.12822v1

### L11c. The authors' own grader, PrecepTron, fine-tuned on a small number of physician-scored cases, matched or exceeded physician agreement on three of five tasks, so no single model, base or fine-tuned, matched on all five (the abstract's 'physician-level consistent scoring across tasks' is looser than the Results and must not be cited)

- Status: VERIFIED (read 2026-09-27)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L11c
- Quote: "PrecepTron approaches physician inter-rater agreement, matching or exceeding it on three of five tasks (NEJM CPCs, NEJM Healer, and BIDMC ER)"
- Source: Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results (PrecepTron section), Results
- URL: https://arxiv.org/pdf/2609.12822v1

### L12. The GPT-4.1 grader using physician-written criteria scored above the average physician in five of seven themes (vendor-authored)

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L12
- Quote: "GPT-4.1 as a grader exceeds the random baseline for all themes as shown in Table 5. It exceeds the average physician score in five out of seven themes"
- Source: Arora et al. (OpenAI), HealthBench, arXiv 2505.08775v1, §8.1 + Table 5, page-text offset 42.6%
- URL: https://arxiv.org/pdf/2505.08775v1

### L13. Using the expert rubric, five of seven graders passed between 38.0% and 45.7% of the proofs experts failed and wrongly failed between 5.0% and 12.3% of the proofs experts passed; the other two passed 63.6% and 74.8% of the failed proofs (false fails 3.8% and 2.5%)

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L13
- Quote: "82.7% 43.0% 5.2% 82.0% 45.7% 5.0% 81.8% 38.0% 8.8% 80.9% 41.4% 8.7% 78.9% 39.6% 12.3% 77.0% 63.6% 3.8% 74.2% 74.8% 2.5%"
- Source: Gonzalez et al., QEDBench (ICML 2026), arXiv 2602.20629v3, Fig. 5, page-text offset 5.4%
- URL: https://arxiv.org/pdf/2602.20629v3

### L13a. The paper defines leniency as the false-positive rate and harshness as the false-negative rate against the expert rubric pass mark

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L13a
- Quote: "We decompose errors into Leniency Rate(False Positives) and Harshness Rate(False"
- Source: Gonzalez et al., QEDBench, Fig. 5 caption, page-text offset 5.5%
- URL: https://arxiv.org/pdf/2602.20629v3

### L13b. The paper reports significant positive bias for certain frontier graders it names (Claude Opus 4.5, DeepSeek-V3, Qwen 2.5 Max, Llama 4 Maverick); it does not say this of all graders

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L13b
- Quote: "frontier evaluators like Claude Opus 4.5, DeepSeek-V3, Qwen 2.5 Max, and Llama 4 Maverick exhibit significant positive bias"
- Source: Gonzalez et al., QEDBench, abstract, page-text offset 0.5%
- URL: https://arxiv.org/pdf/2602.20629v3

### L13c. QEDBench's leniency and harshness figures are for graders using the expert rubric

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L13c
- Quote: "Judge Reliability Metrics (Expert Rubric Only)"
- Source: Gonzalez et al., QEDBench (ICML 2026), arXiv 2602.20629v3, Fig. 5 label, page-text offset 5.4%
- URL: https://arxiv.org/pdf/2602.20629v3

### L14. Order-swap flip rate (MT-Bench / JudgeBench): Gemini 3.1 Pro 0.035 / 0.020; Claude Opus 4.6 0.038 / 0.022

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L14
- Quote: "Gemini 3.1 Pro 0.977 0.989 0.035 0.978 0.989 0.020 Claude Sonnet 4 0.960 0.976 0.074 0.941 0.966 0.146 GPT-4o-mini 0.959 0.975 0.127 0.896 0.939 0.380 Claude Opus 4.6 0.958 0.979 0.038 0.974 0.986 0.022"
- Source: Norman, Rivera, Hughes, Reliability without Validity (preprint), arXiv 2606.19544v1, App. Table 6, page-text offset 78.0%
- URL: https://arxiv.org/pdf/2606.19544v1

### L14a. Order-swap flip rate (MT-Bench / JudgeBench): GPT-5.4 0.114 / 0.105

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L14a
- Quote: "GPT-5.4 0.932 0.965 0.114 0.933 0.966 0.105"
- Source: Norman, Rivera, Hughes, Reliability without Validity (preprint), arXiv 2606.19544v1, App. Table 6, page-text offset 78.8%
- URL: https://arxiv.org/pdf/2606.19544v1

### L15. 2026 JudgeBench, chance-corrected agreement (Cohen's kappa): Gemini 3.1 Pro 0.841, Claude Opus 4.6 0.875, Claude Sonnet 4.6 0.782, GPT-5.4 0.606; exact match 0.964, 0.956, 0.920, 0.812 (one order, ties excluded)

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L15
- Quote: "Gemini 3.1 Pro 0.849 0.511 33.8 0.964 0.841 12.3 0.956 0.898 5.9 Claude Opus 4.6 0.848 0.489 35.9 0.956 0.875 8.1 0.943 0.879 6.4 DeepSeek V3.2 0.845 0.486 35.9 0.791 0.545 24.5 0.921 0.826 9.5 Claude Sonnet 4.6 0.851 0.484 36.7 0.920 0.782 13.8 0.942 0.871 7.1 Llama 3.3 70B 0.841 0.465 37.6 0.664 0.283 38.1 0.892 0.769 12.3 Kimi K2.5 0.846 0.461 38.5 0.864 0.720 14.5 0.937 0.873 6.4 GPT-5.4 0.836 0.457 38.0 0.812 0.606"
- Source: Norman, Rivera, Hughes, Reliability without Validity (preprint), arXiv 2606.19544v1, §4.1 Table 2, page-text offset 24.2%
- URL: https://arxiv.org/pdf/2606.19544v1

### L15b. The same authors warn that raw agreement does not correct for chance and overstates how well a grader discriminates

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L15b
- Quote: "This family of metrics does not correct for chance"
- Source: Norman, Rivera, Hughes (preprint), arXiv 2606.19544v1, §2, page-text offset 4.0%
- URL: https://arxiv.org/pdf/2606.19544v1

### L16. Human annotators preferred the markdown side 57% of the time; four of five judges preferred it 73%–97% on the same pairs

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L16
- Quote: "human annotators prefer the markdown side only 57% of the time, while four of the five judges prefer it 73%-97% on the same pairs"
- Source: Soumik, bias mitigation in LLM-as-a-Judge (TMLR 06/2026), arXiv 2604.23178v2, page-text offset 44.2%
- URL: https://arxiv.org/pdf/2604.23178v2

### L16a. The human comparison was two independent annotators on a 30-pair subsample: they preferred markdown 57% of the time on average, while four of the five judges preferred it 73%-97%

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L16a
- Quote: "two independent annotators on a 30-pair subsample preferred markdown only 57% of the time on average, while four of the five judges preferred markdown 73%-97%"
- Source: Soumik, bias mitigation in LLM-as-a-Judge (TMLR 06/2026), arXiv 2604.23178v2, §1 / App. G, page-text offset 28.5%
- URL: https://arxiv.org/pdf/2604.23178v2

### L16b. The style test compared markdown with plain prose

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L16b
- Quote: "STYLE (markdown vs. plain prose)"
- Source: Soumik, bias mitigation in LLM-as-a-Judge (TMLR 06/2026), arXiv 2604.23178v2, §3, page-text offset 22.0%
- URL: https://arxiv.org/pdf/2604.23178v2

### L17. One study finds most measured self-preference is judges being unsure on hard items (its own summary: 89.6%, substantially reducing but not eliminating the evidence)

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L17
- Quote: "evaluator uncertainty accounts for an average of 89.6%"
- Source: Roytburg et al., Are LLM Evaluators Really Narcissists? (ICML 2026), arXiv 2601.22548v4, §1, page-text offset 9.2% · numbers present in text: 89.6%
- URL: https://arxiv.org/pdf/2601.22548v4

### L18. GPT-5 as judge: false-positive rate on its own wrong GSM8K answers 82.0 [72.4, 90.7], 67.3 points above its rate on others' wrong answers

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L18
- Quote: "GPT-5 82.0[72.4, 90.7]+67.3"
- Source: Zhang et al. (Microsoft/MIT), Can We Trust LLM Judges (preprint), arXiv 2609.12002v1, App. B Table 10 (GSM8K), page-text offset 87.4%
- URL: https://arxiv.org/pdf/2609.12002v1

### L19. Across three QA sets (NQ, TQA, HPQA), the panel (PoLL) beat GPT-4 alone on all three (0.763 vs 0.627; 0.906 vs 0.841; 0.867 vs 0.830), edged its best member on two (0.763 vs 0.749; 0.906 vs 0.902) and fell below Haiku on the third (0.867 vs 0.873)

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L19
- Quote: "GPT-4 0.627 0.841 0.830 CMD-R 0.734 0.902 0.815 Haiku 0.749 0.894 0.873 GPT-3.5 0.726 0.859 0.833 PoLL 0.763 0.906 0.867"
- Source: Verga et al. (Cohere), Replacing Judges with Juries, arXiv 2404.18796v2, Table 1, page-text offset 18.5%
- URL: https://arxiv.org/pdf/2404.18796v2

### L21. When the judge is no better than the model it grades, correcting it with human labels can at most halve the labels you need

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L21
- Quote: "no debiasing method can decrease the required amount of ground truth labels by more than half"
- Source: Dorner, Nastl, Hardt (ICLR 2025 Oral), arXiv 2410.13341v4, abstract, page-text offset 1.4%
- URL: https://arxiv.org/pdf/2410.13341v4

### L22. In the original JudgeBench (2024), many strong graders, GPT-4o among them, scored only slightly better than random on hard correctness pairs

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L22
- Quote: "performing just slightly better than random guessing"
- Source: Tan et al., JudgeBench (ICLR 2025), arXiv 2410.12784v2, abstract, page-text offset 1.8%
- URL: https://arxiv.org/pdf/2410.12784v2

### L22a. JudgeBench found a grader's ability to verify answers highly correlated with its ability to solve the problem itself

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L22a
- Quote: "the ability of the judge to verify the solution pairs is highly correlated with its ability to solve the problem itself"
- Source: Tan et al., JudgeBench (ICLR 2025), arXiv 2410.12784v2, §4.4, page-text offset 41.3%
- URL: https://arxiv.org/pdf/2410.12784v2

### L23. PRIOR ART: Eugene Yan's August 2024 review of LLM-evaluators drew on about two dozen papers and already covers reference-based evaluation and position bias

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L23
- Quote: "Drawing from two dozen papers"
- Source: Eugene Yan, Evaluating the Effectiveness of LLM-Evaluators (aka LLM-as-Judge), Aug 2024, page-text offset 1.5%
- URL: https://eugeneyan.com/writing/llm-evaluators/

### L24. PRIOR ART: Hamel Husain's guide (Oct 2024) has a principal domain expert make pass/fail judgments and tracks the judge's agreement with them; the warning that agreement misleads on imbalanced data was added in a 2025 revision

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L24
- Quote: "I also tracked agreement rates over time to ensure we were converging on a good prompt"
- Source: Hamel Husain, Using LLM-as-a-Judge For Evaluation: A Complete Guide (Oct 2024), page-text offset 60.7%
- URL: https://hamel.dev/blog/posts/llm-judge/

### L24a. Husain's guide, as revised in September 2026, advises about 100 examples per failure mode to validate an automated grader and says below 60 the confidence intervals are often too wide; the October 2024 original did not contain this advice

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L24a
- Quote: "Aim for about 100 examples per failure mode, with enough Pass and Fail examples to measure both classes. Below 60 examples, the confidence intervals are often too wide"
- Source: Hamel Husain, Using LLM-as-a-Judge For Evaluation (Oct 2024), page-text offset 50.8%
- URL: https://hamel.dev/blog/posts/llm-judge/

### L25. In code review against a written requirement, GPT-4o wrongly failed correct code (false-negative rate) 26.2% (HumanEval) and 35.9% (MBPP) when judging directly, rising to 73.2% and 87.9% when also asked to explain and repair

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L25
- Quote: "GPT-4o achieves a relatively low FNR in HumanEval (26.2%) and MBPP (35.9%), but once explana- tions and repairs are required, the FNR increases sharply to 73.2% in HumanEval and 87.9% in MBPP"
- Source: Jin & Chen, Are LLMs Reliable Code Reviewers? Systematic Overcorrection (Automated Software Engineering 33, art. 90, 26 Jun 2026; read as arXiv 2603.00539v1, and the published text prints the same sentence), §5.2, page-text offset 38.7% · numbers present in text: 26.2, 35.9, 73.2, 87.9
- URL: https://arxiv.org/pdf/2603.00539v1

### L25a. In that paper, FNR is rejecting correct code (a false fail) and FPR is accepting buggy code (a false pass)

- Status: VERIFIED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L25a
- Quote: "FNR reflects over-correction (rejecting correct implementations), while FPR reflects unsafe acceptance"
- Source: Jin & Chen, Are LLMs Reliable Code Reviewers? Systematic Overcorrection (Automated Software Engineering 33, art. 90, 2026; read as arXiv 2603.00539v1), Table 2 caption, page-text offset 37.2%
- URL: https://arxiv.org/pdf/2603.00539v1

### P1. The prior registered before any paper was read: graders would hold on easy, preference-style chat grading and fail on correctness of hard items they cannot solve themselves and in specialised domains

- Status: EXECUTED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-P1
- Quote: judges HOLD on easy, preference-style chat grading and FAIL on (a) correctness of hard items the judge cannot itself solve and (b) specialised domains
- Source: first-party: pre-registration, 19:57 BST 2026-09-26

### J1. As run, scored to the precision each question asked for (one pass, n = 24 arithmetic-style questions with one right answer each, listed in Line/judge-check/items.py, 2026-09-26 21:02 BST): Claude Haiku 4.5, answering without tools, got 4 of 24 wrong. Haiku 4.5 grading those answers without a key passed 3 of its 4 wrong answers; given the verified key, it passed 0 of 4

- Status: EXECUTED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J1
- Quote: HEADLINE haiku passed 3/4 of its own wrong answers without a key, 0/4 with one
- Source: first-party: Judge Check run

### J2. Claude Sonnet 5 grading the same answers without a key agreed with the key on 23 of 24: no false passes, one false fail (a correct final answer reached by an unjustified leap, which it failed although told to grade the final answer); with the key, 24 of 24

- Status: EXECUTED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J2
- Quote: no key agree 23/24 · false passes 0 · false fails 1 · unparsed 0 / with key agree 24/24
- Source: first-party: Judge Check run

### J3. Re-dressing the same answers in markdown changed 2 of 24 of Haiku's verdicts (both from pass to fail on wrong answers) and none of Sonnet's

- Status: EXECUTED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J3
- Quote: markdown re-dress changed 2/24 verdicts / markdown re-dress changed 0/24 verdicts
- Source: first-party: Judge Check run

### J4. As run (the precision each question asked for): Haiku as judge could not solve 2 of the 24 items itself and, without a key, passed the wrong answer on both; on H6 (digit sum of 50!) it passed a wrong answer although it solved the item correctly when asked. Under the registered rule it could not solve 1 item (H11) and passed the wrong answer on it; H6 holds under both rules

- Status: EXECUTED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J4
- Quote: solved it itself: 22/24 · no-key agree on solved 21/22 · on unsolved 0/2
- Source: first-party: Judge Check run

### J5. The run cost $2.83. As run, predictions P1, P4 and P5 passed, P3 passed for Haiku, P2 and Sonnet's P3 were unscorable. Under the pre-registered scoring rule (score.py --registered), P1 failed for both graders, P5 failed, Sonnet's P3 failed, P4 passed, Haiku's P3 passed

- Status: EXECUTED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J5
- Quote: cost $2.8326 · Predictions (pre-registered)
- Source: first-party: Judge Check run

### T1. Of 19 studies we found that report which way AI graders err against human or ground-truth labels, 12 lean toward passing flawed work (4 of them weakly, adversarially or in one direction only), 3 toward failing sound work (code review against a spec; older-model essay grading; a frontier panel on clinical diagnoses) and 4 are mixed

- Status: EXECUTED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-T1
- Quote: Tally (Part 2, excluding 7b): Supports 12 (1, 2, 3, 4, 5, 6, 9, 13, 14, 16, 18, 19; of these, 13, 16, 18 and 19 are weak, adversarial or one-directional) · Contradicts 3 (7, 8, 11) · Mixed 4 (10, 12, 15, 17)
- Source: first-party: error-direction sweep, 2026-09-26

### T2. The evidence base is 77 distinct studies: 63 in the four sweeps and 14 more in the error-direction check

- Status: EXECUTED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-T2
- Quote: PAPERS TOTAL 77 distinct studies
- Source: first-party: paper count

### J6. Under the pre-registered scoring rule (0.001 relative tolerance), Haiku's answers were wrong on 2 of 24; Haiku grading them passed both wrong answers without a key and neither with one; both graders, given the key, failed the two answers that missed the precision the question asked for

- Status: EXECUTED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J6
- Quote: HEADLINE haiku passed 2/2 of its own wrong answers without a key, 0/2 with one
- Source: first-party: Judge Check run, rescored

### W1. Our earlier piece A Green Score Is Not Evidence reported a groundedness metric that scored its best with the evidence removed

- Status: EXECUTED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-W1
- Quote: A groundedness metric scored its best with the evidence removed.
- Source: first-party: published article (subtitle)
- URL: https://harryfloyd.substack.com/p/a-green-score-is-not-evidence

### W2. Our earlier piece Your AI Looks Best Where You Can Check It Least argued that the work that is hardest to check is where AI output looks best

- Status: EXECUTED (read 2026-09-26)
- Row link: https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-W2
- Quote: Your AI Looks Best Where You Can Check It Least
- Source: first-party: published article (title)
- URL: https://harryfloyd.substack.com/p/your-ai-looks-best-where-you-check-least

## Cite

- A claim: name the original source (and its locator) first, then the row it was checked in, e.g. "<source>, <locator>. Checked in: Claim Ledger, "Your AI Grader Is Only as Good as Its Answer Key", The Durability Curve, row <id>, https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-<id>."
- The essay: Floyd, Harry (2026). Your AI Grader Is Only as Good as Its Answer Key. The Durability Curve. https://durabilitycurve.com/blog/your-ai-grader-answer-key/
- This ledger: Floyd, Harry (2026). Claim Ledger: Your AI Grader Is Only as Good as Its Answer Key. The Durability Curve. https://durabilitycurve.com/claims/your-ai-grader-answer-key/
- Downloads: https://durabilitycurve.com/claims/your-ai-grader-answer-key.csv · https://durabilitycurve.com/claims/your-ai-grader-answer-key.json

Licence: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Reuse it, quote it or train on it, with credit to The Durability Curve and a link. Open to AI (robots.txt: search=yes, ai-input=yes, ai-train=yes).
