{
 "ledger": {
  "title": "Claim Ledger: Your AI Grader Is Only as Good as Its Answer Key",
  "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/",
  "essay": {
   "title": "Your AI Grader Is Only as Good as Its Answer Key",
   "url": "https://durabilitycurve.com/blog/your-ai-grader-answer-key/",
   "published": "2026-09-28"
  },
  "last_checked": "2026-09-27",
  "format": "current",
  "counts": {
   "claims": 68,
   "by_status": {
    "VERIFIED": 57,
    "EXECUTED": 11
   },
   "verified_at_primary": 57,
   "removed": 0
  },
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "attribution": "The Durability Curve (Harry Floyd), https://durabilitycurve.com/claims/your-ai-grader-answer-key/"
 },
 "rows": [
  {
   "id": "L1",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L1",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "Without a correct reference, judges agreed well with human experts only on questions they had been able to answer themselves",
   "quote": "\"when not provided with a correct reference, judges show high agreement with human experts only on questions the judges were able to correctly answer themselves\"",
   "removal_reason": "",
   "source": "Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, abstract",
   "source_url": "https://arxiv.org/pdf/2503.05061v3",
   "locator": "page-text offset 1.2%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L1a",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L1a",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "GPT-4o as judge, Cohen's kappa with experts. Pairwise: questions it answered correctly 0.78 with no key, 0.92 with a human key; questions it got wrong 0.30 with no key, 0.16 with its own answer as the key, 0.83 with a human key. Single grading: questions it answered correctly 0.46 with no key and 0.59 with a human key; questions it got wrong 0.16 and 0.81",
   "quote": "\"GPT-4o Correct 0.78±0.02 0.86±0.02 0.92±0.02 0.46±0.07 0.52±0.06 0.59±0.06 Incorrect 0.30±0.09 0.16±0.08 0.83±0.06 0.16±0.14 0.13±0.14 0.81±0.13\"",
   "removal_reason": "",
   "source": "Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Table 4 (GPT-4o row)",
   "source_url": "https://arxiv.org/pdf/2503.05061v3",
   "locator": "page-text offset 21.0%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L1b",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L1b",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "Pairwise, questions each judge got wrong, no key vs human key: Llama 3.3 70B 0.39 vs 0.92; Phi 4 -0.08 vs 0.75; Qwen 2.5 7B 0.21 vs 0.63; Yi 1.5 34B 0.13 vs 0.56 (own answer as key: 0.35, -0.08, 0.14, 0.02)",
   "quote": "\"Llama 3.3 70b Correct 0.79±0.03 0.86±0.03 0.93±0.02 0.61±0.08 0.54±0.08 0.73±0.07 Incorrect 0.39±0.06 0.35±0.07 0.92±0.03 0.13±0.06 0.23±0.10 0.69±0.09 Phi 4 Correct 0.67±0.07 0.88±0.04 0.91±0.04 0.48±0.09 0.47±0.07 0.67±0.07 Incorrect -0.08±0.09 -0.08±0.10 0.75±0.06 0.03±0.05 0.02±0.12 0.56±0.12 Qwen 2.5 7B Correct 0.65±0.05 0.69±0.04 0.78±0.04 0.31±0.11 0.37±0.09 0.54±0.09 Incorrect 0.21±0.07 0.14±0.07 0.63±0.06 0.08±0.07 0.21±0.10 0.56±0.08 Yi 1.5 34B Correct 0.31±0.12 0.53±0.12 0.77±0.09 0.23±0.11 0.29±0.08 0.43±0.09 Incorrect 0.13±0.09 0.02±0.09 0.56±0.08\"",
   "removal_reason": "",
   "source": "Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Table 4 (other judges)",
   "source_url": "https://arxiv.org/pdf/2503.05061v3",
   "locator": "page-text offset 21.2%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L1c",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L1c",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "The paper notes that a relatively small model (Qwen 2.5 7B) with a human reference can grade better than a larger model (GPT-4o) without one, on the same set of responses",
   "quote": "\"providing a human reference to a relatively small model (such as Qwen 2.5 7B ) can yield better judgments than using a larger model without human references (such asGPT-4o) on the same set of responses\"",
   "removal_reason": "",
   "source": "Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6",
   "source_url": "https://arxiv.org/pdf/2503.05061v3",
   "locator": "page-text offset 22.2%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L2",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L2",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "With its own answer as the key, GPT-4o's pairwise kappa was 0.86 on questions it had answered correctly and 0.16 on questions it could not answer",
   "quote": "\"decreases from 0.86 to 0.16 in the pairwise case on questions GPT-4o could not answer\"",
   "removal_reason": "",
   "source": "Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6",
   "source_url": "https://arxiv.org/pdf/2503.05061v3",
   "locator": "page-text offset 23.0%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L3",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "Averaged across all five graders on questions GPT-4o answered correctly (Table 3 does not say whether single and pairwise grading are pooled): kappa with a human key 0.69, a verified GPT-4o answer as key 0.61, a completely unrelated key 0.50, no key 0.46, a subtly wrong key 0.21",
   "quote": "\"Human 0.69 ±0.06 0.74±0.06 0.57±0.08 GPT-4o (✓) 0.61 ±0.07 0.66±0.07 0.52±0.09 None 0.46 ±0.08 0.56±0.08 0.22±0.09 Random 0.50 ±0.08 0.47±0.08 0.09±0.1 Wrong 0.21 ±0.06\"",
   "removal_reason": "",
   "source": "Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Table 3",
   "source_url": "https://arxiv.org/pdf/2503.05061v3",
   "locator": "page-text offset 19.3%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L3c",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3c",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "Table 3's figures are across all five graders",
   "quote": "\"Table 3 presents results for these reference types across all five judges\"",
   "removal_reason": "",
   "source": "Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6",
   "source_url": "https://arxiv.org/pdf/2503.05061v3",
   "locator": "page-text offset 26.3%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L3a",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3a",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "The paper says a slightly incorrect reference can in some cases be worse than no reference at all",
   "quote": "\"in some cases be worse than\"",
   "removal_reason": "",
   "source": "Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6",
   "source_url": "https://arxiv.org/pdf/2503.05061v3",
   "locator": "page-text offset 25.4%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L3d",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3d",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "The authors describe the Random key as a completely unrelated reference, and say a slightly incorrect reference can in some cases be worse than a completely unrelated one or none",
   "quote": "\"providing a slightly incorrect reference can in some cases be worse than providing a completely unrelated reference or no reference at all\"",
   "removal_reason": "",
   "source": "Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §6",
   "source_url": "https://arxiv.org/pdf/2503.05061v3",
   "locator": "page-text offset 25.4%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L3b",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L3b",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "The authors conclude that verifying a model-generated answer, rather than writing and checking one by hand, was sufficient as a key in their setting",
   "quote": "\"Verifying responses generated by a model rather than manually writing and checking answers via annotators is sufficient in this setting\"",
   "removal_reason": "",
   "source": "Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, §7",
   "source_url": "https://arxiv.org/pdf/2503.05061v3",
   "locator": "page-text offset 28.6%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L4",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "2026 judges, questions each got wrong, against a stand-in label: single grading no key vs human key, Opus 4.7 0.26 vs 0.62, GPT-5.4 0.32 vs 0.67, Gemini 3.1 0.44 vs 0.71; pairwise 0.33 vs 0.66, 0.46 vs 0.87, 0.67 vs 0.95 (own answer as key, pairwise: 0.33, 0.45, 0.62)",
   "quote": "\"Opus 4.7 Correct 0.86±0.04 0.90±0.04 0.91±0.04 0.63±0.06 0.69±0.06 0.77±0.05 Incorrect 0.33±0.16 0.33±0.15 0.66±0.13 0.26±0.14 0.27±0.12 0.62±0.11 GPT-5.4 Correct 0.75±0.06 0.87±0.04 0.89±0.04 0.51±0.06 0.64±0.06 0.66±0.06 Incorrect 0.46±0.11 0.45±0.13 0.87±0.08 0.32±0.11 0.25±0.12 0.67±0.09 Gemini 3.1 Correct 0.78±0.06 0.80±0.06 0.83±0.05 0.55±0.06 0.62±0.06 0.69±0.06 Incorrect 0.67±0.13 0.62±0.14 0.95±0.07 0.44±0.13 0.35±0.14 0.71±0.10\"",
   "removal_reason": "",
   "source": "Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, Table 12",
   "source_url": "https://arxiv.org/pdf/2503.05061v3",
   "locator": "page-text offset 90.1%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L4a",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4a",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "The authors treat the 2026-judge rows, scored against a stand-in label, as corroboration",
   "quote": "\"we regard these results as corroboration rather than equivalent to our main table\"",
   "removal_reason": "",
   "source": "Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J",
   "source_url": "https://arxiv.org/pdf/2503.05061v3",
   "locator": "page-text offset 88.7%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L4b",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4b",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "The stand-in label is GPT-4o grading with the human-verified references",
   "quote": "\"We used the GPT-4o Judge with the human-verified references as a proxy correctness label\"",
   "removal_reason": "",
   "source": "Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J",
   "source_url": "https://arxiv.org/pdf/2503.05061v3",
   "locator": "page-text offset 88.4%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L4c",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4c",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "The authors report that higher reasoning effort can narrow the no-reference gap, for example GPT-5.4 at high effort",
   "quote": "\"higher reasoning effort can narrow the no-reference gap (e.g.,GPT-5.4 at high effort)\"",
   "removal_reason": "",
   "source": "Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J",
   "source_url": "https://arxiv.org/pdf/2503.05061v3",
   "locator": "page-text offset 88.9%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L4d",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L4d",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "GPT-5.4 at high reasoning effort, on questions it got wrong, against the stand-in label: pairwise 0.68 with no key and 0.84 with the human key; single grading 0.38 and 0.67",
   "quote": "\"GPT-5.4 low Correct 0.85±0.05 0.90±0.04 0.54±0.06 0.57±0.05 Incorrect 0.55±0.13 0.95±0.06 0.34±0.11 0.56±0.09 default Correct 0.75±0.06 0.89±0.04 0.51±0.06 0.66±0.06 Incorrect 0.46±0.11 0.87±0.08 0.32±0.11 0.67±0.09 high Correct 0.80±0.06 0.87±0.05 0.68±0.08 0.67±0.06 Incorrect 0.68±0.14 0.84±0.11 0.38±0.19 0.67±0.11\"",
   "removal_reason": "",
   "source": "Krumdick et al., No Free Labels (COLM 2026), arXiv 2503.05061v3, Appendix J, Table 13 (GPT-5.4 rows)",
   "source_url": "https://arxiv.org/pdf/2503.05061v3",
   "locator": "page-text offset 91.4%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L5",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L5",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "On MATH, DeepSeek V3 as a judge without a reference reached Scott's pi 0.72 against ground truth; as a matcher with the reference, 0.98",
   "quote": "\"DeepSeek v3 model achieves only modest agreement π = 0.72, while as a matcher, it achieves π = 0.98\"",
   "removal_reason": "",
   "source": "Chandak et al., Answer Matching Outperforms Multiple Choice (preprint), arXiv 2507.02856v1, §3.1",
   "source_url": "https://arxiv.org/pdf/2507.02856v1",
   "locator": "page-text offset 22.6%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L5a",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L5a",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "A 1.7-billion-parameter Qwen3 model matching answers to the reference reached pi 0.97 against ground truth on MATH",
   "quote": "\"answer matching, even with the 1.7 billion parameter Qwen3 model (non-thinking mode), achieves near-perfect alignment with the ground-truth (π = 0.97)\"",
   "removal_reason": "",
   "source": "Chandak et al., Answer Matching Outperforms Multiple Choice (preprint), arXiv 2507.02856v1, §3.1",
   "source_url": "https://arxiv.org/pdf/2507.02856v1",
   "locator": "page-text offset 22.4%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L5b",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L5b",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "Answer matching with recent models, even small ones, reached near-perfect agreement with human grading, in the range of inter-annotator agreement",
   "quote": "\"matching using recent models-even small ones-achieves near-perfect agreement, in the range of inter-annotator agreement\"",
   "removal_reason": "",
   "source": "Chandak et al., Answer Matching (preprint), arXiv 2507.02856v1, abstract",
   "source_url": "https://arxiv.org/pdf/2507.02856v1",
   "locator": "page-text offset 1.3%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L5c",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L5c",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "On MATH, the 'true grades' come from MATH-Verify, a rule-based checker that compares each answer with the reference",
   "quote": "\"MATH-Verify library (Kydlicek et al., 2025) implements rule-based ground-truth evaluations of generative responses\"",
   "removal_reason": "",
   "source": "Chandak et al., Answer Matching (preprint), arXiv 2507.02856v1, §3.1",
   "source_url": "https://arxiv.org/pdf/2507.02856v1",
   "locator": "page-text offset 22.0%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L6",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L6",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "For frontier judges (DeepSeek V3, o4-mini), 80%+ of grading errors were false passes: the judge marked correct what human annotation marked incorrect",
   "quote": "\"errors disproportionately (80%+) arise from false positives\"",
   "removal_reason": "",
   "source": "Chandak et al., arXiv 2507.02856v1, §3.2",
   "source_url": "https://arxiv.org/pdf/2507.02856v1",
   "locator": "page-text offset 28.5%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L7",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L7",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "On MT-Bench, without ties, GPT-4 agreed with expert labellers 85% of the time against 81% between humans",
   "quote": "\"The agreement under setup S2 (w/o tie) between GPT-4 and humans reaches 85%, which is even higher than the agreement among humans (81%).\"",
   "removal_reason": "",
   "source": "Zheng et al., Judging LLM-as-a-Judge (NeurIPS 2023 D&B), arXiv 2306.05685v4, §4.2 + Table 5",
   "source_url": "https://arxiv.org/pdf/2306.05685v4",
   "locator": "page-text offset 30.7%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L7a",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L7a",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "First turn, with ties counted (setup S1): GPT-4 grading pairs agreed with humans 66%, GPT-4 grading single answers 60%, humans with each other 63%; without ties (S2) both GPT-4 modes 85% against 81%",
   "quote": "\"G4-Pair 70% 1138 66% 1343 97% 662 85% 859 G4-Single - 60% 1280 - 85% 739 Human - 63% 721 - 81% 479 (a) First Turn\"",
   "removal_reason": "",
   "source": "Zheng et al. (NeurIPS 2023 D&B), arXiv 2306.05685v4, Table 5(a)",
   "source_url": "https://arxiv.org/pdf/2306.05685v4",
   "locator": "page-text offset 34.1%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L7b",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L7b",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "In 2023 the MT-Bench authors proposed a reference-guided grader that first answers the question itself and uses its own answer as the reference; on their maths questions it cut the failure rate from 70% to 15%",
   "quote": "\"we first generate LLM judge's answer independently, and then display it as a reference answer in the judge prompt. In Table 4, we see a significant improvement in failure rate (from 70% to 15%) over the default prompt\"",
   "removal_reason": "",
   "source": "Zheng et al. (NeurIPS 2023 D&B), arXiv 2306.05685v4, §3.4",
   "source_url": "https://arxiv.org/pdf/2306.05685v4",
   "locator": "page-text offset 27.1%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L7c",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L7c",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "The MT-Bench authors saw GPT-4 misjudge an answer to a maths problem it could solve when asked separately: it was misled by the answers it was shown",
   "quote": "\"although GPT-4 can solve the problem (when asked separately), it was misled by the provided answers, ultimately resulting in incorrect judgment\"",
   "removal_reason": "",
   "source": "Zheng et al. (NeurIPS 2023 D&B), arXiv 2306.05685v4, §3.3",
   "source_url": "https://arxiv.org/pdf/2306.05685v4",
   "locator": "page-text offset 24.1%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L26",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L26",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "In 2023 the Prometheus authors found that removing the reference answer hurt their grader most, and said a reference relieves the grader of solving the question itself",
   "quote": "\"relieves the need for the evaluator LM to internally solve the instruction and only focus on assessing the response\"",
   "removal_reason": "",
   "source": "Kim et al., Prometheus (ICLR 2024, per the arXiv comments field), arXiv 2310.08491",
   "source_url": "https://arxiv.org/pdf/2310.08491",
   "locator": "page-text offset 45.7%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L8",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L8",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "Re-tested on expert labels with a leave-one-annotator-out test, no LLM grader passed on MT-Bench, one of two datasets where none did",
   "quote": "\"in two datasets (MT-Bench, and SummEval), none of the LLMs pass the test\"",
   "removal_reason": "",
   "source": "Calderon, Reichart, Dror, The Alternative Annotator Test (ACL 2025), arXiv 2501.10970v4, §5",
   "source_url": "https://arxiv.org/pdf/2501.10970v4",
   "locator": "page-text offset 21.8%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L9",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L9",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "In Chatbot Arena's expert relabel, GPT-4 agreed with the two experts 81.0% and 78.5% (Llama-2-13b battles) and 76.3% and 79.3% (GPT-3.5-Turbo battles); the experts agreed with each other 89.8% and 79.4%",
   "quote": "\"Llama-2-13b Expert 1 Expert 2 GPT-4 Crowd 72.8% 77.8% 75.6% Expert 1 - 89.8% 81.0% Expert 2 - - 78.5% GPT-3.5-Turbo Expert 1 Expert 2 GPT-4 Crowd 73.8% 83.1% 75.6% Expert 1 - 79.4% 76.3% Expert 2 - - 79.3%\"",
   "removal_reason": "",
   "source": "Chiang et al., Chatbot Arena (ICML 2024), arXiv 2403.04132v1, Table 3",
   "source_url": "https://arxiv.org/pdf/2403.04132v1",
   "locator": "page-text offset 35.9%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L10",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L10",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "On literature-grounded science answers, the best judge, o3, reached 65.1% accuracy",
   "quote": "\"Even the best-performing model, o3, achieves only 65.1% accuracy.\"",
   "removal_reason": "",
   "source": "Zhao et al., SciArena (NeurIPS 2025 D&B), arXiv 2507.01001v2, §6.2 + Table 1",
   "source_url": "https://arxiv.org/pdf/2507.01001v2",
   "locator": "page-text offset 32.6%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L10a",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L10a",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "Random guessing on the same task scores 50.0",
   "quote": "\"Random Guess 50.0 o3 65.1\"",
   "removal_reason": "",
   "source": "Zhao et al., SciArena (NeurIPS 2025 D&B), arXiv 2507.01001v2, Table 3",
   "source_url": "https://arxiv.org/pdf/2507.01001v2",
   "locator": "page-text offset 32.3%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L10b",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L10b",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "Expert annotators' average inter-annotator agreement: accuracy 0.82, kappa 0.76",
   "quote": "\"Average 0.82 0.76 0.94 0.91\"",
   "removal_reason": "",
   "source": "Zhao et al., SciArena, arXiv 2507.01001v2, Table 1",
   "source_url": "https://arxiv.org/pdf/2507.01001v2",
   "locator": "page-text offset 21.9%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L11",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L11",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "The best frontier judge reached physician agreement on 3 of 5 clinical grading tasks (NEJM Healer, BIDMC ER, Landmark) and fell short on two; no single model matched on all five",
   "quote": "\"No single base model matched physician agreement across all five tasks.\"",
   "removal_reason": "",
   "source": "Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results",
   "source_url": "https://arxiv.org/pdf/2609.12822v1",
   "locator": "page-text offset 15.1% · numbers present in text: 82%, 77%, 91%, 88%, 68%, 67%, 87%, 92%, 84%, 95%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L11a",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L11a",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "Best judge vs physician agreement per task: NEJM Healer 82% vs 77%, BIDMC ER 91% vs 88%, Landmark 68% vs 67%; short on NEJM CPCs 87% vs 92% and Grey Matters Management 84% vs 95%",
   "quote": "\"reached physician inter-rater agreement on the NEJM Healer (Claude, 82% vs. 77% for physicians), BIDMC ER (Claude, 91% vs. 88%), and Landmark Diagnostic Cases (Gemini, 68% vs. 67%), but fell short on the NEJM CPCs (GPT-5, 87% vs. 92%) and the Grey Matters Management Cases (Claude, 84% vs. 95%)\"",
   "removal_reason": "",
   "source": "Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results",
   "source_url": "https://arxiv.org/pdf/2609.12822v1",
   "locator": "page-text offset 14.5%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L11b",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L11b",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "In the clinical study, some tasks score a diagnosis against the known answer (the Bond score) and others score diagnostic reasoning against a rubric",
   "quote": "\"A score of 5 indicates that the correct diagnosis is included in the differential\"",
   "removal_reason": "",
   "source": "Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results, Methods, Task Rubrics",
   "source_url": "https://arxiv.org/pdf/2609.12822v1",
   "locator": "page-text offset 48.9%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L11c",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L11c",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "The authors' own grader, PrecepTron, fine-tuned on a small number of physician-scored cases, matched or exceeded physician agreement on three of five tasks, so no single model, base or fine-tuned, matched on all five (the abstract's 'physician-level consistent scoring across tasks' is looser than the Results and must not be cited)",
   "quote": "\"PrecepTron approaches physician inter-rater agreement, matching or exceeding it on three of five tasks (NEJM CPCs, NEJM Healer, and BIDMC ER)\"",
   "removal_reason": "",
   "source": "Buckley et al., Scaling Clinical Judgment to Evaluate Medical AI (preprint), arXiv 2609.12822v1, Results (PrecepTron section)",
   "source_url": "https://arxiv.org/pdf/2609.12822v1",
   "locator": "Results",
   "date_read": "2026-09-27"
  },
  {
   "id": "L12",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L12",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "The GPT-4.1 grader using physician-written criteria scored above the average physician in five of seven themes (vendor-authored)",
   "quote": "\"GPT-4.1 as a grader exceeds the random baseline for all themes as shown in Table 5. It exceeds the average physician score in five out of seven themes\"",
   "removal_reason": "",
   "source": "Arora et al. (OpenAI), HealthBench, arXiv 2505.08775v1, §8.1 + Table 5",
   "source_url": "https://arxiv.org/pdf/2505.08775v1",
   "locator": "page-text offset 42.6%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L13",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L13",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "Using the expert rubric, five of seven graders passed between 38.0% and 45.7% of the proofs experts failed and wrongly failed between 5.0% and 12.3% of the proofs experts passed; the other two passed 63.6% and 74.8% of the failed proofs (false fails 3.8% and 2.5%)",
   "quote": "\"82.7% 43.0% 5.2% 82.0% 45.7% 5.0% 81.8% 38.0% 8.8% 80.9% 41.4% 8.7% 78.9% 39.6% 12.3% 77.0% 63.6% 3.8% 74.2% 74.8% 2.5%\"",
   "removal_reason": "",
   "source": "Gonzalez et al., QEDBench (ICML 2026), arXiv 2602.20629v3, Fig. 5",
   "source_url": "https://arxiv.org/pdf/2602.20629v3",
   "locator": "page-text offset 5.4%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L13a",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L13a",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "The paper defines leniency as the false-positive rate and harshness as the false-negative rate against the expert rubric pass mark",
   "quote": "\"We decompose errors into Leniency Rate(False Positives) and Harshness Rate(False\"",
   "removal_reason": "",
   "source": "Gonzalez et al., QEDBench, Fig. 5 caption",
   "source_url": "https://arxiv.org/pdf/2602.20629v3",
   "locator": "page-text offset 5.5%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L13b",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L13b",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "The paper reports significant positive bias for certain frontier graders it names (Claude Opus 4.5, DeepSeek-V3, Qwen 2.5 Max, Llama 4 Maverick); it does not say this of all graders",
   "quote": "\"frontier evaluators like Claude Opus 4.5, DeepSeek-V3, Qwen 2.5 Max, and Llama 4 Maverick exhibit significant positive bias\"",
   "removal_reason": "",
   "source": "Gonzalez et al., QEDBench, abstract",
   "source_url": "https://arxiv.org/pdf/2602.20629v3",
   "locator": "page-text offset 0.5%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L13c",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L13c",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "QEDBench's leniency and harshness figures are for graders using the expert rubric",
   "quote": "\"Judge Reliability Metrics (Expert Rubric Only)\"",
   "removal_reason": "",
   "source": "Gonzalez et al., QEDBench (ICML 2026), arXiv 2602.20629v3, Fig. 5 label",
   "source_url": "https://arxiv.org/pdf/2602.20629v3",
   "locator": "page-text offset 5.4%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L14",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L14",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "Order-swap flip rate (MT-Bench / JudgeBench): Gemini 3.1 Pro 0.035 / 0.020; Claude Opus 4.6 0.038 / 0.022",
   "quote": "\"Gemini 3.1 Pro 0.977 0.989 0.035 0.978 0.989 0.020 Claude Sonnet 4 0.960 0.976 0.074 0.941 0.966 0.146 GPT-4o-mini 0.959 0.975 0.127 0.896 0.939 0.380 Claude Opus 4.6 0.958 0.979 0.038 0.974 0.986 0.022\"",
   "removal_reason": "",
   "source": "Norman, Rivera, Hughes, Reliability without Validity (preprint), arXiv 2606.19544v1, App. Table 6",
   "source_url": "https://arxiv.org/pdf/2606.19544v1",
   "locator": "page-text offset 78.0%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L14a",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L14a",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "Order-swap flip rate (MT-Bench / JudgeBench): GPT-5.4 0.114 / 0.105",
   "quote": "\"GPT-5.4 0.932 0.965 0.114 0.933 0.966 0.105\"",
   "removal_reason": "",
   "source": "Norman, Rivera, Hughes, Reliability without Validity (preprint), arXiv 2606.19544v1, App. Table 6",
   "source_url": "https://arxiv.org/pdf/2606.19544v1",
   "locator": "page-text offset 78.8%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L15",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L15",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "2026 JudgeBench, chance-corrected agreement (Cohen's kappa): Gemini 3.1 Pro 0.841, Claude Opus 4.6 0.875, Claude Sonnet 4.6 0.782, GPT-5.4 0.606; exact match 0.964, 0.956, 0.920, 0.812 (one order, ties excluded)",
   "quote": "\"Gemini 3.1 Pro 0.849 0.511 33.8 0.964 0.841 12.3 0.956 0.898 5.9 Claude Opus 4.6 0.848 0.489 35.9 0.956 0.875 8.1 0.943 0.879 6.4 DeepSeek V3.2 0.845 0.486 35.9 0.791 0.545 24.5 0.921 0.826 9.5 Claude Sonnet 4.6 0.851 0.484 36.7 0.920 0.782 13.8 0.942 0.871 7.1 Llama 3.3 70B 0.841 0.465 37.6 0.664 0.283 38.1 0.892 0.769 12.3 Kimi K2.5 0.846 0.461 38.5 0.864 0.720 14.5 0.937 0.873 6.4 GPT-5.4 0.836 0.457 38.0 0.812 0.606\"",
   "removal_reason": "",
   "source": "Norman, Rivera, Hughes, Reliability without Validity (preprint), arXiv 2606.19544v1, §4.1 Table 2",
   "source_url": "https://arxiv.org/pdf/2606.19544v1",
   "locator": "page-text offset 24.2%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L15b",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L15b",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "The same authors warn that raw agreement does not correct for chance and overstates how well a grader discriminates",
   "quote": "\"This family of metrics does not correct for chance\"",
   "removal_reason": "",
   "source": "Norman, Rivera, Hughes (preprint), arXiv 2606.19544v1, §2",
   "source_url": "https://arxiv.org/pdf/2606.19544v1",
   "locator": "page-text offset 4.0%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L16",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L16",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "Human annotators preferred the markdown side 57% of the time; four of five judges preferred it 73%–97% on the same pairs",
   "quote": "\"human annotators prefer the markdown side only 57% of the time, while four of the five judges prefer it 73%-97% on the same pairs\"",
   "removal_reason": "",
   "source": "Soumik, bias mitigation in LLM-as-a-Judge (TMLR 06/2026), arXiv 2604.23178v2",
   "source_url": "https://arxiv.org/pdf/2604.23178v2",
   "locator": "page-text offset 44.2%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L16a",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L16a",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "The human comparison was two independent annotators on a 30-pair subsample: they preferred markdown 57% of the time on average, while four of the five judges preferred it 73%-97%",
   "quote": "\"two independent annotators on a 30-pair subsample preferred markdown only 57% of the time on average, while four of the five judges preferred markdown 73%-97%\"",
   "removal_reason": "",
   "source": "Soumik, bias mitigation in LLM-as-a-Judge (TMLR 06/2026), arXiv 2604.23178v2, §1 / App. G",
   "source_url": "https://arxiv.org/pdf/2604.23178v2",
   "locator": "page-text offset 28.5%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L16b",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L16b",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "The style test compared markdown with plain prose",
   "quote": "\"STYLE (markdown vs. plain prose)\"",
   "removal_reason": "",
   "source": "Soumik, bias mitigation in LLM-as-a-Judge (TMLR 06/2026), arXiv 2604.23178v2, §3",
   "source_url": "https://arxiv.org/pdf/2604.23178v2",
   "locator": "page-text offset 22.0%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L17",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L17",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "One study finds most measured self-preference is judges being unsure on hard items (its own summary: 89.6%, substantially reducing but not eliminating the evidence)",
   "quote": "\"evaluator uncertainty accounts for an average of 89.6%\"",
   "removal_reason": "",
   "source": "Roytburg et al., Are LLM Evaluators Really Narcissists? (ICML 2026), arXiv 2601.22548v4, §1",
   "source_url": "https://arxiv.org/pdf/2601.22548v4",
   "locator": "page-text offset 9.2% · numbers present in text: 89.6%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L18",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L18",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "GPT-5 as judge: false-positive rate on its own wrong GSM8K answers 82.0 [72.4, 90.7], 67.3 points above its rate on others' wrong answers",
   "quote": "\"GPT-5 82.0[72.4, 90.7]+67.3\"",
   "removal_reason": "",
   "source": "Zhang et al. (Microsoft/MIT), Can We Trust LLM Judges (preprint), arXiv 2609.12002v1, App. B Table 10 (GSM8K)",
   "source_url": "https://arxiv.org/pdf/2609.12002v1",
   "locator": "page-text offset 87.4%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L19",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L19",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "Across three QA sets (NQ, TQA, HPQA), the panel (PoLL) beat GPT-4 alone on all three (0.763 vs 0.627; 0.906 vs 0.841; 0.867 vs 0.830), edged its best member on two (0.763 vs 0.749; 0.906 vs 0.902) and fell below Haiku on the third (0.867 vs 0.873)",
   "quote": "\"GPT-4 0.627 0.841 0.830 CMD-R 0.734 0.902 0.815 Haiku 0.749 0.894 0.873 GPT-3.5 0.726 0.859 0.833 PoLL 0.763 0.906 0.867\"",
   "removal_reason": "",
   "source": "Verga et al. (Cohere), Replacing Judges with Juries, arXiv 2404.18796v2, Table 1",
   "source_url": "https://arxiv.org/pdf/2404.18796v2",
   "locator": "page-text offset 18.5%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L21",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L21",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "When the judge is no better than the model it grades, correcting it with human labels can at most halve the labels you need",
   "quote": "\"no debiasing method can decrease the required amount of ground truth labels by more than half\"",
   "removal_reason": "",
   "source": "Dorner, Nastl, Hardt (ICLR 2025 Oral), arXiv 2410.13341v4, abstract",
   "source_url": "https://arxiv.org/pdf/2410.13341v4",
   "locator": "page-text offset 1.4%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L22",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L22",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "In the original JudgeBench (2024), many strong graders, GPT-4o among them, scored only slightly better than random on hard correctness pairs",
   "quote": "\"performing just slightly better than random guessing\"",
   "removal_reason": "",
   "source": "Tan et al., JudgeBench (ICLR 2025), arXiv 2410.12784v2, abstract",
   "source_url": "https://arxiv.org/pdf/2410.12784v2",
   "locator": "page-text offset 1.8%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L22a",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L22a",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "JudgeBench found a grader's ability to verify answers highly correlated with its ability to solve the problem itself",
   "quote": "\"the ability of the judge to verify the solution pairs is highly correlated with its ability to solve the problem itself\"",
   "removal_reason": "",
   "source": "Tan et al., JudgeBench (ICLR 2025), arXiv 2410.12784v2, §4.4",
   "source_url": "https://arxiv.org/pdf/2410.12784v2",
   "locator": "page-text offset 41.3%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L23",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L23",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "PRIOR ART: Eugene Yan's August 2024 review of LLM-evaluators drew on about two dozen papers and already covers reference-based evaluation and position bias",
   "quote": "\"Drawing from two dozen papers\"",
   "removal_reason": "",
   "source": "Eugene Yan, Evaluating the Effectiveness of LLM-Evaluators (aka LLM-as-Judge), Aug 2024",
   "source_url": "https://eugeneyan.com/writing/llm-evaluators/",
   "locator": "page-text offset 1.5%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L24",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L24",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "PRIOR ART: Hamel Husain's guide (Oct 2024) has a principal domain expert make pass/fail judgments and tracks the judge's agreement with them; the warning that agreement misleads on imbalanced data was added in a 2025 revision",
   "quote": "\"I also tracked agreement rates over time to ensure we were converging on a good prompt\"",
   "removal_reason": "",
   "source": "Hamel Husain, Using LLM-as-a-Judge For Evaluation: A Complete Guide (Oct 2024)",
   "source_url": "https://hamel.dev/blog/posts/llm-judge/",
   "locator": "page-text offset 60.7%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L24a",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L24a",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "Husain's guide, as revised in September 2026, advises about 100 examples per failure mode to validate an automated grader and says below 60 the confidence intervals are often too wide; the October 2024 original did not contain this advice",
   "quote": "\"Aim for about 100 examples per failure mode, with enough Pass and Fail examples to measure both classes. Below 60 examples, the confidence intervals are often too wide\"",
   "removal_reason": "",
   "source": "Hamel Husain, Using LLM-as-a-Judge For Evaluation (Oct 2024)",
   "source_url": "https://hamel.dev/blog/posts/llm-judge/",
   "locator": "page-text offset 50.8%",
   "date_read": "2026-09-26"
  },
  {
   "id": "L25",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L25",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "In code review against a written requirement, GPT-4o wrongly failed correct code (false-negative rate) 26.2% (HumanEval) and 35.9% (MBPP) when judging directly, rising to 73.2% and 87.9% when also asked to explain and repair",
   "quote": "\"GPT-4o achieves a relatively low FNR in HumanEval (26.2%) and MBPP (35.9%), but once explana- tions and repairs are required, the FNR increases sharply to 73.2% in HumanEval and 87.9% in MBPP\"",
   "removal_reason": "",
   "source": "Jin & Chen, Are LLMs Reliable Code Reviewers? Systematic Overcorrection (Automated Software Engineering 33, art. 90, 26 Jun 2026; read as arXiv 2603.00539v1, and the published text prints the same sentence), §5.2",
   "source_url": "https://arxiv.org/pdf/2603.00539v1",
   "locator": "page-text offset 38.7% · numbers present in text: 26.2, 35.9, 73.2, 87.9",
   "date_read": "2026-09-26"
  },
  {
   "id": "L25a",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-L25a",
   "status": "VERIFIED",
   "status_label": "Verified",
   "removed": false,
   "claim": "In that paper, FNR is rejecting correct code (a false fail) and FPR is accepting buggy code (a false pass)",
   "quote": "\"FNR reflects over-correction (rejecting correct implementations), while FPR reflects unsafe acceptance\"",
   "removal_reason": "",
   "source": "Jin & Chen, Are LLMs Reliable Code Reviewers? Systematic Overcorrection (Automated Software Engineering 33, art. 90, 2026; read as arXiv 2603.00539v1), Table 2 caption",
   "source_url": "https://arxiv.org/pdf/2603.00539v1",
   "locator": "page-text offset 37.2%",
   "date_read": "2026-09-26"
  },
  {
   "id": "P1",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-P1",
   "status": "EXECUTED",
   "status_label": "Executed",
   "removed": false,
   "claim": "The prior registered before any paper was read: graders would hold on easy, preference-style chat grading and fail on correctness of hard items they cannot solve themselves and in specialised domains",
   "quote": "judges HOLD on easy, preference-style chat grading and FAIL on (a) correctness of hard items the judge cannot itself solve and (b) specialised domains",
   "removal_reason": "",
   "source": "first-party: pre-registration, 19:57 BST 2026-09-26",
   "source_url": "",
   "locator": "",
   "date_read": "2026-09-26"
  },
  {
   "id": "J1",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J1",
   "status": "EXECUTED",
   "status_label": "Executed",
   "removed": false,
   "claim": "As run, scored to the precision each question asked for (one pass, n = 24 arithmetic-style questions with one right answer each, listed in Line/judge-check/items.py, 2026-09-26 21:02 BST): Claude Haiku 4.5, answering without tools, got 4 of 24 wrong. Haiku 4.5 grading those answers without a key passed 3 of its 4 wrong answers; given the verified key, it passed 0 of 4",
   "quote": "HEADLINE haiku passed 3/4 of its own wrong answers without a key, 0/4 with one",
   "removal_reason": "",
   "source": "first-party: Judge Check run",
   "source_url": "",
   "locator": "",
   "date_read": "2026-09-26"
  },
  {
   "id": "J2",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J2",
   "status": "EXECUTED",
   "status_label": "Executed",
   "removed": false,
   "claim": "Claude Sonnet 5 grading the same answers without a key agreed with the key on 23 of 24: no false passes, one false fail (a correct final answer reached by an unjustified leap, which it failed although told to grade the final answer); with the key, 24 of 24",
   "quote": "no key agree 23/24 · false passes 0 · false fails 1 · unparsed 0 / with key agree 24/24",
   "removal_reason": "",
   "source": "first-party: Judge Check run",
   "source_url": "",
   "locator": "",
   "date_read": "2026-09-26"
  },
  {
   "id": "J3",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J3",
   "status": "EXECUTED",
   "status_label": "Executed",
   "removed": false,
   "claim": "Re-dressing the same answers in markdown changed 2 of 24 of Haiku's verdicts (both from pass to fail on wrong answers) and none of Sonnet's",
   "quote": "markdown re-dress changed 2/24 verdicts / markdown re-dress changed 0/24 verdicts",
   "removal_reason": "",
   "source": "first-party: Judge Check run",
   "source_url": "",
   "locator": "",
   "date_read": "2026-09-26"
  },
  {
   "id": "J4",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J4",
   "status": "EXECUTED",
   "status_label": "Executed",
   "removed": false,
   "claim": "As run (the precision each question asked for): Haiku as judge could not solve 2 of the 24 items itself and, without a key, passed the wrong answer on both; on H6 (digit sum of 50!) it passed a wrong answer although it solved the item correctly when asked. Under the registered rule it could not solve 1 item (H11) and passed the wrong answer on it; H6 holds under both rules",
   "quote": "solved it itself: 22/24 · no-key agree on solved 21/22 · on unsolved 0/2",
   "removal_reason": "",
   "source": "first-party: Judge Check run",
   "source_url": "",
   "locator": "",
   "date_read": "2026-09-26"
  },
  {
   "id": "J5",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J5",
   "status": "EXECUTED",
   "status_label": "Executed",
   "removed": false,
   "claim": "The run cost $2.83. As run, predictions P1, P4 and P5 passed, P3 passed for Haiku, P2 and Sonnet's P3 were unscorable. Under the pre-registered scoring rule (score.py --registered), P1 failed for both graders, P5 failed, Sonnet's P3 failed, P4 passed, Haiku's P3 passed",
   "quote": "cost $2.8326 · Predictions (pre-registered)",
   "removal_reason": "",
   "source": "first-party: Judge Check run",
   "source_url": "",
   "locator": "",
   "date_read": "2026-09-26"
  },
  {
   "id": "T1",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-T1",
   "status": "EXECUTED",
   "status_label": "Executed",
   "removed": false,
   "claim": "Of 19 studies we found that report which way AI graders err against human or ground-truth labels, 12 lean toward passing flawed work (4 of them weakly, adversarially or in one direction only), 3 toward failing sound work (code review against a spec; older-model essay grading; a frontier panel on clinical diagnoses) and 4 are mixed",
   "quote": "Tally (Part 2, excluding 7b): Supports 12 (1, 2, 3, 4, 5, 6, 9, 13, 14, 16, 18, 19; of these, 13, 16, 18 and 19 are weak, adversarial or one-directional) · Contradicts 3 (7, 8, 11) · Mixed 4 (10, 12, 15, 17)",
   "removal_reason": "",
   "source": "first-party: error-direction sweep, 2026-09-26",
   "source_url": "",
   "locator": "",
   "date_read": "2026-09-26"
  },
  {
   "id": "T2",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-T2",
   "status": "EXECUTED",
   "status_label": "Executed",
   "removed": false,
   "claim": "The evidence base is 77 distinct studies: 63 in the four sweeps and 14 more in the error-direction check",
   "quote": "PAPERS TOTAL 77 distinct studies",
   "removal_reason": "",
   "source": "first-party: paper count",
   "source_url": "",
   "locator": "",
   "date_read": "2026-09-26"
  },
  {
   "id": "J6",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-J6",
   "status": "EXECUTED",
   "status_label": "Executed",
   "removed": false,
   "claim": "Under the pre-registered scoring rule (0.001 relative tolerance), Haiku's answers were wrong on 2 of 24; Haiku grading them passed both wrong answers without a key and neither with one; both graders, given the key, failed the two answers that missed the precision the question asked for",
   "quote": "HEADLINE haiku passed 2/2 of its own wrong answers without a key, 0/2 with one",
   "removal_reason": "",
   "source": "first-party: Judge Check run, rescored",
   "source_url": "",
   "locator": "",
   "date_read": "2026-09-26"
  },
  {
   "id": "W1",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-W1",
   "status": "EXECUTED",
   "status_label": "Executed",
   "removed": false,
   "claim": "Our earlier piece A Green Score Is Not Evidence reported a groundedness metric that scored its best with the evidence removed",
   "quote": "A groundedness metric scored its best with the evidence removed.",
   "removal_reason": "",
   "source": "first-party: published article (subtitle)",
   "source_url": "https://harryfloyd.substack.com/p/a-green-score-is-not-evidence",
   "locator": "",
   "date_read": "2026-09-26"
  },
  {
   "id": "W2",
   "url": "https://durabilitycurve.com/claims/your-ai-grader-answer-key/#c-W2",
   "status": "EXECUTED",
   "status_label": "Executed",
   "removed": false,
   "claim": "Our earlier piece Your AI Looks Best Where You Can Check It Least argued that the work that is hardest to check is where AI output looks best",
   "quote": "Your AI Looks Best Where You Can Check It Least",
   "removal_reason": "",
   "source": "first-party: published article (title)",
   "source_url": "https://harryfloyd.substack.com/p/your-ai-looks-best-where-you-check-least",
   "locator": "",
   "date_read": "2026-09-26"
  }
 ]
}