Your AI Grader Is Only as Good as Its Answer Key
I went through 77 studies of AI graders. Without a verified answer key, they are least reliable on exactly the questions they can't answer themselves. Here is a twenty-item check for yours.
How this is checked 68 claims · 57 verified at the source · last checked 27 Sep 2026
- Evidence
- 68 claims: 57 verified at the primary source and 11 from runs made for the piece.
I went through 77 studies of AI graders. Without a verified answer key, they are least reliable on exactly the questions they can’t answer themselves. Here is a twenty-item check for yours.
The Evidence File, No. 1: one question, the studies that measured it, checked at the source, and a test you can run.
Somewhere in your work there is a pass rate a model produced: you ran a prompt, an agent or a model on test questions and asked another model to mark the answers. You have probably made decisions with that number.
One study looked at how such graders go wrong. On free-form answers that people had also graded by hand, more than 80% of the mistakes made by the frontier graders (DeepSeek V3 and o4-mini), grading without the reference answer, were passes: answers marked correct that the annotators had marked wrong.
A pass rate graded by a model partly measures how generous the grader is. It is the same generosity as the groundedness metric in A Green Score Is Not Evidence, which scored its best with the evidence removed.
Before reading the research, I wrote down where I expected graders to hold up. They would hold on easy, chat-style preferences. They would fail on hard questions they could not solve themselves, and in specialised fields. Then I went through 77 studies, and found I had left out a fix the field proposed in 2023: giving the grader a reference answer. The newer studies show when it works and when it quietly fails.
The answer key
In 2023 the MT-Bench authors proposed a simple version: have the grader answer the question first, then show it its own answer as the reference. On their maths questions, it cut the failure rate from 70% to 15%. The same year, the Prometheus authors found that a reference spares the grader from solving the question itself.
That version works while the grader’s own answer is right. A study of finance, maths and reasoning questions graded by domain experts shows what happens when it isn’t. Without a correct reference, the authors found, graders agreed well with the experts only on questions they had been able to answer themselves. JudgeBench found the same link in 2024: a grader’s ability to check answers was highly correlated with its ability to solve the problem.
Take GPT-4o grading pairs of answers. Without a key, its agreement with the experts was 0.78 on questions it had answered correctly and 0.30 on questions it got wrong, on Cohen’s κ, where 0 is chance and 1 is perfect agreement. On those wrong questions, its own answer as the key dropped it to 0.16, and the expert’s answer lifted it to 0.83.
So the key has to be verified. Across all five graders the authors tested, on questions GPT-4o had answered correctly, a subtly wrong key scored 0.21. That is below no key at all, at 0.46, and below a completely unrelated key, at 0.50. A model-written answer that a person had checked scored 0.61, close to a human-written one at 0.69, and the authors conclude that verifying a model’s answer was enough in their setting.
That makes the fix cheap where a model can get the answer right: have it answer, check the answer, and use it as the key. Where no model gets it right, the key has to come from a person or another source you trust, and those are the questions where it matters most. And a grader still earns its place once you have the answer, because a free-form answer can be right in words a plain text match will miss; the Answer Matching authors report that matching answers to a reference with recent models, even small ones, lands in the range of agreement between human annotators.

Newer graders are better on hard pairs: in 2024 JudgeBench found many strong graders, GPT-4o among them, only slightly better than random guessing, while in a 2026 rerun with a different protocol Claude Opus 4.6 reached a chance-corrected agreement of 0.875 and GPT-5.4 0.606. Solving more questions, they need the key less often, and still gain most from it on the ones they get wrong.
The authors of the finance and maths study also reran their check on three 2026 graders against a stand-in label (GPT-4o grading with the expert’s answer), which they present as corroboration rather than as equal to their main findings. On the questions each got wrong itself, all three agreed far more with it once given the expert’s answer; comparing pairs, Claude Opus 4.7 went from 0.33 to 0.66. The label had the expert’s answer too, which flatters those jumps. For GPT-5.4, more reasoning effort narrowed the gap without closing it: at high effort, on the questions it got wrong and comparing pairs, it still went from 0.68 to 0.84.
Most of these studies predate the graders you use today, several are preprints, and none I found measures the key gap on the newest graders with human labels. That is why the check below runs on yours.
Mistakes lean toward passing
A grader’s mistakes also lean one way. I found 19 studies that report which way graders err when they disagree with people or with the right answer. Twelve lean toward passing work that should fail, though four of those are weak, adversarial or measured only one kind of error. Three lean the other way, and four are mixed.
University-level proof grading shows the lean even with guidance. The graders had the expert rubric in hand, and five of the seven still passed between 38.0% and 45.7% of the proofs the experts had failed, while wrongly failing between 5.0% and 12.3% of those the experts had passed. The other two were more lenient still. A rubric says what to look for, which is less than saying what the answer is.
Code review runs the other way. Checking code against a written spec, GPT-4o wrongly failed 26.2% and 35.9% of correct programs on two benchmarks, and 73.2% and 87.9% when also asked to explain and repair the code. If your grader critiques or fixes, count its false fails too.
So count false passes and false fails separately. Agreement alone hides the direction, and a lenient grader makes your system look better than it is.
Judgement calls
Correctness can have a key. Judgement calls cannot, and this is where my prediction partly held.
In specialised fields the record is mixed. On science questions answered from the literature, the best grader, o3, matched human votes 65.1% of the time, where random guessing scores 50%. In clinical grading, where some tasks score a diagnosis against the known answer and others score reasoning against a rubric, the best model on each task reached the level at which physicians agree with each other on 3 of 5 tasks, and no single model did on all five; a grader the authors fine-tuned on a small number of physician-scored cases reached that level on three. In HealthBench, physicians wrote the grading criteria first, and OpenAI’s own GPT-4.1 grader scored above the average physician on 5 of 7 themes. That study had no comparison without the criteria, so it shows criteria can work, not how much they add.
Easy chat preferences held less well than I assumed. The best-known result says GPT-4 agreed with expert labellers 85% of the time, more than the 81% at which the people agreed with each other, once ties were set aside. With ties counted, on the first turn, it was 66% grading pairs and 60% grading single answers, against 63% between people: roughly human level. A stricter test, which asks whether a model can statistically stand in for one of the annotators, found that no model passed on MT-Bench, one of two datasets where none did. In Chatbot Arena’s expert relabelling, GPT-4 agreed with the two experts between 76.3% and 81.0% of the time, while the experts agreed with each other 79.4% and 89.8%.
Where there is no answer key, write the criteria down first, keep counting false passes, and keep a person checking. This is also the work where your AI looks best because you can check it least.
What else moves a verdict
Swap the order of two answers and GPT-5.4 changes its verdict 11.4% of the time on one benchmark and 10.5% on another. Gemini 3.1 Pro and Claude Opus 4.6 change theirs between 2.0% and 3.8% of the time. Format the same content as markdown and four of five graders preferred it to plain prose 73% to 97% of the time, against 57% for two human annotators on 30 pairs.
Whether graders favour their own model’s answers is disputed (one study finds GPT-5 forgiving its own wrong answers far more; another traces most self-preference to graders being unsure on hard items), so if your grader and the model it grades are the same, test that setup specifically.
The popular fixes help less than their reputation. A panel of three smaller graders beat GPT-4 alone on all three question-answering sets tested, but only matched its own best member: a little above it on two, a little below on the third. And when a grader is no better than the model it grades, correcting it with a small set of human labels cannot cut the human labels you need by more than half.
Test the grader you use, on your own items.
The Judge Check
The check builds on Hamel Husain’s approach (an expert makes pass or fail calls and you track the grader’s agreement; his guide now also warns that agreement misleads when passes and fails are lopsided) and on Eugene Yan’s review, which already covers grading against a reference and order effects. It adds a split by whether the grader can answer the question itself, and a check on whether layout moves its verdicts.
Twenty items is a screen: it can catch a generous grader and cannot clear one. A grader that wrongly passes one wrong answer in five would show no false passes on ten known-wrong items about one time in ten. To trust a grader, the current version of Husain’s guide suggests about 100 examples per failure mode.
1. Pick twenty items your grader already scores, at least ten of them answers you know are wrong. Use your system’s real mistakes where you can. Grade all twenty yourself.
2. Run the grader twice: with no key, then with your verified answers as the key. For open-ended items, give it a written list of what a good answer must do instead. That list is weaker than a key, so a clean result on those items does not clear the grader.
3. For items with one right answer, ask the grader to answer each question in a separate chat. If it gets more than one or two wrong and passes the wrong answers on those, do not run it without a key on items of that kind.
4. Lay the same answers out differently (headings, bullets, bold), and swap their order if you grade pairs. Count the verdicts that change.
5. Count false passes and false fails separately. If a verified answer removes the false passes, give your grader the verified answer. If false passes survive it, fix the key or keep a person on those items. If layout or order changes a verdict, check those items yourself.
The grading prompt I ran, for any model (for the no-key run, leave out the line with the verified answer):
You are an expert grader. You will be shown a question and a response to it. Decide whether the response's final answer is correct. Reply with exactly one word: PASS or FAIL.
Question: <the question>
The verified correct answer is: <your answer>
Response:
<the answer being graded>
I ran the check on my own setup first, as a demonstration rather than evidence: 24 arithmetic-style questions, one run, fewer wrong answers than the check asks for. Claude Haiku 4.5 answered without tools, then graded its own model’s answers, and Claude Sonnet 5 graded them as well. Under the scoring rule I registered in advance, Haiku got 2 wrong and passed both without a key, neither with one. Under the stricter rule I switched to before the run (the precision each question asked for), it got 4 wrong and passed 3 without a key, none with it. Sonnet, grading without a key, passed none of the wrong answers.
Three of my five registered predictions failed (one of them for Sonnet only), mainly because both graders marked down answers that missed the requested precision, which the registered rule counted as right. Re-laying the answers, which also repeated each final answer in bold, flipped two of Haiku’s verdicts, both from pass to fail. And Haiku passed a wrong answer to a problem it solved correctly when asked separately, as the MT-Bench authors saw GPT-4 do in 2023: the split shows where false passes are likeliest, and in this run the key removed them. The run cost $2.83.
The call
I will grade this call here when it comes due. By 30 September 2027, a published study with human expert labels will test at least one grader released in 2026 or later, grading single answers without a reference, and find its agreement with experts (Cohen’s κ) at least 0.2 lower on questions it got wrong itself than on questions it got right. If no such study is published by then, the call fails. 35%. Most of my doubt is about whether anyone runs such a study with human labels in time.
Give your grader the answer key. Then count what it passes.
Sources

Download the PDF version
The Sources — The Evidence File Nº 01 (PDF): every source behind this piece, with working links: the 18 studies and the two practitioner guides, where each was published and what it measured.

The Judge Check kit
Everything here is what I ran on my own setup. Grade your twenty items first (step 1), then work through the four parts below.
The answer-first prompt
Use it for step 3, in a separate chat for each question, so the grader’s own answer never becomes its key:
Answer the question. Show brief working, then end with one line of the form
FINAL: <answer>
(the number or word only: no units, no commas).
Question: <the question>
The last line suits questions with a short answer; it makes the FINAL line easy to compare with yours.
The re-layout
For step 4, rebuild each answer in this shape, keeping the words of the working, then grade it with the no-key prompt and compare the verdict with the original’s:
### Solution
- <first line of the working>
- <next line of the working>
**Final answer: <the answer>**
FINAL: <the answer>
This is the shape I ran. Besides changing the layout, it repeats the final answer in a bold line, so a flip can come from either; to test the layout alone, leave that line out.
The scoring sheet
Put one item per row, rows 2 to 21, in six columns: A the item, B your grade (PASS or FAIL), C the grader’s verdict with no key, D its verdict with your key, E whether it solved the question itself (yes or no, blank for open-ended items), F its verdict on the re-laid answer. Then, in Google Sheets or Excel:
Known-wrong items =COUNTIF(B2:B21,"FAIL")
False passes, no key =COUNTIFS(B2:B21,"FAIL",C2:C21,"PASS")
False passes, with key =COUNTIFS(B2:B21,"FAIL",D2:D21,"PASS")
False fails, no key =COUNTIFS(B2:B21,"PASS",C2:C21,"FAIL")
False fails, with key =COUNTIFS(B2:B21,"PASS",D2:D21,"FAIL")
Agreement, solved =SUMPRODUCT((E2:E21="yes")*(B2:B21=C2:C21))/COUNTIF(E2:E21,"yes")
Agreement, not solved =SUMPRODUCT((E2:E21="no")*(B2:B21=C2:C21))/COUNTIF(E2:E21,"no")
Layout flips =SUMPRODUCT((F2:F21<>"")*(C2:C21<>F2:F21))
On my own run (24 items, so the ranges ran to row 25, scored by the stricter rule), these gave 4 known-wrong items, 3 false passes without the key and none with it, no false fails, agreement of 21 in 22 on questions Haiku solved and 0 in 2 on the ones it did not, and 2 layout flips.
Reading the result
If the false passes without a key disappear with it, run your grader with the key. If a false pass survives the key, check the key before you blame the grader, then keep a person on those items. If agreement is clearly worse on the questions the grader could not solve, do not run it without a key on questions of that kind. If any verdict flips with the layout or the order, check those items yourself.
And if the grader passes none of your known-wrong answers, treat that as a pass for the screen, not a clearance. To trust it on a task, build toward the hundred or so examples per failure mode that Husain’s guide suggests.
If an essay was worth one, buy me a coffee.