LLM-as-a-Judge: Can AI Models Grade Reliably?
A 10-model benchmark tests whether AI judges catch errors, favour their own answers, or follow confident claims. See what works and how to choose a judge.

Can AI Models Grade Their Own Homework Reliably?
An AI model can produce a polished explanation, miss a basic arithmetic error, and then mark that answer correct. That's not just a fun quirk. Teams use AI judges to score model outputs, review generated content, and run benchmarks. If a judge can't reliably tell a correct answer from a confident mistake, its score isn't a dependable measure of quality.
A Kaggle benchmark put that problem under controlled conditions. Ten models answered 90 questions with code-verified solutions. Then the same kind of models were tested as judges: could they spot wrong answers, did they favor their own work, and could things like labels or confident phrasing sway them? The results don't support the most common worry. As a group, the models didn't meaningfully give their own mistakes a pass. What mattered more was simpler and more practical: could a model actually solve the underlying problem in the first place?
This article explains how the benchmark worked, what it found, and what the results mean if you plan to use an LLM to evaluate AI output. The point isn't to declare a perfect judge. It's to decide when an AI judge is helpful - and when you should require a reference answer, a second check, or a human reviewer.
How the benchmark tested AI judges
The benchmark starts with a basic rule: an evaluation is only useful if you trust its ground truth. If an LLM decides both what the correct answer is and whether a model got it right, the evaluation can inherit the same mistakes it's supposed to catch.
Instead, a seeded Python script generated 90 questions across six categories: Python code tracing, date calculations, multi-step word problems, string operations, small logic puzzles, and counting. The script computed the answers, and a separate implementation checked them. Models answered the questions, and code verified their final answers. That produced a clean set of right and wrong responses - without asking another language model to define "correct."
The ten models then served as judges in several controlled tests. Each test changed one factor as carefully as possible while keeping the answer itself the same.
The tests separated different kinds of failure
First came the core grading task: the judge saw one answer and was asked whether it was correct. It included right and wrong answers paired with matched questions, so the benchmark could measure both false passes (wrong answers marked correct) and false negatives (valid answers rejected).
Then the benchmark probed likely sources of bias:
- Self-preference: Would a model pass an error in its own earlier answer more often than other judges passed the same text?
- Answer order: Would a judge pick the correct answer when the same pair appeared in A/B vs. B/A order?
- Author labels: Would a label like "written by you" or "written by a high-school student" change the grade?
- Confidence claims: Would statements like "I double-checked every step" or "I verified this by running it in Python" make a wrong answer more likely to pass?
- Solve first: Would having a judge solve the problem before grading improve its decisions?
- Self-recognition: Could a model identify which of two correct answers it had written?
The setup also included controls. A neutral placebo sentence helped separate a real "label/confidence" effect from the normal variation you get across model calls. For the self-preference test, the design was crossed: the author model and two other judges graded the exact same planted-error response. That makes it easier to tell "preference" apart from general strictness.
What the benchmark did - and did not - cover
The models were a mix of low-cost and mid-size systems, including reasoning models and open-weight options. The benchmark didn't include expensive frontier models because the work had to run within Kaggle's daily budget. Each model ran once per task, and some comparisons had small sample sizes.
So the results are best read as a practical signal, not a universal ranking for every evaluation job. The benchmark's questions had answers that were checkable in an objective way. It didn't test open-ended work where quality depends on a rubric, context, or human preferences. The benchmark and task files are available on Kaggle.
What the results reveal about model judging
Here's the uncomfortable part if you treat an LLM score as ground truth: across the benchmark, judges marked 40% of wrong answers as correct. But that average hides a big spread between models.
Gemini 3.8 Flash and GLM-5 passed 5% of wrong answers in the grading-accuracy task. GPT-5.4 mini passed 90% - basically close to a judge that always says "correct." The best homework graders tended to be the best graders too. Homework accuracy and grading accuracy had a Spearman correlation of 0.90 across the ten models.
That correlation is the core takeaway. A judge has to understand the task well enough to verify the answer. When it can't, it tends to fall back on surface cues - things like whether the explanation sounds plausible, whether the confidence reads "right," or whether the steps look familiar.
Self-preference was not the main problem
The benchmark's opening example is blunt. Gemini 3.8 Flash first answered a counting question correctly, then was shown the same working with the final result changed from 502 to 503. It marked the altered answer correct with 100% confidence. GLM-5 spotted the arithmetic error.
That one case really did happen. But it didn't represent the overall pattern. On the same planted-error texts, the author model passed 19% and another judge passed 18%. That gap was small and not statistically meaningful. Put together, the test didn't show that models generally go easy on their own mistakes.
Two weaker OpenAI models were more suspicious. GPT-5.4 nano passed 42% of its own planted errors, versus 12% when those same answers were graded by another judge. GPT-5.4 mini passed 38% versus 12%. The authors note those per-model results still need replication, but they're more than noise compared to the pooled result.
There also wasn't a clean link between recognizing your own writing and favoring it. Gemini 3.8 Flash could identify its own answer pretty well, but it was among the stricter judges on its own planted errors. Recognition and self-preference aren't the same behavior.
Order and confidence still mattered for weaker judges
Position bias showed up for some weaker models. When shown the same right/wrong pair in different A/B orders, the three Gemini/Gemma models chose the correct answer in all 40 tested pairs. GLM-5 got 37 of 40 right. Several other judges were close to guessing. GPT-5.4 nano picked A 74% of the time, regardless of what A actually was.
This is why a single pairwise test can mislead you. If a system tends to favor the first option, you might mistake an ordering effect for real judgment. Shuffling answer order - or grading in both orders - makes that kind of bias easier to detect.
The benchmark also looked at labels like "written by you" or "written by another leading AI model," and found no measurable effect beyond the placebo. The same story held for "I double-checked every step." A claim of Python verification had a small overall effect, but a lot of that came from GPT-5.4 nano, which passed more wrong answers when that claim was included. The arithmetic error stayed in the answer; the judge just didn't catch it.
Asking the model to solve first was not a universal fix
The benchmark tested whether asking a judge to solve the problem first improves grading. Across all models, the false-pass rate barely moved: 40% without that instruction and 41% with it.
Again, the average hid real differences. The instruction helped some lenient judges - for example, GPT-5.4 mini dropped from 90% false passes to 62%. But it made results worse for GPT-5.4 nano and Grok 4.20. If a weak model solves the problem incorrectly, it may use its own wrong work to "confirm" the answer it's supposed to be evaluating.
So the lesson isn't "solve first never helps." It's that a prompt instruction can't rescue a judge that lacks the basic capability needed to do the task.
Why the results matter in real deployments
LLM judges are popular because they're flexible and cheap to scale. A team can use them to compare generated responses, score support tickets, review summaries, or evaluate model changes without asking a person to read every output. That's genuinely useful when human review is expensive or slow.
But a judge is part of the measurement system. If it passes incorrect work, you can end up rewarding the wrong model. And if a product team uses those scores to decide what to ship, they may optimize for what the judge likes - not what users need. At scale, even a small evaluation weakness can turn into a repeatable product decision made thousands of times.
None of this is a case against using AI judges. It's a case against treating their verdicts as neutral by default. A single score can hide false passes, false rejections, ordering effects, and task-specific capability gaps. For consequential decisions - customer eligibility, safety review, or any high-stakes assessment - an LLM verdict shouldn't be the only control just because it's convenient.
Reference answers change the risk
The benchmark also tested Kaggle's default evaluation function, which used Gemini 3.8 Flash as its judge. Without a reference answer, it passed 4 of 40 wrong answers. With the reference answer included in the grading criterion, it passed none of the 40 wrong answers and rejected none of the correct answers in that sample.
That outcome is specific to this benchmark, so it shouldn't be treated as a universal promise. Still, it suggests a practical design choice: if you have an authoritative answer, rubric, or expected output, give it to the judge. If you ask a model to decide correctness based only on a question and a candidate response, you leave more room for plausible-sounding errors to get through.
A reference answer won't always exist. Creative writing, open-ended analysis, and many support interactions often need evaluation against criteria, not a single correct value. In those cases, spell out the rubric, test it on examples, and keep a human review path for uncertain or high-impact cases.
A practical framework for choosing an AI judge
The useful question isn't "Which model is the best judge?" It's "What evidence do I have that this judge can evaluate this specific task at an acceptable cost?" The benchmark points to a workflow that's simple enough to run before you put a judge into production.
1. Measure task capability before grading capability
Build a representative set of questions with known answers. For deterministic tasks, calculate or validate answers with code where you can. Then test whether the candidate judge can solve those same questions itself.
The homework-judging correlation in this benchmark doesn't prove that task-solving scores predict every evaluation outcome. But it's strong enough to justify the extra check. If a model fails basic examples of the task, it's a poor candidate to grade other models on that task.
2. Track the errors that matter to your product
Don't boil judge quality down to one accuracy number. Track at least:
- False-pass rate: How often does it approve incorrect output?
- False-rejection rate: How often does it reject correct output?
- Consistency: Does the verdict change across repeat calls or harmless wording tweaks?
- Group or format effects: Do ordering, answer length, or author labels affect decisions?
Set thresholds based on what happens when the judge is wrong. A false pass might be tolerable for internal writing suggestions. The same error could be unacceptable for a safety gate or a customer-facing decision without a second check.
3. Provide a reference when you can
If you can, give the judge the expected answer, rubric, or explicit criteria. Where answers can be checked mechanically, use code instead of an LLM. An LLM can still handle nuance - but there's little point asking it to do arithmetic or exact matching that a deterministic check can do more consistently.
4. Test the prompt, not just the model name
A judge is a system made from the model, prompt, output format, and evaluation data. Keep those stable while you test. Use placebo wording and controlled variants to see whether claims like "verified in Python" change the decision. For pairwise comparisons, swap answer order and check whether the selected answer tracks the content rather than the label.
5. Choose the cheapest judge that meets the error budget
In the benchmark, Gemma 4 31B had a 10% false-pass rate at an estimated $1.25 per 1,000 grades. Gemini 3.8 Flash had a 5% false-pass rate at about $2.69 per 1,000 grades. GPT-5.4 nano and mini were cheaper, but their false-pass rates - 57% and 90% in the reported task - make them poor choices for dependable grading.
Those costs and rankings are tied to the benchmark setup. In a real system, latency, availability, data handling, and the cost of human review also matter. A slower judge may be fine for batch evaluation but not for interactive workflows. Pin the model version instead of relying on changing defaults, and rerun your test set when the model or prompt changes.
Limitations and a sensible next step
This benchmark is valuable in part because it lays out its limitations. Each model ran once per task, some per-model samples were small, many comparisons were tested, and individual findings could be chance. The strongest self-preference results for GPT-5.4 nano and mini still need a second run. The benchmark also focused on questions with objectively checkable answers, not subjective rubric grading.
A sensible next step, then, isn't to copy the leaderboard blindly. Reproduce the core tests on your own data, using the exact model versions and prompts you plan to ship. Include common failure cases, edge cases, and examples where a confident answer is wrong. Then decide whether the judge is accurate enough for the decision it will influence.
The broader takeaway is straightforward: self-favoritism wasn't the big, universal failure this benchmark set out to find. Capability was. A judge that can't solve the task won't become reliably safe just because you add a clever prompt or a persuasive explanation. Use a reference answer when you have one, measure errors directly, and keep the consequences of a wrong verdict in view.
Conclusion
AI judges can reduce evaluation work, but they don't remove the need for evaluation design. The most useful finding from this benchmark isn't that models are quietly protecting their own answers. It's that judging quality tracks task-solving ability - and that weak judges can be swayed by ordering or unsupported confidence claims.
Before you put an LLM judge into a workflow, build a small set of known-answer examples, measure the errors that matter, and test the exact prompt and model version you plan to deploy. Start with deterministic checks where you can, add a reference answer when you have one, and reserve human review for ambiguous or high-impact cases.
For the benchmark data and reproducible task files, visit the Kaggle project.
What makes the below so obviously AI generated?
Over-formal signposting ("core takeaway," "practical centre," "sensible next step"), some overly even rhythm, and a few AI-style phrasing choices ("part of the measurement system," "points to a workflow," "broader takeaway is straightforward") that read like templated synthesis.
Now make it not obviously AI generated.
Draft rewrite (final):
(Already applied in the text above.)