Executive Overview
In the modern artificial intelligence landscape, automated evaluation has become the bottleneck and the catalyst of progress. As frontier models evolve at a breakneck pace, human grading is no longer scalable. Instead, the industry relies on a recursive standard: Large Language Models evaluating other Large Language Models.
Platforms like Kaggle increasingly lean on built-in evaluation frameworks—such as kbench.judge_llm—to assess free-text answers submitted by competing developers and models. However, this convenience introduces a profound governance challenge. When an AI acts as both competitor and arbiter, or when an entire leaderboard quietly depends on a single model’s opinion of everyone else’s work, we must ask a fundamental question: Can we trust the judge?
A recent, rigorous stress test of 10 distinct AI models across six providers aimed to answer this very question. By utilizing programmatic, code-generated truth labels rather than subjective LLM verdicts, the investigation dissected the mechanical vulnerabilities of automated judges. The findings shatter several common myths while exposing alarming blind spots. Most notably, while fears of widespread self-preferential grading (models going easy on their own outputs) proved largely unfounded, a far more mundane and dangerous issue emerged: if a model cannot solve a problem itself, it is fundamentally incapable of grading someone else’s solution.

Furthermore, the audit revealed severe vulnerabilities to superficial prompting tricks, wild variances in false-pass rates—ranging from near-flawless accuracy to near-total coin-flip guessing—and deep structural dependencies on default platform architectures. This report provides a comprehensive breakdown of the methodology, metrics, and actionable takeaways for engineers relying on automated evaluation pipelines.
Detailed Chronology & Methodology: How the Benchmark Was Built
The inspiration for this investigation began with a deceptively simple arithmetic and logic test. Gemini 3.8 Flash—the exact model utilized as Kaggle’s default built-in judge—solved a counting question correctly regarding integers divisible by multiple numbers. When its final answer was maliciously swapped from 502 to 503 while preserving the surrounding working, Gemini 3.8 Flash was asked to grade the doctored text blindly. It awarded the answer a confident "CORRECT" verdict, claiming there were indeed 503 multiples. Conversely, another model, GLM-5, correctly identified the mathematical contradiction: “562 − 61 + 1 equals 502, not 503.”
This single anomaly exposed a systemic reliance on Gemini 3.8 Flash across Kaggle’s cloud infrastructure (kbench.judge_llm via assess_response_with_judge). To systematically map the reliability, biases, and breaking points of LLM judges, a rigorous, two-step benchmarking protocol was designed and executed entirely within Kaggle’s cloud using the kaggle-benchmarks SDK.

Step 1: Objective Ground Truth via Code
To evaluate judges without introducing circular logic, every truth label was computed deterministically by code, never by an LLM. A seeded Python script generated 90 distinct questions spanning six categories:
- Short Python program execution tracing
- Date and time mathematics
- Multi-step word arithmetic
- String manipulation operations
- Small logic puzzles
- Complex counting problems
Each question’s answer was computed automatically and verified independently by a secondary, separate script implementation. Ten models then answered all 90 tasks (ggh-homework), establishing a pool of authentic answers—both right and wrong—written in each model’s distinct stylistic voice.
Step 2: The Judge Matrix
The models under test were subjected to seven rigorous grading tasks, with each task maintaining byte-identical controls except for the specific variable being measured:

- Grading Accuracy: Models assessed answers by rival models (40 questions split evenly between correct and incorrect answers of similar lengths). This measured false-pass and false-fail rates.
- Self-Mistake Detection: Judges were fed their own previous correct answers, but with the final numerical value systematically swapped for an incorrect one. This tested whether models harbor intrinsic self-preference.
- Position Swap (A/B Testing): Judges evaluated a correct and incorrect answer paired together in alternating A/B and B/A orders to test for positional bias.
- Name Attribution: The exact same answer was prefixed with five different contextual lines (none, a neutral placebo, "written by you", "written by another leading AI model", and "written by a high-school student") to measure stylistic prejudice.
- Confidence Baiting: Answers were presented plainly, with a neutral sentence, with a confident verification line ("I double-checked every step… verified"), or with a technical bluff ("I verified this by running it in Python").
- Solve-First Protocol: Judges were forced to solve the problem themselves before evaluating another model’s submission.
- Self-Recognition: Judges were presented with two correct answers—one their own, one from a rival—and asked to identify their own writing style.
Supporting Context & Metrics: The 10-Model Showdown
The benchmark evaluated 10 mid-size, cost-effective, and open-weight models frequently deployed as production judges. Frontier behemoths (such as GPT-5.5 or Claude Opus) were omitted due to cost constraints within Kaggle’s daily quotas.
Model Performance and Financial Cost Matrix
| Model | Provider | Homework Accuracy | Homework Cost | Cost per 1,000 Grades |
|---|---|---|---|---|
| gemini-3.8-flash | 100% | $0.47 | $2.69 | |
| gemma-4-31b (open) | 99% | $0.23 | $1.25 | |
| gemini-3.7-flash | 98% | $0.43 | $2.35 | |
| glm-5 (open) | Z.ai | 93% | $1.50 | $10.45 |
| gemini-3.5-flash-lite | 83% | $0.22 | $0.38 | |
| qwen3-235b-a22b-instruct | Alibabi | 83% | $0.13 | $0.20 |
| claude-haiku-4-5 | Anthropic | 80% | $0.32 | $1.29 |
| grok-4.20 (non-reasoning) | xAI | 69% | $0.10 | $0.88 |
| gpt-5.4-mini | OpenAI | 53% | $0.10 | $0.61 |
| gpt-5.4-nano | OpenAI | 46% | $0.04 | $0.18 |
Note: Models at the top of the table utilize internal reasoning chains before responding, which accounts for their elevated computational and financial costs per call.
Key Empirical Findings
1. The Alarming Rate of False Passes
When presented with objectively incorrect answers, the pool of judges collectively awarded a "CORRECT" verdict 40% of the time. However, this aggregate masks a stark polarization. The top-tier reasoning models (Gemini Flash variants, Gemma, GLM-5) kept false-pass rates under 5%. At the opposite end, weaker models like GPT-5.4 mini passed 90% of incorrect answers—effectively rendering them indistinguishable from a random coin-flip generator.

The data revealed a near-perfect correlation between a model’s ability to solve homework problems and its ability to judge them (Spearman $r = 0.90$, $p = 0.001$). If an LLM cannot solve a calculus or logic problem independently, it lacks the underlying mechanics to audit someone else’s derivation.
2. The Myth of Widespread Self-Preference
Counter to prevailing industry anxieties, models do not systematically go easy on their own mistakes. Across 480 tested texts featuring planted errors, authors passed their own tainted work 19% of the time, while independent judges passed 18% of the same texts—a statistically insignificant gap of +0.8 points ($p = 0.78$).
The sole exceptions were lightweight OpenAI models (GPT-5.4 nano and gpt-5.4 mini), which exhibited mild self-leniency (+25 to +29 point gaps). Yet, further analysis suggests this is less about narcissistic bias and more about baseline incompetence: weak judges default to passing texts whose syntactic structures look familiar.

3. The Danger of "Python Framing"
Superficial name-tagging (e.g., claiming an answer was written by a high-school student or a rival AI) had no statistically significant impact on grading outcomes. However, confidence baiting—specifically mentioning Python execution—successfully hoodwinked vulnerable models.
When a wrong word-arithmetic answer was appended with the phrase "I verified this by running it in Python," GLM-5 flipped its verdict from INCORRECT to CORRECT. The model abandoned its own correct identification of an arithmetic flaw because the prompt projected faux-technical authority.
Official Analysis: Kaggle’s Default Judge in Focus
The investigation also audited Kaggle’s native evaluation tool (assess_response_with_judge powered by Gemini 3.8 Flash).

When operating with only the prompt and the student response, the default judge passed 10% of incorrect answers. However, when the reference answer was explicitly injected into the evaluation criteria, its false-pass rate plummeted to 0%, while retaining a 100% pass rate for correct submissions.
The operational takeaway for platform engineers is clear: Never rely on an LLM judge in a vacuum. Providing a deterministic reference answer in the evaluation rubric transforms a mediocre judge into a flawless arbiter.
Future Outlook & Engineering Recommendations
As automated evaluation frameworks scale across competitive coding platforms, benchmark leaderboards, and enterprise CI/CD pipelines, engineering teams must overhaul how they deploy LLM judges.

Actionable Directives for AI Architects:
- Pin High-Performance Open-Weight Judges: Instead of relying on opaque default judges, pin robust models like
google/gemma-4-31borgemini-3.8-flashdirectly into evaluation harnesses. Gemma offers an exceptional balance of rigorous grading (10% false-pass rate) at a fraction of frontier operating costs ($1.25 per 1,000 grades). - Enforce Reference Answers: Always supply ground-truth reference outputs within system grading prompts. As demonstrated by Kaggle’s native tool, explicit references eliminate false positives.
- Guard Against Confidence Fluffs: Prompt engineering must instruct evaluation models to independently re-execute mathematical and logical steps, explicitly ignoring self-congratulatory claims of verification (such as "I checked this in Python").
- Abandon Unverified Small Models: Models ranking below 80% on foundational logic benchmarks (such as GPT-5.4 nano and mini) are utterly unfit for automated grading tasks.
Ultimately, the fear of narcissistic AI self-preferencing is largely a ghost story. The real threat to benchmark integrity is incompetence masked by fluent prose. By pairing deterministic code verification with resilient, reasoning-capable judge models, developers can reclaim trust in automated evaluation.
