In this post12 sections
  1. Why a model grading a model makes interviewers wary
  2. Position bias, and how to cancel it
  3. Verbosity bias and self-preference
  4. Rubric drift and prompt sensitivity
  5. Calibrating a judge against human labels
  6. Measuring agreement with Cohen’s kappa
  7. Grading what it cannot do, and grading its own loop
  8. When not to use a judge at all
  9. What to say in the interview
  10. Questions people ask
  11. Keep reading
  12. More from the blog

Your works, the demo looks good, and now you need a number. So you reach for a second model to grade the first, and the interviewer asks the obvious question: who grades the grader? This post answers it for the AI round; for the other rounds, start with the FDE interview guide.

The short answer: the main LLM-as-a-judge pitfalls are position bias, verbosity bias, self-preference, rubric drift, prompt sensitivity and run-to-run noise. The fix is to treat the judge as a measuring instrument. Check it against human labels with a chance-corrected agreement score before you trust it, pin its prompt and model, and use plain code wherever an answer can be checked exactly.

Here is the whole list at a glance.

PitfallHow to catch it
Position biasSwap the order
Verbosity biasPad a good answer
Self-preferenceCompare with people on each family’s outputs
Rubric driftRe-run the golden set on any change
Prompt sensitivityReword the prompt, re-run
Run-to-run noiseJudge the same set twice
Grading its own loopKeep a human-labeled holdout

Below: each pitfall with its check, a short script, when to skip the judge, and the words to say.

Why a model grading a model makes interviewers wary

A model judge sounds circular because it partly is. If the system and the grader share a blind spot, the grader passes the mistake, and the score rises while quality stays flat. The failure is quiet: the dashboard looks healthy, and nothing on it tells you the grader is wrong.

Postings name it as part of the job, which makes it fair game in the AI round:

  • Scale AI’s Frontier Agents postings list evaluation harnesses built from offline benchmarks, online experiments, golden datasets, regression suites and LLM-as-a-Judge. Source 1Frontier Agents Engineer (Forward Deployed Engineering)PublisherScale AI (Greenhouse)Source typecompany job posting
  • Cursor’s FDE postings make the FDE own production quality, including evals, debugging model behavior and failure modes. Source 2Forward Deployed EngineerPublisherCursor (Ashby job board)Source typecompany job board
  • Decagon’s Agent Deployment Engineer posting asks for evaluation and regression-testing frameworks that validate agent behavior against ground truth and guard against performance drift. Source 3Agent Deployment Engineer @ Decagon (San Francisco)PublisherDecagon (Ashby job board)Source typecompany job posting

It shows up in take-homes too. One candidate reported, in a public GitHub repository created in August 2026, what the repository calls the Hippocratic AI coding assignment for an AI Agent Deployment Engineer role: turn a simple bedtime-story request into a story for ages 5 to 10 using prompting, and add an LLM judge. Source 4Hippocratic AI Coding AssignmentPublisherakshay-menta (GitHub)Source typecandidate’s take-home repository We use that task as the running example below. If your loop includes a take-home, the post on the OpenAI FDE take-home ends with a routine that works for any format.

Position bias, and how to cancel it

Show a judge two answers and ask which is better, and it may favor whichever sits in a particular slot, regardless of content. The paper by Zheng and colleagues, Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, documents this, along with verbosity bias and self-enhancement bias.

The check is cheap: ask twice, swapping the order, and only accept a verdict that survives the swap.

def judge(first, second):
    # Stand-in for an LLM call.
    # Returns "first" or "second",
    # and always picks the first slot.
    return "first"

def compare(a, b):
    one = judge(a, b)  # a shown first
    two = judge(b, a)  # b shown first
    if (one, two) == ("first", "second"):
        return "a"
    if (one, two) == ("second", "first"):
        return "b"
    return "inconsistent"  # followed order

print(compare("answer A", "answer B"))

It prints inconsistent, because the stand-in judge only ever picks the first slot. With a real judge, the share of inconsistent verdicts is a property of your judge you can report. If it is high, the pairwise setup is not measuring quality.

Three fixes, from cheapest:

  • Swap and reconcile. Count an inconsistent pair as a tie, never as a win.
  • Randomize order across the set, so any leftover bias spreads evenly over both systems.
  • Score one answer at a time against a rubric instead of comparing pairs. Pointwise scoring has no slot to be biased toward.

Verbosity bias and self-preference

Verbosity bias means the longer answer tends to win. For a bedtime story that is the wrong direction: a story that drags on is worse for a tired child, not better.

To test it, take an output your labelers marked as good and pad it: restate the moral, add a paragraph of scenery, repeat the ending. If the judge’s score rises, it is rewarding length. Then:

  • Put length in the rubric on purpose. “Fail if the story repeats a scene or runs past the requested length.”
  • Give the judge a reference answer, so “more” is measured against “enough”.
  • Grade each criterion pass or fail instead of asking for overall quality, which is where length sneaks in.

Self-preference means a judge tends to rate its own outputs, and plausibly its own family’s, more highly (Panickssery and colleagues show evaluators favoring their own generations). The usual fix is a judge from a different family from the system it grades, and a check against people on each family’s outputs.

A take-home may fix the model. One candidate reported, in August 2026, that the bedtime-story assignment says not to change the OpenAI model. Source 4Hippocratic AI Coding AssignmentPublisherakshay-menta (GitHub)Source typecandidate’s take-home repository Whether that covers the judge is worth one clarifying question. If your judge has to share a family with the system, say so in the write-up, and give human spot checks more weight.

Rubric drift and prompt sensitivity

Rubric drift is when the thing you measure changes under you. You tighten a criterion on Tuesday, and the score on Wednesday is not comparable to Monday’s. Or the model provider updates the version behind a floating alias, and the same prompt now grades differently. Either way, a chart that goes up may just mean the ruler shrank.

Research calls one version of this criteria drift: grading outputs changes what you think the criteria should be.

The fix is to version the judge like code. Every score gets stored with the exact judge that produced it:

judge:
  model: pinned-snapshot-id   # never a floating alias
  prompt_version: story-judge-v3
  rubric_version: v3
  temperature: 0
  scale: pass_fail

Pinning does not make the judge repeatable. Even at temperature zero, the same input can get a different verdict on another call. So run the judge twice on the and report how many verdicts flip. That flip rate is the noise floor: a score change smaller than it is not a change.

When any line changes, re-run the frozen golden set with the new judge before you compare numbers across the change. It doubles as a regression suite for the judge.

Prompt sensitivity is the judge’s version of the same problem: small wording changes move scores. Reordering the criteria, changing the example, or switching from a wide number scale to pass or fail can all shift results. Two habits help:

  • Prefer binary criteria with sharp wording. Vague criteria give the judge room to wander.
  • Paraphrase test. Reword the judge prompt without changing its meaning and re-run a sample. If verdicts flip, the rubric is carrying less of the decision than you thought.

Here is the difference sharp wording makes, for the bedtime-story judge:

  • Vague: “Is the story appropriate for young children?”
  • Sharp: “Fail if the story contains injury, death, or a threat to a character that is not resolved before the end. Otherwise pass.”

A person and a model can apply the sharp version the same way.

Calibrating a judge against human labels

A judge’s score means nothing until you know how often it agrees with people. This loop is the one we teach; no employer publishes a process for it.

Calibrate before you trust

  • Sample real outputs, and include known failures on purpose. A sample of all-good outputs cannot show whether the judge catches bad ones.
  • Have people label the sample with the same rubric the judge gets, without seeing the judge’s verdicts.
  • Have two people label an overlapping slice. Their agreement with each other is your reference point: a judge that agrees with them about as well as they agree with each other is doing as well as you can measure.
  • Run the judge on the same items and measure agreement with a chance-corrected score.
  • Read every disagreement. Each one is a vague criterion, a labeling mistake or a real judge error.
  • Fix the rubric or the prompt, then measure again on items you did not tune on.

If you tweak the judge prompt until it matches your calibration labels, you have fitted the judge to those items, and the agreement score flatters it. Keep a held-out slice for the final number.

Measuring agreement with Cohen’s kappa

Raw agreement, the share of items where judge and human match, flatters a judge whenever one label is common. If most stories pass, a judge that passes everything agrees with people most of the time and catches nothing.

Inter-rater agreement scores fix this by subtracting the agreement you would expect by chance. For two raters, Cohen’s kappa is (p_o - p_e) / (1 - p_e), where p_o is observed agreement and p_e is the agreement expected from each rater’s label frequencies alone. No library needed:

def kappa(a, b):
    n = len(a)
    p_o = sum(x == y for x, y in zip(a, b)) / n
    a_yes, b_yes = sum(a) / n, sum(b) / n
    p_e = a_yes * b_yes + (1 - a_yes) * (1 - b_yes)
    return p_o, p_e, (p_o - p_e) / (1 - p_e)

human = [1, 1, 1, 1, 1, 1, 0, 0, 0, 0]  # 1 = pass
judge = [1, 1, 1, 1, 1, 0, 1, 0, 0, 0]
fmt = "obs %.2f  chance %.2f  k %.3f"
print(fmt % kappa(human, judge))

skewed = [1] * 9 + [0]
lazy = [1] * 10  # passes everything
print(fmt % kappa(skewed, lazy))

It prints:

obs 0.80  chance 0.52  k 0.583
obs 0.90  chance 0.90  k 0.000

The first line matches scikit-learn’s cohen_kappa_score on the same labels. The second line is the one to remember: against a set where 9 of 10 stories pass, a judge that passes everything reaches 0.90 raw agreement, higher than the honest judge’s 0.80, and a kappa of 0, because chance would have got every one of its matches right.

Report a pair of extra numbers beside kappa. Of the outputs people failed, how many did the judge fail? Of the outputs people passed, how many did the judge pass? In the toy data above, people failed 4 stories and the judge caught 3; people passed 6 and the judge passed 5. The lazy judge catches 0 of the 1 failure. Call them the catch rate on failures and the keep rate on passes. A judge with a high keep rate and a low catch rate scores well on agreement and still misses what you built it to find.

How to read kappa in the round:

  • Compare it to your people, not to a table. Bands such as Landis and Koch’s are a convention, not a standard. The working bar is the kappa your two human labelers reach with each other on this rubric. A judge that matches that is doing as well as the task allows.
  • Say the sample is small. On a handful of items, one label moves kappa a long way. In the toy data, one extra disagreement can drop kappa from 0.583 to 0.348, and removing one can lift it to 0.783, depending on which label flips. The toy data above proves the arithmetic, not a judge.
  • Break it down by criterion. A judge can agree well on “resolved ending” and poorly on “vocabulary for young children”. Use it only where it earns trust.

Grading what it cannot do, and grading its own loop

A judge that cannot solve a math problem cannot reliably grade the solution. The Zheng paper lists limited reasoning ability among the judge’s weaknesses, next to the three biases above. Give the judge a reference answer to compare against, or move the check to code.

The second trap is quieter. If the same judge drives rewrites, say a generate, judge and revise loop for the bedtime story, or your own prompt tuning, the storyteller is being tuned to please that judge. Its score on the final outputs then flatters the system, because the system learned the judge’s taste, not the reader’s.

The fix:

  • Keep a human-labeled holdout the loop never sees, and never tune against it.
  • Report that number as the headline, not the judge’s score on outputs it helped shape.
  • Say which judge drove the loop in the write-up, so a reviewer knows which scores are self-graded.

When not to use a judge at all

The best judge is often no judge. If code can check the answer exactly, code is cheaper, faster, deterministic, and has no bias to calibrate.

QuestionUse
Is the output valid JSON?Code
Does the total match?Code
Do the tests pass?Code
Is the quote in the source?Code
Is the tone right?Judge
Is the claim supported?Judge
Is it safe for a child?Judge and people

For the bedtime story, length and required fields are code checks. Whether a scene would frighten a young child is a judgment call, so it goes to the judge, with people reviewing a sample every week.

Skip the judge, or put a person in its place, in these cases too:

  • When its error rate is larger than the error you are trying to measure. If the system gets a category wrong rarely and the judge disagrees with people more often than that, the judge’s noise swamps the signal.
  • When a single miss is expensive. For a medical or financial decision, a judge can sort a queue, but a person makes the call.

What to say in the interview

Name the failure modes before you are asked. The lesson on labeling assumptions and surfacing failure modes covers that habit in general; here is how it sounds for a judge.

Split the work.

“I use code for everything checkable: the output parses, the length is in range. For what code can’t check, like whether a scene is too scary, I use an LLM judge with a pass or fail rubric per criterion.”

Prove the judge.

“Two people label a sample, including stories I know are bad. I measure the judge’s kappa against their labels and against their agreement with each other, and I know its flip rate on repeat runs.”

Guard it.

“I swap order for any pairwise comparison and test with padded answers for length bias. The judge model, prompt and rubric are pinned and versioned. Where it disagrees with people most, on vocabulary, a person reviews a weekly sample.”

Then have a sentence ready for each follow-up:

  • “Why not just use a stronger model as the judge?” A stronger judge still has biases, and I would still need to measure its agreement with people. Strength is a guess until the kappa says otherwise.
  • “How many labels do you need?” Enough that the number stops moving. I bootstrap kappa over the labeled items, grow the set until the interval is narrow enough to decide with, and add labels first where the judge is weakest.
  • “The customer asks why they should trust the score.” Show them the agreement with their own experts, and what a person still reviews. The lesson on explaining AI limits to non-technical leaders covers that conversation.

Common mistakes, and the fix for each:

  • Quoting raw agreement. Quote kappa, with the human-to-human figure beside it.
  • Calibrating on all-pass samples. Seed known failures.
  • A floating model alias. Pin the snapshot and store it with every score.
  • Judging what code can check. Move it to code.

The judge is one piece of a larger answer to “how do you know it works?”, which the post on LLM evaluation in the interview walks through. Where the labels come from is the subject of the post on building a golden set for LLM evals.

Practice it where it bites: design an eval harness for a support agent lists “trusts a model grader that was never checked against human labels” as a pitfall, and how do you know it works? is the question this post feeds. Then score model outputs against a ground-truth file with Pro.

Treat the judge as an instrument, and you will have the answer before the interviewer asks.

GlossaryAgentA system in which a model chooses steps and tool calls to complete a task, within limits the design sets.More on AgentGlossaryInter-rater agreementHow consistently two reviewers, or a reviewer and a model judge, label the same items, corrected for agreement by chance.More on Inter-rater agreementGlossaryForward deployed engineerA software engineer who builds and ships production systems inside a customer’s problem and environment, accountable to that customer’s outcome.More on Forward deployed engineerGlossaryGolden setA curated, labeled set of examples, sampled to cover the cases that matter, used to evaluate a system the same way over time.More on Golden set

Questions people ask

What are the main LLM-as-a-judge biases?

Position bias (preferring the answer shown first or second), verbosity bias (preferring longer answers) and self-preference (rating its own outputs, and plausibly its own family’s, higher). Judges are also sensitive to small prompt changes, so pin the prompt and the model version, and expect some verdicts to change between runs even at temperature zero.

How do you validate an LLM judge?

Have people label a sample of outputs with the same rubric, run the judge on the same sample, and measure agreement with a chance-corrected score such as Cohen’s kappa. Read the disagreements, fix the rubric or the prompt, and measure again before you rely on the judge.

When should you not use an LLM judge?

When the answer can be checked in code, such as an exact match, valid JSON, a correct number or a passing test. Use a judge for qualities code cannot check, such as tone or whether an answer is supported by its source, and keep people reviewing a sample.

Keep reading