Definition

AI judge (LLM judge)

An AI judge is a model that grades another model's output against a written rubric, so you can score work with no single right number.

By Shane Butler, Co-founder, AI Analyst Lab. Published Sep 26, 2026, updated Sep 26, 2026.

In analytics a judge has a narrow job. Revenue, retention and conversion have right answers, and those go to a golden set. The judge gets what has no single right number, like whether a chart supports its title or whether a memo's recommendation follows from its own numbers.

A judge is a model grading a model, so before you trust it you check its verdicts against a person's labels. That check is called calibration.

The rubric: seven pass or fail rules for a chart

This rubric is from the chart judge exercise in our course. Each rule is a yes or no with a concrete anchor on each side, and a chart passes only when every rule passes. The first five rules came from fixing one chart, one flaw at a time. The last two were added after the first calibration run.

RulePassFail
action-title Pass: The title states the takeaway with a direction or size, like "Revenue grew 447% in 2024" Fail: The title only names the axes, like "Revenue by Month"
focus-color Pass: One deliberate accent color on a warm or neutral background Fail: Default blue, or no intentional focus color
direct-label Pass: The final value is labeled on the end of the line Fail: The reader has to read the value off the axis
decluttered Pass: No top or right border, light gridlines, no marker on every point Fail: A full box border, heavy gridlines, or a dot on every point
zero-baseline Pass: The value axis starts at zero Fail: It starts above zero and exaggerates the differences
title-accurate Pass: The size the title claims matches the data Fail: The title exaggerates
clean-numbers Pass: Numbers are rounded sensibly Fail: Absurd precision, like cents on a value in the millions

The judge prompt

This is judge.md from the exercise, minus a logging instruction at the end. Blind means the judge does not see earlier scores or the human labels.

You score one chart against the rubric
in RUBRIC.md. Run BLIND: do not read
scores.csv before you score. For each
axis in RUBRIC.md:
1. Reason one sentence about the chart
   against that axis's pass/fail anchors.
2. Return a verdict: pass or fail.

Rules:
- One pass/fail per axis. Never a 1-5.
- A chart's overall verdict is pass only
  if every axis passes.
- Judge the chart only; do not invent
  data you cannot see.

The judge passed a chart claiming 900% growth

The labeled set is 14 line charts of monthly revenue for NovaMart, the synthetic online store we use in the course. We built the charts with known flaws and labeled them from those flaws: 4 pass and 10 fail. We added two of the failing charts on purpose, so that a judge built from the first five rules would pass them. One says "Revenue exploded 900% across 2024" when revenue grew about 447%. The other prints every number to the cent.

Line chart of NovaMart monthly revenue in 2024, rising steadily from January to about $0.5M in December, titled Revenue grew 447% across 2024, with one orange line and the final value labeled at its end.
Labeled pass. It passes all seven rules.
The same line chart of NovaMart monthly revenue in 2024, with the title changed to Revenue exploded 900% across 2024. Revenue grew about 447%.
Labeled fail. Same data, and a title that claims 900%.

We ran the judge on all 14 charts on Jul 1, 2026, while building the exercise, and a script lined its verdicts up against the labels.

RunRulesJudge passedAgreed with labels
Run 1First five612 of 14
Run 2All seven414 of 14
Run 3, repeatAll seven414 of 14

In run 1 the judge passed all 4 good charts and the 2 planted ones (recall 1.00, precision 0.67). The action-title rule only checked that the title stated a takeaway, and nothing checked that the takeaway was true. No rule said anything about rounding. The judge was doing what the rubric told it to.

Adding two rules closed the gap. Run 2 used all seven, failed both charts and agreed with every label. Run 3, a few minutes later, came back the same.

We knew the answers before we started, so perfect agreement here shows the method works. It does not show this judge is right about your charts. On your own outputs the disagreements are the useful part: each one is a rule you had not written yet or a label you got wrong. A rubric that agrees with you on charts has not been tested on memos, so each new rubric needs its own labeled set.

A judge is one check among several. How to check an AI data analyst's answer covers the rest, and why AI gives different answers covers why a rerun can move.

In the courses

In Agentic Analytics: Build an AI Analyst, week 3 (AI evals for AI analytics), you build a chart judge rule by rule and check it against a labeled set. The free AI evals course covers judges for AI features in general, and how to build an AI data analyst shows where the judge sits in the system.

Questions people ask about AI judges

What is an AI judge?

An AI judge is a model that grades another model's output against a written rubric, so you can score work with no single right number.

When should I use an AI judge instead of a golden set?

Use a golden set wherever the answer can be recomputed, like a revenue total or a conversion rate. Use a judge only where there is no single right number: whether a chart supports its title, whether a memo's recommendation follows from its own numbers. Golden set

How do you calibrate an AI judge?

Label a set of outputs by hand as pass or fail. Run the judge on the same set without showing it the labels. Line up its verdicts against yours, read every disagreement, and fix the rubric or the label. Rerun until the disagreements stop teaching you anything new.

How many human labels do I need before I trust the judge?

Our teaching set is 14 charts, which is enough to show the method and too small to trust for a real decision. The rule of thumb we teach for any eval set is 50 to 75 cases per kind of output, taken from real work. The set also has to contain the failures you care about, or the judge will agree with you for the wrong reason.

What agreement rate is good enough?

In our first run the judge agreed on 12 of 14 charts, and both misses were charts a person failed and the judge passed. That direction is the costly one, because a chart that passes rarely gets a second look. Read every disagreement. Each one either becomes a new rule or shows that the label was wrong.

Why pass or fail instead of a 1 to 5 score?

Two people can agree on a yes or no with a concrete anchor. On a 1 to 5 scale, a 3 and a 4 is an argument neither side can settle. For more detail, add more pass or fail rules.

Is the judge consistent from run to run?

Test it. Score one output five times and look for any rule whose verdict changes. A rule that flips is worded too loosely for the model to read the same way twice, so you sharpen its pass and fail wording and run it again. We have not published run-to-run numbers for this judge.

Next cohorts start Oct 19 and Nov 2.

AI Analytics for Everyone
$1,800 · Oct 19 · ★ 4.9/5
Enroll on Maven
Agentic Analytics: Build an AI Analyst
$2,500 · Nov 2 · ★ 4.9/5
Enroll on Maven
Or come to a free workshop this Wednesday. Register free