In analytics a judge has a narrow job. Revenue, retention and conversion have right answers, and those go to a golden set. The judge gets what has no single right number, like whether a chart supports its title or whether a memo's recommendation follows from its own numbers.
A judge is a model grading a model, so before you trust it you check its verdicts against a person's labels. That check is called calibration.
The rubric: seven pass or fail rules for a chart
This rubric is from the chart judge exercise in our course. Each rule is a yes or no with a concrete anchor on each side, and a chart passes only when every rule passes. The first five rules came from fixing one chart, one flaw at a time. The last two were added after the first calibration run.
| Rule | Pass | Fail |
|---|---|---|
| action-title | Pass: The title states the takeaway with a direction or size, like "Revenue grew 447% in 2024" | Fail: The title only names the axes, like "Revenue by Month" |
| focus-color | Pass: One deliberate accent color on a warm or neutral background | Fail: Default blue, or no intentional focus color |
| direct-label | Pass: The final value is labeled on the end of the line | Fail: The reader has to read the value off the axis |
| decluttered | Pass: No top or right border, light gridlines, no marker on every point | Fail: A full box border, heavy gridlines, or a dot on every point |
| zero-baseline | Pass: The value axis starts at zero | Fail: It starts above zero and exaggerates the differences |
| title-accurate | Pass: The size the title claims matches the data | Fail: The title exaggerates |
| clean-numbers | Pass: Numbers are rounded sensibly | Fail: Absurd precision, like cents on a value in the millions |
The judge prompt
This is judge.md from the exercise, minus a logging instruction at the end. Blind means the judge does not see earlier scores or the human labels.
You score one chart against the rubric
in RUBRIC.md. Run BLIND: do not read
scores.csv before you score. For each
axis in RUBRIC.md:
1. Reason one sentence about the chart
against that axis's pass/fail anchors.
2. Return a verdict: pass or fail.
Rules:
- One pass/fail per axis. Never a 1-5.
- A chart's overall verdict is pass only
if every axis passes.
- Judge the chart only; do not invent
data you cannot see. The judge passed a chart claiming 900% growth
The labeled set is 14 line charts of monthly revenue for NovaMart, the synthetic online store we use in the course. We built the charts with known flaws and labeled them from those flaws: 4 pass and 10 fail. We added two of the failing charts on purpose, so that a judge built from the first five rules would pass them. One says "Revenue exploded 900% across 2024" when revenue grew about 447%. The other prints every number to the cent.
We ran the judge on all 14 charts on Jul 1, 2026, while building the exercise, and a script lined its verdicts up against the labels.
| Run | Rules | Judge passed | Agreed with labels |
|---|---|---|---|
| Run 1 | First five | 6 | 12 of 14 |
| Run 2 | All seven | 4 | 14 of 14 |
| Run 3, repeat | All seven | 4 | 14 of 14 |
In run 1 the judge passed all 4 good charts and the 2 planted ones (recall 1.00, precision 0.67). The action-title rule only checked that the title stated a takeaway, and nothing checked that the takeaway was true. No rule said anything about rounding. The judge was doing what the rubric told it to.
Adding two rules closed the gap. Run 2 used all seven, failed both charts and agreed with every label. Run 3, a few minutes later, came back the same.
We knew the answers before we started, so perfect agreement here shows the method works. It does not show this judge is right about your charts. On your own outputs the disagreements are the useful part: each one is a rule you had not written yet or a label you got wrong. A rubric that agrees with you on charts has not been tested on memos, so each new rubric needs its own labeled set.
A judge is one check among several. How to check an AI data analyst's answer covers the rest, and why AI gives different answers covers why a rerun can move.
In the courses
In Agentic Analytics: Build an AI Analyst, week 3 (AI evals for AI analytics), you build a chart judge rule by rule and check it against a labeled set. The free AI evals course covers judges for AI features in general, and how to build an AI data analyst shows where the judge sits in the system.