Guide

How to check an AI data analyst's answer

Ask it questions you already know the answer to. Ask the same question more than once. Read the definition it picked and the SQL it ran. Then get the number a second way.

Shane Butler By Shane Butler Sep 26, 2026 Updated Sep 26, 2026

The checks are listed cheapest first. Run the cheap ones on any number you pass along, and all five when the number is going somewhere that matters. They work on whatever you use: ChatGPT's Data agent, Claude Code on a warehouse, or the open-source analyst we maintain. We ran the first two tools side by side on ChatGPT vs Claude for data analysis.

You need them because a wrong number looks like a right one. When we sent one stakeholder question about active customers to four fresh ChatGPT chats, we got three different numbers, and no chat asked what active meant. That test is written up on why AI gives different answers.

1. Ask questions you already know the answer to

Start with numbers your company already trusts: reconciled finance figures, the metrics the board sees every quarter. Ask the AI analyst for those and compare.

A list of questions with verified answers is a golden set. The first one can be a handful of questions in a text file. Keep the question, the verified answer, the definition and the query together, so when it misses you can see which step missed. When a question that drives a decision has no trusted answer, work it out once by hand and add it. That is slow, and it is what gives you a real accuracy number for the questions you care about.

Copy this once per question and fill in the brackets, or download the template file.

- question: "[the question, in the words a stakeholder uses]"
  verified_answer: "[the number]"
  verified_by: "[who checked it, and against what: finance close, board deck]"
  definition: "[numerator, denominator, grain, filters, time window]"
  query: |
    [the SQL that produced the verified answer]
  tolerance: "[how far off still counts as a match]"

Be careful with numbers that are only familiar. A metric can be wrong for years while a whole company quotes it.

This check catches wrong joins, wrong filters, wrong tables and a metric defined differently from yours. It misses every question you have no answer for, which is most of the ones worth asking.

2. Ask the same question more than once

Open fresh chats, turn memory off and send the same words. It needs no answer key, so it is cheap. It also catches the least.

If five runs give five numbers, the metric is undefined, and the runs have just listed the candidate definitions for you. That gap between runs is the reliability spread. If five runs agree, the system has settled on one answer, and a wrong query settles just as well as a right one.

A wrong query can be perfectly stable. It'll hand you the same wrong number all day if you execute that query over and over again.

From the recorded session Validate Claude Code Analytics Output, June 17, 2026.

In that session we ran "how are we doing on checkout conversion" five times against a warehouse where the definition sat in a metric dictionary. All five runs came back at 33.2 percent. So the system was consistent with one definition. Whether that was the definition the stakeholder wanted is a separate question.

3. Read the definition it chose

When you do not supply a definition, the analyst chooses its own. Most say which somewhere: in the answer, in the working notes, or only in the SQL. Find it and check whether your company defines the metric the same way.

The fix is to write the definition down where the analyst reads it before it computes anything. We call that a metric contract: the metric's name, numerator, denominator, grain (what one row is), filters and time window. What happened when we tested one on three models is on why AI gives different answers.

This check catches an analyst answering a different question from the one you meant. It misses a query that computes the agreed definition badly.

4. Read the SQL

Ask for the query behind the number and read it the way you would read a junior analyst's query. Look for three things first. A join that changes the grain: an order-level amount summed after a join to line items gets counted once per line. A missing status filter, so cancelled and returned orders count as revenue. A window you did not expect, like a partial month at the end of a trend, or 90 days when you meant 30.

It helps when the query gets written down whether or not anyone asks. In the public ai-analyst repo, a hook writes every Snowflake query the analyst runs to a log, with the date, the tables it touched and the rows it returned. That record is the analysis trace. Without it, an analyst asked "what did you run?" guesses from whatever is left in its context.

This check catches a wrong query. It misses a correct query on the wrong definition, and anything wrong in the data itself.

5. Get the number a second way

Compute the same number by a different route and see whether the two meet: bottom-up from line items against top-down from the orders table, a second model, a dashboard the team already trusts. That is triangulation. When the routes meet, the number is stronger, though still not proven. When they miss, the gap tells you where to dig.

In our tests of ChatGPT's Data agent, we asked it to rebuild August revenue bottom-up from line items. It returned $230,284.84 and said nothing about the $278,993.42 it had been reporting all week. Asked whether the two reconcile, it produced a reconciliation to the cent: scope, fees, taxes, tips, discounts, refunds. Both numbers were right for what they measured. The silence about the gap was the failure, and the second route is what exposed it.

A $6.9 million total that should have been $3.6 million

In the June 24 session Pressure-Test Any AI Analysis, Shane pasted in a prepared revenue-by-category query of the kind an agent writes, on NovaMart, a synthetic e-commerce dataset. Its total was $6.9 million. Claude Code, set up with an audit skill and a semantic layer, was asked whether the number could be trusted.

It found the join first. The orders table has about 47,000 rows and order_items about 75,000. The query joined the two and summed total_amount, a column on the orders table, once for every line item. Every order with more than one item was counted more than once.

…our total comes out as $6.9 million, when actually it should be something more like $3.6 million. So the query overstates the revenue by 87%.

Then it found two definition problems against the semantic layer. There was no status filter, so cancelled and returned orders counted as revenue, roughly another 15 percent too high if the stakeholder wanted completed orders only. And total_amount includes the delivery charge and is net of discounts, so line_total was the better column. The audit ended with a corrected query, which we ran.

I've seen, like, human analysts make that mistake. I've made that mistake myself.

Reading the SQL found the join. Reading the definition found the filter and the column. A known answer would have caught it too, if finance had a reconciled revenue figure for the period. A repeat run would not have, because the same query returns the same $6.9 million on every run.

Other questions

How do I check an AI data analyst's answer when I have no answer key?

Checks 2 to 5 need no reference: repeat runs, the definition, the SQL and a second route. None of them proves the answer. Then build an answer key for the handful of questions that drive decisions, starting from numbers finance has already reconciled.

How much checking does a number need?

Match it to the stakes. A curious stakeholder's side question gets one run and a glance at the definition. A number going to the board, or into a decision that is expensive to reverse, gets all five checks and a second person reading the trace.

Is this the same as evaluating AI data analyst tools?

No. Tool evaluations compare products on features such as schema understanding, SQL transparency and governance. This page checks a single answer from whichever tool you use. The checks also work as a test when you are choosing: run the same known-answer questions through each candidate and read the traces.

We teach these checks in week 3 of Agentic Analytics: Build an AI Analyst.

AI Analytics for Everyone
$1,800 · Oct 19 · ★ 4.9/5
Enroll on Maven
Agentic Analytics: Build an AI Analyst
$2,500 · Nov 2 · ★ 4.9/5
Enroll on Maven
Or come to a free workshop this Wednesday. Register free