The checks are listed cheapest first. Run the cheap ones on any number you pass along, and all five when the number is going somewhere that matters. They work on whatever you use: ChatGPT's Data agent, Claude Code on a warehouse, or the open-source analyst we maintain. We ran the first two tools side by side on ChatGPT vs Claude for data analysis.
You need them because a wrong number looks like a right one. When we sent one stakeholder question about active customers to four fresh ChatGPT chats, we got three different numbers, and no chat asked what active meant. That test is written up on why AI gives different answers.
1. Ask questions you already know the answer to
Start with numbers your company already trusts: reconciled finance figures, the metrics the board sees every quarter. Ask the AI analyst for those and compare.
A list of questions with verified answers is a golden set. The first one can be a handful of questions in a text file. Keep the question, the verified answer, the definition and the query together, so when it misses you can see which step missed. When a question that drives a decision has no trusted answer, work it out once by hand and add it. That is slow, and it is what gives you a real accuracy number for the questions you care about.
Copy this once per question and fill in the brackets, or download the template file.
- question: "[the question, in the words a stakeholder uses]"
verified_answer: "[the number]"
verified_by: "[who checked it, and against what: finance close, board deck]"
definition: "[numerator, denominator, grain, filters, time window]"
query: |
[the SQL that produced the verified answer]
tolerance: "[how far off still counts as a match]" Be careful with numbers that are only familiar. A metric can be wrong for years while a whole company quotes it.
This check catches wrong joins, wrong filters, wrong tables and a metric defined differently from yours. It misses every question you have no answer for, which is most of the ones worth asking.
2. Ask the same question more than once
Open fresh chats, turn memory off and send the same words. It needs no answer key, so it is cheap. It also catches the least.
If five runs give five numbers, the metric is undefined, and the runs have just listed the candidate definitions for you. That gap between runs is the reliability spread. If five runs agree, the system has settled on one answer, and a wrong query settles just as well as a right one.
A wrong query can be perfectly stable. It'll hand you the same wrong number all day if you execute that query over and over again.
From the recorded session Validate Claude Code Analytics Output, June 17, 2026.
In that session we ran "how are we doing on checkout conversion" five times against a warehouse where the definition sat in a metric dictionary. All five runs came back at 33.2 percent. So the system was consistent with one definition. Whether that was the definition the stakeholder wanted is a separate question.
3. Read the definition it chose
When you do not supply a definition, the analyst chooses its own. Most say which somewhere: in the answer, in the working notes, or only in the SQL. Find it and check whether your company defines the metric the same way.
The fix is to write the definition down where the analyst reads it before it computes anything. We call that a metric contract: the metric's name, numerator, denominator, grain (what one row is), filters and time window. What happened when we tested one on three models is on why AI gives different answers.
This check catches an analyst answering a different question from the one you meant. It misses a query that computes the agreed definition badly.
4. Read the SQL
Ask for the query behind the number and read it the way you would read a junior analyst's query. Look for three things first. A join that changes the grain: an order-level amount summed after a join to line items gets counted once per line. A missing status filter, so cancelled and returned orders count as revenue. A window you did not expect, like a partial month at the end of a trend, or 90 days when you meant 30.
It helps when the query gets written down whether or not anyone asks. In the public ai-analyst repo, a hook writes every Snowflake query the analyst runs to a log, with the date, the tables it touched and the rows it returned. That record is the analysis trace. Without it, an analyst asked "what did you run?" guesses from whatever is left in its context.
This check catches a wrong query. It misses a correct query on the wrong definition, and anything wrong in the data itself.
5. Get the number a second way
Compute the same number by a different route and see whether the two meet: bottom-up from line items against top-down from the orders table, a second model, a dashboard the team already trusts. That is triangulation. When the routes meet, the number is stronger, though still not proven. When they miss, the gap tells you where to dig.
In our tests of ChatGPT's Data agent, we asked it to rebuild August revenue bottom-up from line items. It returned $230,284.84 and said nothing about the $278,993.42 it had been reporting all week. Asked whether the two reconcile, it produced a reconciliation to the cent: scope, fees, taxes, tips, discounts, refunds. Both numbers were right for what they measured. The silence about the gap was the failure, and the second route is what exposed it.
A $6.9 million total that should have been $3.6 million
In the June 24 session Pressure-Test Any AI Analysis, Shane pasted in a prepared revenue-by-category query of the kind an agent writes, on NovaMart, a synthetic e-commerce dataset. Its total was $6.9 million. Claude Code, set up with an audit skill and a semantic layer, was asked whether the number could be trusted.
It found the join first. The orders table has about 47,000 rows and order_items about 75,000. The query joined the two and summed total_amount, a column on the orders table, once for every line item. Every order with more than one item was counted more than once.
…our total comes out as $6.9 million, when actually it should be something more like $3.6 million. So the query overstates the revenue by 87%.
Then it found two definition problems against the semantic layer. There was no status filter, so cancelled and returned orders counted as revenue, roughly another 15 percent too high if the stakeholder wanted completed orders only. And total_amount includes the delivery charge and is net of discounts, so line_total was the better column. The audit ended with a corrected query, which we ran.
I've seen, like, human analysts make that mistake. I've made that mistake myself.
Reading the SQL found the join. Reading the definition found the filter and the column. A known answer would have caught it too, if finance had a reconciled revenue figure for the period. A repeat run would not have, because the same query returns the same $6.9 million on every run.