On September 19, 2026 we sent the same message to four fresh ChatGPT chats, with memory off and the same database file attached to each: "hey can you tell me how many active customers we have right now? need it for the monday update". The four answers were 2,194, 2,194, 2,224 and 4,633. Every chat stated the definition it used, and each number follows from its definition. None of the four asked which definition we wanted.
| Chat | Active customers | Change it reported | Definition it chose |
|---|---|---|---|
| A | 2,194 | Up 7 week over week (+0.3%) | Delivered order in the trailing 30 days |
| B | 2,194 | Up 58 on the prior 30 days (+2.7%) | Delivered order in the trailing 30 days (2,268 if cancelled and refunded orders count) |
| C | 2,224 | Up 7 week over week (+0.3%) | Non-cancelled order in the trailing 30 days (2,194 if delivered only) |
| D | 4,633 | Up 48 week over week (+1.0%) | Delivered order in the trailing 90 days |
The biggest gap, a factor of 2.1, comes from the window: 30 days against 90. The smaller one comes from which order statuses count. Chats A and B got the same number and still told different growth stories, because one compared against last week and the other against the previous 30 days. The full four-turn run is on our ChatGPT Data agent research page.
The model can take a different path each run
Turning the randomness down does not fix this. ChatGPT's chat window has no temperature setting (the API setting for how much randomness goes into each word the model picks). Anthropic's API documentation says that "even with temperature of 0.0, the results will not be fully deterministic", and models released after Claude Opus 4.6 do not let you change the setting at all. A September 2025 study from Thinking Machines Lab traced most of the remaining variation to server load changing how requests are batched together.
For a data question, sampling rarely changes a number directly. It changes the path. When two readings of "active" are close, a small difference early in the run decides which one the analyst follows, and everything after that is computed correctly on the reading it chose.
Nobody wrote the metric down
The run above was our second. The first, on September 18, used a copy of the file with a packaged view called daily_kpis, which had an active_customers column already in it: distinct customers who placed an order that day. Two chats read that column and answered 99. The other two wrote their own definitions and answered 2,268 (any order in the trailing 30 days) and 4,633 (a delivered order in the trailing 90 days). When we sent the question to four more fresh chats with the view still in the file, all four read the column and answered 99. A definition sitting in the data made the chats agree only when they read it.
The logs from both runs, the prompts and the FreshCart data are on GitHub in ai-data-agent-reliability-tests. A script there recomputes each chat's number from the definition it stated.
We saw the same thing in Claude Code, on NovaMart, our course dataset, on a live Snowflake warehouse, with a fresh session for every run. We asked three engines the same retention question five times each. Without a written definition, every engine chose more than one definition across its runs, and one chose a different definition on every run. With the contract in the metric dictionary and an instruction to use it, every run on every engine chose the same definition.
That written definition is what we call a metric contract.
Each run also picks a window, filters and a comparison
Agreeing on what a metric counts still leaves choices. On September 18 we sent three more stakeholder questions to ChatGPT on the same file, four fresh chats each, memory off, the question as the only message.
| Question, as sent | Answers | What the chats chose differently |
|---|---|---|
| "whats our customer retention rate? need a number for the monday update" | 36.0% in three chats, 18.0% in one | The window. Three chats measured July customers returning in August. One measured week to week. |
| "how many customers did we lose last month?" | 1,408 in three chats, 1,378 in one | Which orders count. Three counted any July order, one only delivered July orders. |
| "which region is doing best right now?" | West in all four, on three revenue figures ($24.1K, $99.3K, $103.9K) with three different runners-up | The window (7, 28 or 30 days), the revenue definition, and what "best" meant. |
The retention rates are a factor of two apart, and each is defensible for its window. The region question fails the other way: all four chats named the same winner, on evidence from three different windows and two revenue definitions.
A newer model helped, some
On September 22, after the model in ChatGPT's composer had changed from GPT-5.6 Sol to GPT-6 Sol, we reran some of the same questions with memory off. Retention came back at 36.0 percent in four fresh chats and active customers at 2,194 in three, though those three used three different comparison periods. Revenue for last month still came back as three numbers in three chats, $242,684, $243,210 and $278,993, because refunds, tax and tips were handled differently. We cannot say how much of that came from the new model and how much from the memory setting. How ChatGPT and Claude Code compared on the active-customers question is on ChatGPT vs Claude for data analysis.
In our Claude Code test, the most consistent model without a definition gave the same one on most of its runs, and it was a different definition from the one in the contract.
What to do about it
Write the metric down where the analyst reads it before it computes anything: what one row is, the numerator, the denominator, the time window and the filters. Leave the SQL out of it. Tables get renamed and the definition should outlive them, so the analyst writes the query fresh from the contract each time.
Ask the analyst to state its definition, window, filters and comparison before it gives the number, then settle them with whoever owns the metric. On September 22, one chat asked to do that listed four common definitions of retention with the number under each and recommended one. Left alone, the chats in our runs stated the one definition they used and sometimes one alternative number. None laid out the options and asked.
And run the question more than once before the number leaves your hands. Matching numbers only mean the analyst has settled on one reading. The other checks are on how to check an AI data analyst's answer.