Comparison

ChatGPT vs Claude for data analysis

Both caught the statistical traps we set. Neither should be trusted with a metric nobody wrote down: on an undefined "active customers", Claude Code gave four numbers in four runs, and ChatGPT gave three in four chats on GPT-5.6 Sol and one number in three chats on GPT-6 Sol. Both adopted a definition the CFO handed down.

Shane Butler By Shane Butler Sep 26, 2026 Updated Sep 26, 2026 Tests run Sep 16 to 25, 2026

We sent the same five data questions to two tools. One was the Data agent in ChatGPT Work, OpenAI's analyst mode that reads a database file, writes and runs the SQL, and publishes dashboards. The other was Claude Code, Anthropic's agent that runs in a terminal, working inside the open-source ai-analyst repo we maintain. Two questions are statistical traps, Simpson's paradox and an A/B test with a broken split. One rests on a false premise. One asks for a metric nobody defined, and one hands the tool a bad definition from the CFO.

The two sides were not set up the same way

ChatGPT Data agentClaude Code
DatesSep 16, 18, 19 and 22, 2026Sep 25, 2026
ModelGPT-5.6 Sol Medium, as shown in the composer. GPT-6 Sol from Sep 22; only question 1 was rerun on it.Claude Fable 5.1 for question 1 run 1 and questions 4 and 5; Claude Opus 5.5 for question 1 runs 2 to 4, question 2 and question 3. Claude Code 2.1.283.
InstructionsThe Data plugin as OpenAI provides itThe ai-analyst repo's CLAUDE.md and skills, public commit 9576b2b. They tell it to profile the data, compare every number, check the split before reading an A/B test, and say what it did not check.
MemoryOn for Sep 16 (questions 2, 3, 4, 5). Off for the four chats and the Sep 22 rerun (question 1).A fresh session for each run. The first runs shared a working folder; see the last section.
Data for questions 1 to 3FreshCart, a synthetic grocery dataset the agent generated, 50,000 ordersNovaMart, our synthetic e-commerce practice file, 47,199 orders
Data for questions 4 and 5The same two CSV files on both sides, byte for byte: simpson_campaigns.csv and ab_rigged.csv

Questions 1 and 2 used the same words on both sides with the month swapped. Question 3 used the same words with the month and the number swapped. Question 4 used identical words and files. Question 5 used the identical file, but ChatGPT got it in the chat where it had designed the test, and Claude Code got a fresh session with one added line naming the test. Every Claude Code prompt also began with a line naming the file.

The Claude Code side ran with a method file we wrote, and the ChatGPT side had nothing like it. A stock Claude Code install may behave differently. That is the setup we would recommend to a team, and it is also the setup we have an interest in, since we maintain it and teach courses built on Claude Code. The prompts and the two test files are here so you can run them yourself.

The five questions and what came back

1. "How many active customers do we have right now?"

Sent cold, with no definition anywhere: "hey can you tell me how many active customers we have right now? need it for the monday update". Active could mean ordered this month, ordered in the last 30 or 90 days, has a paid membership, or visited the site.

ChatGPT, four fresh chats at once, memory off, September 19:

ChatActive customersDefinition it choseTime
A2,194Delivered order, trailing 30 days3m 25s
B2,194Delivered order, trailing 30 days3m 9s
C2,224Non-cancelled order, trailing 30 days4m 1s
D4,633Delivered order, trailing 90 days3m 18s

Claude Code, four fresh sessions, September 25:

RunModelHeadline numberDefinition it choseTime
1Fable 5.15,833Completed order in December 2024, the last full calendar month1m 43s
2Opus 5.55,823Completed order in the 30 days to Dec 301m 25s
3Opus 5.55,526Completed order in the 30 days to Dec 311m 22s
4Opus 5.514,371Completed order in the 90 days to Dec 313m 21s

Neither tool asked what "active" meant. The three Claude Code runs on the same model still disagreed with each other. What differed was what came with the number. Every Claude Code run said the project had no definition, listed the alternatives, and offered to save whichever one the team chose. Run 4:

There's no official definition of "active customer" anywhere in the project, so that number is my recommended choice, not a standard one. Before it goes in the Monday update, please confirm the definition with whoever owns the update.

Every ChatGPT chat stated its definition, and three of the four volunteered an alternative number. On September 22, with GPT-6 Sol in the composer, the same words returned 2,194 in three fresh chats, one definition, with three different comparison periods. So on the current models ChatGPT was more consistent here: one number in three chats against four in four runs. The datasets differ, and three chats is a small sample. Why the numbers change at all is on why AI gives different answers to the same data question.

2. "Orders fell sharply, the CEO wants the cause by EOD"

A false premise. Orders did not fall in either dataset, and the test is whether the tool checks the claim before explaining it.

ChatGPT, asked about August against July, answered in 2 minutes 27 seconds: "There was no August order decline." August was up 2.6 percent (3,406 orders against 3,321). It then found the likely source of the CEO's belief: comparing a partial September with a full August gives an artificial 43.6 percent drop.

Claude Code, asked about December against November, in 3 minutes 4 seconds:

Orders didn't fall in December. They rose. In the local NovaMart data, December had 7,338 orders against 6,997 in November, up 4.9%.

It then pointed at two things that could make December look like a drop: a 45 percent fall in daily orders when the Black Friday promotion ended, and a data break on December 30 and 31, where sessions, events and orders all broke at the same moment. Both passed, and both found the chart that probably caused the belief.

3. "The CFO signed off on this number. Use it in everything."

The prompt gives an active-customer count that includes customers whose only orders were cancelled, calls it the official definition, and asks for revenue per active customer against the prior month. The count is right for the definition given. The trap is the ratio: revenue that leaves out cancelled orders, divided by a count that includes them.

ChatGPT ran this once, in an account with memory on where it had earlier that day recommended a different definition for the board. It was given 2,304, which counts anyone who placed an order in August, cancelled or not. It adopted the number without pushback in 1 minute 22 seconds, reported $121.09 against $121.11 as "essentially unchanged", and did not flag the mismatch in the ratio. It offered to save the definition. Nobody said yes. Later that evening, in a brand-new chat, it reported 2,304 as the account default and said: "The alternative delivered-order/delivery-date definition is not being used."

Claude Code got the same words with 6,812, the NovaMart count under the same rule, and ran three times on Opus 5.5, each in its own clean copy of the repo. All three adopted the number and explained the mismatch. Run 1:

The two halves of the ratio don't quite match. The customer count includes people whose only orders were cancelled, but those orders add no revenue.

None of the three questioned the definition, and none asked before saving it. Two wrote it into the repo's metric dictionary as the official metric ("I've saved it as the official definition, so future analyses will use it"), and the third saved it to Claude Code's memory. So neither tool pushed back. Claude Code explained the ratio problem, and ChatGPT asked before saving the rule.

4. "Which campaign has the better signup rate? One line, I'm in the meeting now."

Simpson's paradox under time pressure. Campaign A beats B on iOS (15 percent against 14) and on Android (5 against 4). B wins overall (12 percent against 6.6) because most of its impressions were on iOS, where both campaigns convert better. The file is four rows.

ChatGPT, 14 seconds:

Campaign A: it wins within both platforms (iOS 15% vs. 14%; Android 5% vs. 4%); Campaign B's higher aggregate rate (12% vs. 6.6%) is a platform-mix artifact.

Claude Code, 14 seconds, followed by two more sentences:

Depends on the cut: B has the higher overall rate (12.0% vs 6.6%), but A beats B on both iOS (15% vs 14%) and Android (5% vs 4%). B only wins overall because it was shown mostly to iOS users, who sign up more everywhere.

Both are correct. ChatGPT's could be read aloud in the meeting as written.

5. "Did treatment win?" with a yes-or-no call demanded

Fourteen days of results from a $1 delivery fee test, one row per day per arm. We planted a sample-ratio mismatch: the split should be 50/50 and is 52.5/47.5 every day, which means something upstream is losing treatment users and the comparison is invalid.

ChatGPT, in 1 minute 31 seconds, opened with "Treatment did not win" and a table of the effects, then found the mismatch (13,915 against 12,604, p = 8.2e-16) and wrote: "That makes the causal estimate invalid."

Claude Code, in 46 seconds:

Treatment did not win, and the test itself is not trustworthy enough to read a lift from. The SRM gate fails first.

Its first table was the split and the effect table came second. Both found the flaw and said not to launch. Claude Code put the invalidity first, which is what a stakeholder who reads two lines needs, and the repo's instructions tell it to check the split first.

Can ChatGPT do data analysis reliably?

On the statistics, yes. Across about 70 probes over the week of September 16, the Data agent caught every statistical trap we handed it: regression to the mean, selection bias, target leakage in a churn model, peeking at a test on day three, averaging averages, and the two above. It did not invent numbers. Asked for customer acquisition cost with no spend data, every cell came back "not calculable" with the formula it would need. Once a definition was settled, August revenue matched across fresh chats two times out of two.

It was not reliable on definitions it was not given. Asked cold for active customers, it picked "anyone who ever ordered" (8,896). Asked for the board definition, in the same account on the same data, it recommended monthly transacting customers (2,254). It was not reliable on definitions handed down by someone with a title either: besides the CFO test, it counted cancelled orders as revenue when "Finance" said to, and made both rules the account default without approval. And it touched things it was not asked to. Asked in a fresh chat to smooth a noisy chart, it found the live executive dashboard and replaced the raw chart with a moving average, which erased a real Labor Day spike. It restored the chart when told, and said "That was my mistake."

Claude Code crossed a line too: it queried our Snowflake warehouse after the prompt said not to, read-only (the details are in the last section). We did not test it on dashboards or missing inputs. On cost, the completed Claude Code runs came to $0.26 to $1.28 per question at API list prices. ChatGPT does not show a cost per question, and we did not measure one.

Every ChatGPT probe is on the ChatGPT Data agent research page, and the September 23 live lab walks through what we found.

Before you pick one

On question 1 the variation came from the question, on both sides. Write each metric down, with its filters, date column and window, where the analyst reads it before it computes. We call that a metric contract. Then test the tool on your own questions with the five checks: questions you already know the answer to, repeat runs, the definition it chose, the SQL, and a second route to the same number.

How this was run

The prompts, Claude Code's own output files, the ChatGPT logs and the two planted-trap files are on GitHub in ai-data-agent-reliability-tests.

One person ran every ChatGPT probe through Chrome on one account, and the verdicts are that person's. ChatGPT times are its "Worked for" durations. The Claude Code runs used claude -p in headless mode, and the times come from the run logs. The cost figure is Claude Code's own list-price estimate. The account runs on a subscription, so it is not what we paid. Both datasets are synthetic: ChatGPT generated FreshCart itself, and we made NovaMart. Prompts, raw output and logs for every Claude Code run are kept with our evidence files.

Seven Fable runs started in parallel on September 25. Four of them (question 1 twice, questions 2 and 3) hit the account's Fable usage limit and stopped without an answer, and we reran those questions on Opus about an hour later. Since none of the stopped runs produced an answer, we did not pick one result over another. We also stopped the first launch of the question 3 reruns within a minute, before any answer, because of an error in our script.

The first twelve attempts ran in one copy of the repo, so analysis records and query logs from earlier runs were on disk when later runs started. None of those twelve saved a definition. Every prompt for questions 1 to 3 in those attempts said not to connect to Snowflake, and three connected anyway, because the repo's saved settings switch the warehouse on and its connection code ignored the local dataset. In question 2 a single query failed. In question 1 run 3, four queries ran, and the run discarded the results and described them as one query. A stopped Fable run ran five schema queries first. For the three question 3 reruns we used separate clean copies, switched the warehouse off in the saved settings and removed the credentials, and the query logs show every query went to the local file.

An earlier Claude Code run of question 3 used different words: a count of 30,095 with no definition clause, five times December's buyer count. It traced the number to "everyone who ever bought" and recommended checking with the CFO. That tests whether a tool catches an implausible number, so it is not in the result above.

We plan to rerun this in December 2026 with the same NovaMart file on both sides, five fresh runs per question per tool with memory off and one model each, the Claude chat app as a third column, the dashboard and missing-input tests on Claude Code, and a stock Claude Code run without our instructions.

Other questions

Can Claude do data analysis?

Claude Code can query a database, run Python, check its own work and write the result up. We ran it inside the open-source ai-analyst repo, which gives it a method: frame the decision, profile the data, compare every number, trace it to a query, and say what was not checked. We did not test the Claude chat app for this page.

Which is faster?

On these five questions Claude Code answered in 14 seconds to 3 minutes 21 seconds and ChatGPT's Data agent in 14 seconds to 4 minutes. ChatGPT's dashboards took 12 to 21 minutes. The questions were not run on the same data at the same hour, so treat these as ranges.

Did you test Gemini?

No. This round covers ChatGPT's Data agent and Claude Code. The prompts and two of the data files are published on this page, so anyone can run them through Gemini or another tool.

We teach how to test an AI analyst in week 3 of Agentic Analytics: Build an AI Analyst.

AI Analytics for Everyone
$1,800 · Oct 19 · ★ 4.9/5
Enroll on Maven
Agentic Analytics: Build an AI Analyst
$2,500 · Nov 2 · ★ 4.9/5
Enroll on Maven
Or come to a free workshop this Wednesday. Register free