Underneath it is a language model running in a loop with tools. It decides what to do, calls a tool, reads the result, and decides again until it has an answer. The same thing is sold as an AI data analyst or a data agent, and the way of working is called agentic analytics, where the loop is laid out step by step.
The model matters less than what sits around it: the written definitions it reads before it counts anything, and the checks it runs on its own output.
What the AI does and what you do
Analysis has two halves. One is execution: find the tables, write the query, make the chart, draft the memo. The other is judgment: decide which question is worth asking, define the metric, check whether the number makes sense, and get a decision made. Both of our courses teach the work as eight steps, split like this.
| Step | The AI analyst | The person |
|---|---|---|
| 01 Goal | Clarifies the decision | Pick what matters |
| 02 Question | Proposes and reframes questions | Pick what is worth asking |
| 03 Metric | Calculates it | Define what counts |
| 04 Analysis | Proposes methods, runs the code | Review assumptions, set the quality bar |
| 05 Insight | Surfaces patterns | Judge what it means for the business |
| 06 Decision | Drafts options | Decide what to act on |
| 07 Story | Drafts the readout | Own what it says |
| 08 Close the loop | Watches the result | Learn and set the next goal |
Our September tests of ChatGPT's data agent showed where that line falls. It caught almost every statistical trap we set: a false premise about an August decline, target leakage in a churn model, a planted sample-ratio mismatch in an A/B test. It also adopted an unverified active-customer number because the prompt said the CFO had signed off on it, and counted cancelled orders as revenue when told Finance wanted that. The statistics are execution. Accepting a bad definition is a judgment failure, and people make the same one.
In the Cowork 101 workshop on August 26, 2026, Shane pointed an analyst built from our plugin at NovaMart CSV files and asked which product categories should get more marketing budget next quarter. It wrote a plan before computing anything, then a brief: revenue grew 41 percent from Q3 to Q4 2024, every category grew but not equally, and electronics needed follow-up before any budget change. It also flagged two of its own assumptions, the revenue definition and excluding returns and cancellations, because no one had written them down for it. That brief is where a person's work starts.
How reliable an AI analyst is
On September 18, 2026, we opened four fresh ChatGPT chats with the Data plugin, memory off, attached the same file, and typed the same words: how many active customers do we have right now. The answers were 99, 4,633, 99 and 2,268, on three definitions. No chat asked which definition we meant. The chats had no written definition to read, so each one chose.
A written definition, which we call a metric contract, changes that. We asked three engines the same retention question five times each. Without a written definition, every engine chose more than one definition across its runs, and one chose a different definition on every run. With the contract in the metric dictionary and an instruction to use it, every run on every engine chose the same definition. The causes of the spread are on why AI gives different answers to the same data question, and every probe from the September tests is on the research page.
A consistent answer can still be wrong. In the Pressure-Test workshop on June 24, 2026, a revenue query overstated the total by 87 percent because an order-level column was summed across line items. It would return the same wrong total on every run.
How to check its answer
Ask questions you already know the answer to. Ask the same question several times and look at the spread. Ask which definition it used, and compare that to the one your team uses. Read the SQL, and look for a join that multiplies rows. Get the number a second way and see if the two agree. Worked examples of all five are on how to check an AI data analyst's answer.
The ones we have used
The AI Analyst repo is our open-source analyst, MIT licensed, and it runs inside Claude Code. It has skill files for the standards it follows, 40 registered pipeline agents for the multi-step jobs, Python helpers for the statistics, and an eval harness. The build is on how to build an AI data analyst. We have also run ChatGPT's data agent and the AI Analyst plugin for Cowork.