The system doing the work is an AI analyst: a language model running in a loop with tools. It queries a warehouse, runs code, reads the metric definitions your team wrote down, and saves what it makes. The model does the execution. The people around it decide what the metrics mean and whether the answer holds.
You will also see it called an AI data agent or autonomous analytics. Augmented analytics, the term some BI vendors use, is something else: a dashboard with an alert or a chat box added, where the definitions, joins and charts were still built by hand ahead of time.
How it differs from a dashboard, a chat assistant and a person
An agentic analyst starts from the question and goes to the data, the way a human analyst does. The difference is that it has no sense that a number looks off unless you build one in as a check.
| Step | Agentic analyst | Human analyst | Dashboard (BI) | Chat assistant |
|---|---|---|---|---|
| Takes the question | Reads the question and asks what decision it serves | Asks what decision it serves | A person picks a dashboard and a filter | A person pastes a table and asks |
| Picks the definition | Reads the written definition, or guesses if there is none | Asks, or uses the team convention | Fixed when the dashboard was built | Guesses from the pasted columns |
| Writes and runs the query | Profiles the tables, writes SQL, runs it | Writes SQL, runs it | Pre-built | Cannot reach the warehouse |
| Checks the number | Reruns, traces and reconciles, if you built those checks | Sanity checks against numbers they know | Only when a tile looks odd | Not checked |
| Writes the memo | A brief with a chart, the definition used, and a recommendation | Writes it | A person reads the tiles and writes it | A paragraph about the pasted table |
| Publishes | Saves to the project, exports to a doc, deck or Slack | Sends it | Already published | Copy and paste |
How the loop runs
The public AI analyst repo runs this loop when you type /analyst and a question. Other agents name the parts differently. What varies between them is how much of steps 2 and 4 exists.
- Frame the question. The agent asks what decision the answer serves before it touches data. "How is retention doing" becomes "should we spend the Q4 budget on winning back lapsed customers," which changes which metric it computes and over what window. When a question is too vague to pick a definition, the agent should stop and ask.
- Read the context. Before it writes SQL, the agent reads a
.knowledge/folder: schema notes, data quirks, the team's metric contracts, and a log of every correction a person has made. Without this step, four fresh ChatGPT chats asked the same question about the same file came back with three different active-customer counts, because each one chose its own definition. - Profile, then query. The agent checks the tables it is about to use (row counts, date ranges, null rates, duplicate keys), writes the SQL, runs it, and pairs every number with a comparison: the prior period, a benchmark, another segment.
- Check the work. The number gets a trace back to the query and the rows. The agent reruns the question to see if the answer holds and reconciles parts to totals. In the public repo, a hook writes a line to a local log for every action the agent takes, so a person can read the exact query later instead of asking the model to remember it.
- Write and publish. A brief with the chart, the definition it used, the caveats and a recommendation, saved into the project and exported where the team reads. The brief names the definition so the reader can disagree with it.
Steps 1, 3 and 5 are what people mean when they say AI can do analysis now. Steps 2 and 4 are what we teach. Without them, step 3 produces a valid query on a guessed definition and step 5 dresses it up. How often that happens, and what fixes it, is on how reliable an AI analyst is.
One run of the loop, from a recorded workshop
In the May 20, 2026 session, Hai Guan ran the public repo in Claude Code against NovaMart, our synthetic course dataset, and asked it a "why did this change" question. Quoted lines are Hai's narration.
He started by asking what the data was about. The agent read the context in the repo and described a simulated online retailer with about 50,000 customers, 47,000 orders, 500 products and a paid membership "kind of like a $99 Amazon" Prime. Then he asked for ticket volume over time as a chart, and it showed a spike in June.
The question itself: "Can you do a root cause analysis on the payment ticket spike in June?" Before explaining the spike, the agent checked that it was real and sized it: "June payment jumped 41% of all tickets, July fell back down, crucially, attempts grew smoothly." Payment attempts growing smoothly rules out the easy explanation, that more payments simply produced more tickets. It cut the tickets by device, app version and order links, hit errors in its own code and fixed them without being asked, and came back with "Defective iOS app release, version 2.3.0." The spike was confined to iOS, the defect was fixed a few days later, and ticket volume went back to its baseline.
Hai's comment at the end: "Now, assuming this is all cool, like, we still have to check the work." The iOS 2.3.0 defect was planted in the data when we generated it, so this was a question with a known answer, and the agent found it. On your own warehouse nothing is planted. A run that finds the real cause and a run that tells a tidy story about a coincidence look the same on screen. The checks that tell them apart are on how to check an AI data analyst's answer.
Trying it
The public repo is MIT licensed and comes with no company data. Clone it, run /setup, connect a source, and ask a real question. The step-by-step build is on how to build an AI data analyst, and the Claude Code setup is on Claude Code for data analysis. We teach the loop in Agentic Analytics: Build an AI Analyst.