One recorded run
This is from the free session Root Cause Analysis in Claude Code, recorded May 20, 2026. Hai Guan runs it on NovaMart, the synthetic online store we teach with.
AI Analyst is a data analyst that runs in Claude Code on your laptop: you connect a data source, ask a question in plain English, and it writes the SQL, checks the data and saves a brief and a chart. It is MIT licensed and contains no company data, so the first step is connecting your own with /setup.
By Shane Butler · Sep 26, 2026. Checked against the repo at commit aeaf850, Sep 26, 2026.
This is from the free session Root Cause Analysis in Claude Code, recorded May 20, 2026. Hai Guan runs it on NovaMart, the synthetic online store we teach with.
Hai asks why payment support tickets spiked in June 2024. The analyst confirms the spike, drills down until it names the root cause, a "defective iOS app release, version 2.3.0", and writes a Google Doc report with the chart, the evidence, a confidence rating and the queries behind it.
Recorded on version 2 of the repo. Version 3, released Aug 27, 2026, rebuilt the tree, so a fresh clone looks different on screen.
CLAUDE.md is the method. Claude Code reads it at the start of every session, and it holds ten rules, from "Frame the decision before touching data" to "Never expose credentials."
Skills are 64 markdown files that apply when they are relevant, so you rarely call one by name. docs/SKILLS.md sorts them into eight kinds:
| Kind | Skills | For example |
|---|---|---|
| Method | 7 | frame the question, compare every number, pair a win with its guardrail metrics |
| Data | 11 | connect a source, profile tables, check data quality |
| Analysis | 8 | plan an analysis, forecast, run the full pipeline |
| Experiments and causal | 4 | A/B test design and readout, sample ratio check, causal methods |
| Checks and evals | 13 | reliability, trace, triangulation, score, eval |
| Charts and exports | 8 | chart style, Google Docs and Slides, Notion, Slack |
| Knowledge and memory | 7 | metric definitions, corrections, saved join patterns |
| Setup | 6 | /setup, switch dataset, hand off a long session |
Pipeline agents are 40 step definitions, listed in agents/registry.yaml, that /run-pipeline runs one after another for a full analysis from question to slide deck.
Python helpers are 151 modules that hold the parts that should not vary between runs: experiment statistics, causal methods, forecasting, validation and provenance logging. In the README's words, "The model reasons; the arithmetic runs in code."
The checks are four commands. /reliability runs one question in five fresh trials and reports how often they agree. Its skill file says: "A wrong analysis can repeat perfectly." /trace-analysis follows a claim back to its query and source. /triangulation compares independent analyses that use different methods. /score-analysis ends in act, investigate, abstain or incomplete.
Evals hold 22 complete-analysis cases on NovaMart in evals/cases/public/. /eval locks each run's output before it is graded, so you can change the system, rerun and compare. The course's held-out answers are not in the repo.
git clone https://github.com/ai-analyst-lab/ai-analyst.git
cd ai-analyst
pip install -e ".[dev]"
claude You need Python 3.10 or later and Claude Code with a Claude subscription. Node.js 18 or later only matters for slides as PDF or HTML. Inside Claude Code, type /setup. It asks in four rounds, two or three questions at a time:
Then ask a real question. The README's example:
/analyst Using the orders table, how did revenue trend last quarter
versus the one before, and where did we lose the most?
Save me a one-page brief with one chart. Connecting Snowflake, BigQuery and the rest is on the Claude Code guide.
On September 25 we ran five data questions through the repo in fresh Claude Code sessions (the full results are on ChatGPT vs Claude for data analysis). It caught both statistical traps. It explained a Simpson's paradox in 14 seconds, and on a rigged A/B test its first line said the test was not trustworthy enough to read a lift from, because the split was 52.5/47.5 when it should be 50/50. When a prompt said orders fell, it checked first and answered that they rose 4.9 percent. Handed a CFO's count of active customers and asked for revenue per active customer, all three runs pointed out that the ratio divided revenue without cancelled orders by a count that included them.
It does not know your metric definitions until you write them down. Asked how many active customers we have right now, in four fresh sessions, three of them on the same model, it gave four numbers, from 5,526 to 14,371, each with a different definition. Every run said the project had no definition and listed the alternatives, but a number still came back. /metric-spec writes the definition, and the metric contract page covers what goes in one. In the CFO test, two of the three runs saved the handed-down definition as the official metric, and none asked first.
The same runs turned up a bug in the repo. The prompts said not to connect to Snowflake, and three of the first twelve attempts queried our warehouse anyway, read-only, because the saved settings switched the warehouse on and the connection code ignored the local dataset. The fix was merged into main on September 26: a named dataset is now honored, and an off switch set in your shell now overrides the saved setting.
BigQuery is implemented, with a setup guide, but its tests run against fake clients and we have not run it against a live project in a recorded session. Check the first few numbers against the BigQuery console, and check the README's "Your data" section for current status. For the checks to run on any answer before you share it, see how to check an AI data analyst's answer.
Yes. The repo is MIT licensed. You need Claude Code with a Claude subscription to run it, and Python 3.10 or later.
The Python tools process your data on your machine, but the prompts and the context the analyst reads go to the model provider, and an export leaves when you run one. The two hooks in .claude/settings.json write local log files only. The README covers this under "What runs on your machine".
The plugin, ai-analyst-plugin, packages the same method for Claude Cowork and Claude Code as skills only, without the pipeline or the eval harness. The repo is the full system.