OpenAI described its in-house data agent on January 29, 2026. It works across more than 600 petabytes and 70,000 datasets for more than 3,500 internal users, and it is not offered to anyone outside OpenAI. The AI Analyst repo is a version one person can run on a laptop. It is MIT licensed, it is what we run, and our course builds on it.
1. Install the repo and open Claude Code
git clone https://github.com/ai-analyst-lab/ai-analyst.git
cd ai-analyst
pip install -e ".[dev]"
claude You need Python 3.10 or later and Claude Code with a Claude subscription. Node.js 18 or later only matters if you want slides as PDF or HTML. The repo comes with small example datasets and synthetic context for NovaMart, the online retailer we teach with. There is no company data in it, no credentials and none of the answer keys used to grade the course.
What is in the repo, with a recorded run, is on the AI Analyst repo page.
2. Run /setup and connect a data source
Inside Claude Code, type /setup. It asks who you are, what you work on and what data to connect. Then it connects a CSV folder, DuckDB, MotherDuck, Postgres, BigQuery or Snowflake and writes the first files in .knowledge/, the folder the analyst reads before every analysis. Connect a copy or a practice database first. What each source needs, and where the secrets go, is in the Claude Code guide.
3. Read CLAUDE.md, then change it
Claude Code reads CLAUDE.md at the start of every session in that folder. The analyst's method lives there as ten numbered rules. Frame the decision before touching data. Profile before trusting: row counts, date range, nulls, duplicate keys. Every number gets a comparison. Say what was not checked. Log corrections so a mistake does not repeat.
Half the rules name a skill, a SKILL.md file under .claude/skills/ that fires when it applies. There are 64 skill folders. When the analyst makes a chart, the visualization skill applies, and you rarely call one by name. To change how the analyst thinks, edit CLAUDE.md. To add a standard, add a skill file. docs/SKILLS.md lists them all.
4. Write down what your metrics mean
This step decides whether two runs of the same question agree. We asked three models the same retention question five times each. Without a written definition, every model chose more than one definition across its runs, and one chose a different definition on every run. With the contract in the metric dictionary and an instruction to use it, every run on every model chose the same definition.
A metric contract is one metric's definition in writing: the name, the grain (one row is a user, an order, a membership), the numerator, the denominator, the filters, the time window, the source table and an owner. /metric-spec writes one and /metrics lists them. Quirks you know about the data, like the column that changed meaning in March, go in the dataset's quirks.md.
A definition says what to count. It does not say which join to count it along. If the analyst sums an order-level column after joining to line items, every order total gets counted once per line. So the semantic layer needs the joins too. The repo keeps join patterns in .knowledge/query-archaeology/, each with its cardinality next to SQL a person has checked, and the analyst reads that folder before it writes new SQL.
5. Ask a real question
/analyst Using the orders table, how did revenue trend last quarter
versus the one before, and where did we lose the most?
Save me a one-page brief with one chart. It frames the decision, profiles the data, runs the comparison, flags anything odd, and writes the brief and chart into outputs/. A quick fact ("how many signed up in March?") gets a query, an answer, a source and a comparison. A why question gets hypotheses and a written investigation. Big jobs go through /run-pipeline, which hands the work to the 40 pipeline agents listed in agents/INDEX.md, each in a fresh Claude Code process.
We ran a first question live on August 26, 2026, in the Cowork 101 workshop, with the plugin version of this analyst on NovaMart and no context layer yet. Category revenue for Q3 against Q4 2024 came back up 41 percent overall, and the analyst flagged two of its own guesses. Shane read them to the room: "So, one is the revenue definition could be incorrect, so we don't have any sort of context layer built into this yet." The other was that it "is currently excluding returns and cancellations, though we may not want to do that." A metric contract settles both before the next run.
6. Check the answer before anyone else sees it
A wrong number arrives with the same confidence as a right one. /trace-analysis follows a claim back to the query and the rows, using a hook that logs every tool action to a local file. /reliability runs the same task in five fresh trials and reports how often they agree. Its own skill file says: "A wrong analysis can repeat perfectly." /triangulation gets the same number by a different method, and /score-analysis ends with act, investigate, abstain or incomplete.
Then do what you would do without AI. Find where the number the decision leans on came from. Make the parts add up to the total, because a join that fanned out shows up there first. Tie it to a number you already know, like last quarter's figure from a deck someone already presented.
In the Pressure-Test Any AI Analysis workshop on June 24, 2026, Shane pasted a revenue-by-category query the analyst had written and asked whether the total could be trusted. The analyst checked for join fan-out first. Orders has 47,000 rows, order items has 75,000, and total_amount is an order-level column being summed across line items.
"So, when we do that join. our total comes out as $6.9 million, when actually it should be something more like $3.6 million. So the query overstates the revenue by 87%. That's, like, a pretty easy… Mistake to make. I've seen, like, human analysts make that mistake. I've made that mistake myself."
It found two more problems in the same pass. Cancelled and returned orders were in the total, "overstating by another 15%." And the order total includes what the customer paid for delivery, so it proposed the line-level total instead. The rest of the method is on how to check an AI data analyst's answer.
7. Build the answer key and keep the score
A golden set is a list of questions with verified answers. The repo keeps them in evals/cases/public/. /eval runs a case, records the setup, and locks the output before it gets graded. Change the system, run the case again, and compare the two runs.
When you correct the analyst, /log-correction saves the fix with the SQL before and after, and the analyst reads those corrections before its next query. A passing reliability run only says the answer repeats. Whether it is right depends on the reference answers, so a person has to review each one.
If you would rather not run a repo, ai-analyst-plugin is the same method packaged for Cowork and Claude Code, without the pipeline or the eval harness. The Claude Code guide compares the browser, Cowork and Claude Code, and what is an AI analyst covers the term.