Onboarding a New Analyst Is Building an AI Analyst
How to build a reliable AI analyst is the same list I would use to train a person. Six things, and what each one becomes in the system.
I’m going to stop telling people how to use AI for data analysis. Instead, I’m going to teach them how to onboard and train a new analyst. Because it is basically the exact same thing.
Here is what I do with a new analyst, and what each step turns into when the analyst is a system.
Give them a real assignment
Who is making what decision, why does it even matter, and what answer would change it? A new analyst who starts from a vague request will produce vague work. So will the system. Understanding the decision becomes a question-framing skill: the first thing the analyst does with any request is turn it into a question that could change what someone does.
Teach them why the metrics matter
What does a metric say about users, and how does it drive the product and the business? Then teach the nuance behind it and where it is easy to screw up: what each row represents, its population, the date window, the filters. Metric knowledge becomes a semantic layer. Written down once, read every time, instead of guessed at in each new conversation.
Show them what good work looks like
A library of past analysis with validated answers. Then review their first attempts against those standards, walking through the feedback piece by piece and helping them not only correct each error but understand why it was an error, so it does not happen again. Reviewed examples become eval cases. The library is the answer key, and the review is the grading.
Make them show their work
For every insight, share the data source, the query, any definitions and assumptions, and the caveats that come along with it. Another analyst should be able to follow their trail and produce the same numbers. Showing work becomes an analysis trace.
Make the basic sanity checks automatic
Count rows before and after joins so you neither duplicate nor drop them. Reconcile numbers against other trusted sources. Rerun the same analysis when the underlying data or the business understanding changes. Sanity checks become deterministic validation, the part of the system that never gets tired.
Teach them to push back early
If a question is vague, the data quality is messy, the metric is unclear, or the assumptions behind the analysis are flimsy, I want to hear about it early, before they keep going down a wrong path and draw conclusions. Asking for help becomes a human review gate.
A lot of building practical agentic analysis is just recreating the way good analysts already work. The technical work, and all of the judgment and friction that come along with working through a problem.
What do the best data scientists or analysts you’ve worked with do almost automatically? Start there. That is the list your system needs.
10+ years in product data science, causal inference and AI evaluation at Stripe, Nextdoor and Ontra.