ChatGPT's Data Agent: Senior Technical Execution, Junior Business Judgment
We ran 70 tests on OpenAI's data agent. It caught almost every statistical trap and failed the same organizational errors people do. Where I would use it and where I would not.
OpenAI’s new data agent is really good. It caught almost every statistical trap we threw at it over 70 tests. Then it failed the same errors people fail in real life.
My read so far: senior-level technical execution with junior-level business judgment. It could save data teams from a lot of fake work, or create a lot more of it.
The technical work holds up
As expected, it ran queries, explored data, built models and experiments, made dashboards. It also handled a bunch of analytical traps I was certain it would screw up.
It challenged false premises and refused to invent answers when data was missing. It caught models cheating with future information and experiments too broken to trust. It recognized misleading comparisons, incomplete periods, tiny samples, and groups too new to measure. It separated patterns from causes, showed its work, and proposed what to test next.
It’s better at analysis than nearly everyone outside data I’ve worked with, and even some folks in data. In grad school, I think it would be top of its class.
It failed the way people fail
But this isn’t school. Doing analysis is only part of our job. The harder part is the messy work inside an organization, where every question is tangled in incentives, hierarchy, history, ego, and human error.
Once we put it in situations data teams experience every day, it failed like humans do. It accepted bad definitions and unverified numbers when a stakeholder wanted them. It produced requested work without asking whether it was useful. It dropped context or buried caveats when asked for a simpler answer. It put numbers behind the story someone already wanted and cleaned up charts to mislead.
Companies already make these mistakes. But good data teams introduce friction. They challenge questions, definitions, and conclusions before someone shares bad work in a meeting and calls it data-driven. In the hands of someone who can’t recognize those failures, AI could produce a mountain of technically correct bullshit that looks credible enough to drive very bad decisions.
It also failed in ways a person would not
It turned unapproved instructions into company knowledge that polluted future work. It carried bad assumptions into new conversations and artifacts without review. It could not reliably tell the difference between permission to investigate and permission to act. It introduced unrelated changes when correcting earlier mistakes.
It did not hallucinate numbers. We are far beyond that. But a correct calculation based on a bad definition is still wrong, and it can still lead to a bad decision.
Most of what data teams get asked for is fake work
Data teams constantly get requests I consider fake work. A leader needs a number for a meeting that won’t get acted on. Another team wants something for a deck so they look like they know what’s going on. An endless list of someone who is curious.
The questions are typically vague, and the person asking often doesn’t really know why they’re asking, especially if it’s on behalf of someone else.
I ask two questions before agreeing to that work. Does it measure what actually matters? Will anyone actually use it to make a decision? Ideally, only the work that answers yes to both makes it to the data team. In reality it is very hard to stop. Sometimes it comes from an important person the management chain is too scared to say no to, even when they know it’s a waste of time. Sometimes it comes from too many directions at once, and filtering the fake stuff out becomes a job of its own. Sometimes we’re just too tired to argue.
This is where agentic analytics could save data teams a lot of time and headache. Give the people asking the tools to answer the fake work themselves, instead of pulling the data team away from work that matters. That comes with its own risk. Harmless curiosity can turn into unreliable or misleading conclusions in the wrong hands, and the tests above show how.
Where I would use it
If nobody’s making an important decision, let people explore the data for themselves. It is a great way to learn, get answers, and be more data-driven.
If the metric and the decision are already well defined, let the data agent handle the technical work, but have someone from the data team review it.
If several definitions could be reasonable and there is no alignment at the company yet, use the agent to calculate each of them and share those as artifacts to align the team.
Where I would not
If the metric does not measure anything that matters, just stop. Data theater props are still props, and they still risk misleading conclusions.
If people will act on the answer but the metric definition or the methodology is unclear or disputed, involve the data team before the agent starts.
Don’t let it publish important numbers, change shared company knowledge, or take action without review. It can cause cascading failures in other analysis downstream.
It’s a good tool. I really like it. It just has the same faults humans do. I’ll be using it, and you should too. Be responsible and don’t get lazy.
AI does not just scale good analysis. It scales your organization’s unresolved definitions, bad incentives, and poor judgment too.
Every test, what we asked, what the agent did and how we counted is on the research page for these tests.
10+ years in product data science, causal inference and AI evaluation at Stripe, Nextdoor and Ontra.