Ten Ways an AI Analyst Breaks
The failures we keep hitting while building AI analysts for real companies. Not the demo failures, the ones that show up in week two.
These are the ten primary challenges you’ll encounter building an AI analyst for your company. None of them show up in the demo. All of them show up once the system meets real data, real permissions, and a second person.
-
Analytical workflows fail to persist from one session to the next, so the same request follows different methods and checks each time.
-
Analytical agents duplicate work or lose important findings during handoffs, and the final answer comes back incomplete or inconsistent.
-
Workflows that perform well on sample data break against real permissions, changing schemas, and messy company data.
-
The same analytical question produces different results across runs, making the output too unstable to trust.
-
The system produces the same wrong result on every run, making a bad analysis look reliable.
-
Queries run successfully while flawed joins, filters, or aggregations produce the wrong result.
-
Company metrics get interpreted with generic assumptions, which changes what the analysis actually measures.
-
Evals misgrade analytical work because the evaluator does not match the standards a human reviewer would apply.
-
Fixing one failure causes analyses that used to work to start failing.
-
Changing the underlying model changes the result, and you are left unsure which answer to trust.
Numbers four and five are the pair to watch. A system that disagrees with itself is obviously untrustworthy. A system that agrees with itself and is wrong looks like it works.
10+ years in product data science, causal inference and AI evaluation at Stripe, Nextdoor and Ontra.