In the behavioral round of a data interview, an interviewer asks how you use AI in your work. In a data role it is a scored question about what you do when a system hands you a confident, well-formatted answer that is wrong.
Who is asking it
In June 2025, Canva's engineering blog said candidates are now expected to use AI tools during interviews. That policy covers engineering roles and does not mention data roles. For data roles, the evidence is people writing up their loops. In a thread we read in August 2026, a senior product analytics data scientist wrote on r/leetcode that the behavioral round of their final loop, run by data scientists, asked directly how they use AI. The same question had caught them unprepared in an earlier loop.
How much postings ask for depends on the role. We re-read 1,294 US postings in four data roles, collected on August 28, 2026, passage by passage. 23.8% of data analyst postings asked the candidate for modern AI such as LLMs, agents or AI-assisted analytics. For data scientists it was 63.1%. The method and the data file are in AI in data job postings, 2026.
What the interviewer is scoring
An AI analyst's wrong answer looks like a right one. The number is low because the agent filtered on the wrong date column, and the person reading the deck cannot tell. The interviewer wants to know whether you have a repeatable way to catch that, and which decisions you do not hand over. Being fast with AI and having no check is a risk to the team's numbers.
"I use Claude for SQL" is a list of tools. A workflow names the task, where AI came in, where it left, and what you did with the output. "I always double-check the output" describes a feeling. A check is something you do the same way every time: recompute a total two independent ways, check row counts after each join, or keep a small golden set of questions with answers you computed yourself and rerun it whenever the prompt or the model changes. More of them are on how to check an AI data analyst's answer. The decision you kept should be one with no single computed answer, like the primary metric for a launch, so there is nothing for a model to check itself against.
The four-check rubric
Each check is pass or fail. We treat four passes as a strong answer and two or fewer as a vague one.
| Check | Pass | Fail |
|---|---|---|
| 1. A specific workflow | Names a recurring task, the tool, and where AI enters and leaves it. The interviewer could repeat it to a colleague. | Tools and speeds with no task, or a task so general it fits any job. |
| 2. A check that catches a wrong answer | A mechanical check, and you can say what it would catch. | "I review everything." "I sanity check it." "I know the data well enough to spot problems." |
| 3. A decision you kept | One decision you did not hand to AI, with a reason about the decision itself, such as there being no computed answer for it. | No boundary, or a boundary that is only policy ("my company does not allow it"). |
| 4. What changed, and one limit | A concrete change you could check, plus one thing the process caught or one thing it is still bad at. | "It made me three times faster," or a story with no friction in it. |
Check 4 is why a productivity multiplier hurts. The interviewer cannot verify it, and the follow-up ("three times faster at what, measured how?") has no good answer.
Three worked answers
These are composites written for this page to show the shape. They are not transcripts of real candidates. Fill every slot with something that happened to you, because the next question is usually "walk me through that".
A data analyst
Interviewer: How do you use AI in your work?
Mostly on the front half of an analysis. Our weekly marketplace report used to take me most of a day. Now I give the agent our metric definitions file, the schema and the question, and it drafts the queries and a first version of the narrative. I take over where the numbers get interpreted.
I can hand off that much because I built a check first. I keep about fifteen questions whose answers I computed myself against a frozen snapshot of the warehouse: revenue by region for a closed quarter, weekly active users for one specific week, a few funnel steps. When I change the prompt, or the model changes underneath me, I rerun all fifteen and diff them against the frozen answers. That caught a date-column swap once. The agent started filtering on delivery date instead of order date, every number came back low, and nothing looked broken.
Where I do not use it is metric definition. We argued about whether checkout conversion was per session or per user, and the agent will give you a defensible number for either. Which one we adopt depends on what the company is trying to do, and no query settles that. I wrote the definition, took it to the PM and the finance partner, and we froze it.
The limit I have hit is that the drafting needs someone who already knows the rough shape of the answer. On a part of the business I did not know well, I could not tell a good draft from a bad one.
The check is a frozen set of known answers with a named catch, and it ends on a real limit.
A product manager
Interviewer: How do you use AI in your work?
I use it to answer my own data questions before I take them to the analytics team. Before a launch review I ask it for the numbers I would otherwise wait two days for: adoption of the new feature by plan tier, and how the first cohort's usage compares to the old flow.
The first thing I learned is that it will answer a vague question with a precise number. I asked for active users and got a figure that used a different definition from the one on our dashboard, and it never said so. So now I do two things. I ask it to state the definition and show the SQL before I read the number, and I check any number I plan to repeat against a dashboard the team already trusts. If they disagree, I take both to the analyst and we work out which definition we mean.
What I keep is the success criteria. What counts as a good launch, and what result would make us roll it back, gets written before anyone pulls data, by me and the analyst together. If the AI helps pick the target after seeing the numbers, the target will fit the numbers.
It has not replaced the analyst. It has changed what I bring them: a specific disagreement between two numbers, where I used to bring a vague question.
We have measured the definition problem in this answer. Four fresh ChatGPT chats given the same request for active customers returned 99, 4,633, 99 and 2,268, and none asked what active meant. Why that happens is on why AI gives different answers to the same data question.
A data scientist
Interviewer: How do you use AI in your work?
Two ways. Day to day, an agent in Claude Code drafts most of my exploratory SQL and the first pass of experiment readouts. And I built the evaluation for the internal analyst agent our business partners use, so I spend time on how it fails.
For my own work, the check is a trace. Every query the agent runs is logged with a timestamp and a purpose line, and before a number goes in a readout I read the query that produced it: the joins and their grain, the status filters, the date window. For the internal agent, I keep a golden set of questions with answers I verified by hand, and I also run the same question five times in fresh sessions. When the five answers differ, the metric is usually undefined, so the fix goes in the metric definitions the agent reads, and I add the question to the golden set so it cannot drift back.
What I do not delegate is causal claims. The agent will describe a before-and-after difference as the effect of a launch if you let it. Whether the design supports a causal reading, and what the decision rule was before the test started, is my call, written into the experiment brief.
The part it still does badly is saying when the data cannot answer the question. Left alone, it answers with whatever tables exist, so the agent now checks that the data can support a question before it runs the analysis.
It passes because it checks at two levels, the analysis trace for one answer and a golden set plus repeat runs for the agent, and it names a weakness with its fix. A written metric contract is the fix for the drifting definition. Repeat runs and golden sets are the core of the free AI evals course.
Questions that come next
We wrote these from the interview reports above and the checks we teach. No employer survey sits behind them.
- Tell me about a time an AI tool gave you a wrong answer. How did you catch it? Pick a failure that looked correct, and name the mechanism that caught it. "It looked off to me" gets exposed on the first follow-up.
- An agent produced this query and this number. What do you do before it goes to the VP? Restate the question with the metric, grain and window first, confirm the data can answer it, recompute by a path that does not reuse the agent's logic, then say what is still unverified.
- Here is some AI-generated SQL. Find the bug. Read in a fixed order: grain, joins, filters, then the aggregate. The bugs that reach an interview return plausible rows, such as a join that multiplies rows or a date column that is close but wrong.
- You change one line of the agent's prompt. How do you know you did not break something else? A regression set: fixed questions with frozen answers against a pinned data snapshot, rerun on every prompt, model or context change.
- How do you stop the agent answering a question the data cannot support? Check first that a table at the right grain exists and its history covers the window. When it does not, the right output says the warehouse does not record it.
Build your answer from last month's work
Pick a task from the last month where you used an AI tool for any part of it, and write one sentence for each of these.
- Where did the AI start and stop in that task?
- What check did you run that would have caught a wrong answer that looked right? If the answer is none, that is what to build this week.
- What part of the task did you keep, and why can it not be delegated?
- What did the process get wrong, or what is it still bad at?
Read the four sentences aloud in that order. That is your answer. Then have someone ask you to walk them through the regression you caught.
Two free recorded workshops practice the rest of the loop: Crack Product Sense and Analytics Interviews with AI and Reframe Any Question, at Work and in Interviews. More is on the AI data interview.