Ask a question once and you get a number. Ask it five times, in fresh sessions, and you get five numbers. The distance between the lowest and the highest is the spread. If the five agree, the system is reliable on that question, whether or not it is right. If they scatter, something in the question is undefined. The model is choosing a definition each run.
Reliability is separate from accuracy. A system can be reliably wrong: the same wrong number every run, from one wrong definition in its context. Ten ways an AI analyst breaks lists that one. A system can also be accurate on average and useless in practice. A metric that moves between Monday and Tuesday cannot be acted on.
We ran that test on NovaMart, our course dataset, on a live warehouse. With no definition in front of the model, different runs gave different answers to "What is our retention rate?" With one metric contract in front of it, every run gave the same number. The only change was a written definition.
That last part is the useful lesson. A wide spread is a map of where your context is missing. Each distinct answer names a definition someone could have meant. Pick the one the business means, write it down, and the spread closes. The model stays the same.
Run this on your own data before you trust anything. The workshop below on building context walks through the test live. Its subtitle is the result: ask an AI the same data question five times and you can get five answers.
In the courses
In Agentic Analytics, week 2 (Design the system, connect MCPs and a warehouse), Friday connects the analyst to a warehouse. Then you watch it disagree with itself on an undefined question. Week 3 turns that into a measurement, and week 4 closes the spread with metric contracts. AI Analytics for Everyone, week 2 (Set up your AI analytical toolkit) has you investigate one question three ways. That is a version of the same test.