Without one, an AI analyst asked for "active customers" or "retention" picks a definition on its own, and a fresh session can pick a different one. With one, it computes the reading your team agreed on and says so.
The term is ours. Vendors call the same idea a metric definition, a governed metric or a measure, usually inside their own product. Ours is a short file in plain words that any agent can read before it computes anything, and it sits beside a semantic layer if you have one. In Agentic Analytics: Build an AI Analyst, week 4 has you write one.
What happens without one
On Sep 18, 2026 we asked four fresh ChatGPT chats, memory off, the same question in the same words about the same grocery delivery database: how many active customers do we have? They returned 99, 99, 2,268 and 4,633. Two counted customers active on one day. One counted anyone who ordered in the last 30 days. One counted anyone with a delivered order in the last 90 days. No chat asked which one we meant, and every query was valid.
The template
Ten fields. The first seven are the definition itself. The last three make it hold up over time.
| Field | What goes in it | Why it is there |
|---|---|---|
| Name | The metric id the agent looks up, and the words people use for it | The agent matches the question to the entry by these words |
| Means | One plain sentence a stakeholder would agree with | The reader can disagree with the meaning without reading SQL |
| Grain | One row per what: a customer, an order, a membership, a session | Customers who ordered and orders placed are two different counts |
| Numerator and denominator | What is counted on top and what it is divided by, or the sum for a total | The formula, in words |
| Window | The time period and how it moves: trailing 30 days, calendar month, day of first order plus 90 | The same count over two windows is two metrics |
| Filters | What is in and what is out: statuses, test accounts, refunds, regions | Where cancelled orders get counted as revenue |
| Source | The tables it is computed from, by name | The SQL itself lives in the semantic layer or a verified query, because tables change |
| Rejected alternatives | The readings you considered and did not choose, with the reason | The fork an agent would otherwise take on its own |
| Owner | The person or team who signs off on changes | Changes go through review, the way code does |
| Check value | Optional. One known value on a known date | The agent can test itself against it; update or remove it when the data moves |
As a file, copy this and fill it in, or download it. YAML is what the public AI analyst repo uses. Markdown with the same headings works just as well, because the agent reads it as text.
metric: name_in_snake_case
aka: ["the words", "people use"]
means: >
One sentence a stakeholder
would agree with.
grain: one row per ...
numerator: ...
# denominator: "none" for a count or sum
denominator: ...
window: ...
filters:
include: [...]
exclude: [...]
# source: table names only; the SQL
# lives elsewhere
source: [table_a, table_b]
rejected:
- reading: ...
why_not: ...
owner: ...
# check_value is optional
check_value:
value: ...
as_of: YYYY-MM-DD A filled example: active customers
This is the contract that would have settled the four chats, written for the same synthetic grocery dataset, FreshCart. Its rejected alternatives are three of the readings the chats took on their own. The check value is the count this definition gave on Sep 16, 2026, the last full day in the file, when two chats in a later run happened to choose it. The owner line is illustrative; FreshCart has no staff.
metric: active_customers
aka:
- active customers
- actives
- active users
means: >
Customers who received at least one
order in the last 30 days.
grain: one row per customer
numerator: >
distinct customers with a delivered
order in the window
denominator: none (a count)
window: >
trailing 30 days, ending on the last
full day of data
filters:
include: [status = delivered]
exclude: [cancelled, refunded]
source: [customers, orders]
rejected:
- reading: active on a single day
why_not: >
a daily count swings with the
day of the week
- reading: >
any order in the trailing 90 days
why_not: >
too slow to show churn inside
a quarter
- reading: >
any order in the window, cancelled
or refunded included
why_not: >
counts people who got nothing
owner: Growth analytics lead
check_value:
value: 2194
as_of: 2026-09-16 How to write one
- Pick the metric people argue about. Start with the one that causes the most back-and-forth in meetings, or the one the agent gave different answers to. One contract is enough to start.
- List the readings. Ask the agent to list every defensible definition of the metric on your data, with the number each one gives. This list becomes your rejected alternatives.
- Choose one with its owner. Pick the reading with the person who owns the decision the metric feeds. Write the meaning in one sentence, then the grain, numerator, denominator, window and filters.
- Write down what you rejected. For each reading you did not choose, one line on why.
- Put it where the agent reads it. Save it in the file the agent loads at the start of every session, for example a metrics index in a .knowledge folder. A definition in a wiki page the agent never loads does nothing.
- Test it. Ask the same question five times in fresh sessions before and after the contract goes in. After, every run should return the same reading and cite the entry.
What changed when the agent had the contract
We asked three engines, two Claude models and one open model, the same retention question five times each, on NovaMart, our synthetic course dataset, with a fresh process for every run. Without a written definition, every engine chose more than one definition across its runs, and one chose a different definition on every run. With the contract in the metric dictionary and an instruction to use it, every run on every engine chose the same definition.
The contract said retention means memberships where is_current is true, divided by all memberships ever created, one row per membership, as a snapshot. One run wrote this query from it, copied from the run's log with only line breaks added:
SELECT
COUNT(*) AS total_memberships,
SUM(CASE WHEN is_current
THEN 1 ELSE 0 END)
AS active_memberships,
ROUND(100.0 * SUM(CASE WHEN is_current
THEN 1 ELSE 0 END) / COUNT(*), 2)
AS retention_pct
FROM memberships A contract makes the answer consistent with what you wrote. Shane put the limit this way in a Jun 17, 2026 workshop: "Reliability just told me the system is consistent with one of those definitions. But it didn't necessarily say that, like, that was the right definition." Whether you wrote the right thing is a conversation with the people who use the number. And a wrong join still produces a wrong total under a perfect contract, which is why it comes with checks on the answer.