Definition

Metric contract

A metric contract is one metric's written definition, its numerator, denominator, grain, filters and time window, which an AI analyst reads before it computes anything.

By Shane Butler, Co-founder, AI Analyst Lab. Published Sep 26, 2026, updated Sep 26, 2026.

Without one, an AI analyst asked for "active customers" or "retention" picks a definition on its own, and a fresh session can pick a different one. With one, it computes the reading your team agreed on and says so.

The term is ours. Vendors call the same idea a metric definition, a governed metric or a measure, usually inside their own product. Ours is a short file in plain words that any agent can read before it computes anything, and it sits beside a semantic layer if you have one. In Agentic Analytics: Build an AI Analyst, week 4 has you write one.

What happens without one

On Sep 18, 2026 we asked four fresh ChatGPT chats, memory off, the same question in the same words about the same grocery delivery database: how many active customers do we have? They returned 99, 99, 2,268 and 4,633. Two counted customers active on one day. One counted anyone who ordered in the last 30 days. One counted anyone with a delivered order in the last 90 days. No chat asked which one we meant, and every query was valid.

The template

Ten fields. The first seven are the definition itself. The last three make it hold up over time.

FieldWhat goes in itWhy it is there
NameThe metric id the agent looks up, and the words people use for itThe agent matches the question to the entry by these words
MeansOne plain sentence a stakeholder would agree withThe reader can disagree with the meaning without reading SQL
GrainOne row per what: a customer, an order, a membership, a sessionCustomers who ordered and orders placed are two different counts
Numerator and denominatorWhat is counted on top and what it is divided by, or the sum for a totalThe formula, in words
WindowThe time period and how it moves: trailing 30 days, calendar month, day of first order plus 90The same count over two windows is two metrics
FiltersWhat is in and what is out: statuses, test accounts, refunds, regionsWhere cancelled orders get counted as revenue
SourceThe tables it is computed from, by nameThe SQL itself lives in the semantic layer or a verified query, because tables change
Rejected alternativesThe readings you considered and did not choose, with the reasonThe fork an agent would otherwise take on its own
OwnerThe person or team who signs off on changesChanges go through review, the way code does
Check valueOptional. One known value on a known dateThe agent can test itself against it; update or remove it when the data moves

As a file, copy this and fill it in, or download it. YAML is what the public AI analyst repo uses. Markdown with the same headings works just as well, because the agent reads it as text.

metric: name_in_snake_case
aka: ["the words", "people use"]
means: >
  One sentence a stakeholder
  would agree with.
grain: one row per ...
numerator: ...
# denominator: "none" for a count or sum
denominator: ...
window: ...
filters:
  include: [...]
  exclude: [...]
# source: table names only; the SQL
# lives elsewhere
source: [table_a, table_b]
rejected:
  - reading: ...
    why_not: ...
owner: ...
# check_value is optional
check_value:
  value: ...
  as_of: YYYY-MM-DD

A filled example: active customers

This is the contract that would have settled the four chats, written for the same synthetic grocery dataset, FreshCart. Its rejected alternatives are three of the readings the chats took on their own. The check value is the count this definition gave on Sep 16, 2026, the last full day in the file, when two chats in a later run happened to choose it. The owner line is illustrative; FreshCart has no staff.

metric: active_customers
aka:
  - active customers
  - actives
  - active users
means: >
  Customers who received at least one
  order in the last 30 days.
grain: one row per customer
numerator: >
  distinct customers with a delivered
  order in the window
denominator: none (a count)
window: >
  trailing 30 days, ending on the last
  full day of data
filters:
  include: [status = delivered]
  exclude: [cancelled, refunded]
source: [customers, orders]
rejected:
  - reading: active on a single day
    why_not: >
      a daily count swings with the
      day of the week
  - reading: >
      any order in the trailing 90 days
    why_not: >
      too slow to show churn inside
      a quarter
  - reading: >
      any order in the window, cancelled
      or refunded included
    why_not: >
      counts people who got nothing
owner: Growth analytics lead
check_value:
  value: 2194
  as_of: 2026-09-16

How to write one

  1. Pick the metric people argue about. Start with the one that causes the most back-and-forth in meetings, or the one the agent gave different answers to. One contract is enough to start.
  2. List the readings. Ask the agent to list every defensible definition of the metric on your data, with the number each one gives. This list becomes your rejected alternatives.
  3. Choose one with its owner. Pick the reading with the person who owns the decision the metric feeds. Write the meaning in one sentence, then the grain, numerator, denominator, window and filters.
  4. Write down what you rejected. For each reading you did not choose, one line on why.
  5. Put it where the agent reads it. Save it in the file the agent loads at the start of every session, for example a metrics index in a .knowledge folder. A definition in a wiki page the agent never loads does nothing.
  6. Test it. Ask the same question five times in fresh sessions before and after the contract goes in. After, every run should return the same reading and cite the entry.

What changed when the agent had the contract

We asked three engines, two Claude models and one open model, the same retention question five times each, on NovaMart, our synthetic course dataset, with a fresh process for every run. Without a written definition, every engine chose more than one definition across its runs, and one chose a different definition on every run. With the contract in the metric dictionary and an instruction to use it, every run on every engine chose the same definition.

The contract said retention means memberships where is_current is true, divided by all memberships ever created, one row per membership, as a snapshot. One run wrote this query from it, copied from the run's log with only line breaks added:

SELECT
  COUNT(*) AS total_memberships,
  SUM(CASE WHEN is_current
      THEN 1 ELSE 0 END)
    AS active_memberships,
  ROUND(100.0 * SUM(CASE WHEN is_current
      THEN 1 ELSE 0 END) / COUNT(*), 2)
    AS retention_pct
FROM memberships

A contract makes the answer consistent with what you wrote. Shane put the limit this way in a Jun 17, 2026 workshop: "Reliability just told me the system is consistent with one of those definitions. But it didn't necessarily say that, like, that was the right definition." Whether you wrote the right thing is a conversation with the people who use the number. And a wrong join still produces a wrong total under a perfect contract, which is why it comes with checks on the answer.

Questions people ask about metric contracts

Is a metric contract the same as a semantic layer?

No. A semantic layer says how to compute things: tables, joins, measures, and the SQL. A metric contract says what one metric means, in plain words, with no SQL. A semantic layer can hold many contracts beside its measures. If you have neither, a folder of contracts in a file the agent reads is the cheaper place to start.

Is it the same as a data contract?

No. A data contract is an agreement between the team that produces a table and the teams that use it: the schema, the types, freshness and quality checks. A metric contract sits one level up, on the business meaning of a number computed from those tables.

Should a metric contract include SQL?

We leave it out. Tables get renamed and columns get split, so SQL written today goes stale while the meaning usually does not. Keep the SQL in the semantic layer or in a verified query, and let the contract name the tables it draws from.

Where do I store it so an AI analyst uses it?

In a file the agent reads before it writes any SQL, every session. In the public AI analyst repo that is the metrics index inside the .knowledge folder. With a chat assistant, paste it at the top of the conversation or put it in the project instructions.

How do I know the agent used my definition?

Ask the same question several times in fresh sessions and check two things: every run uses the same definition, and every run names the contract as its source. If runs disagree or cite nothing, the agent did not read it. Reliability spread

Next cohorts start Oct 19 and Nov 2.

AI Analytics for Everyone
$1,800 · Oct 19 · ★ 4.9/5
Enroll on Maven
Agentic Analytics: Build an AI Analyst
$2,500 · Nov 2 · ★ 4.9/5
Enroll on Maven
Or come to a free workshop this Wednesday. Register free