On Sep 18, 2026 we asked four fresh ChatGPT chats, memory off, how many active customers a grocery delivery database had, and got 99, 99, 2,268 and 4,633. The two chats that returned 99 had found a view in the file that already defined active customers per day and used it. The other two built their own definitions. That view was context, whether anyone meant it that way or not (the full test).
Anthropic's engineering team describes good context engineering as "finding the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome." For an agent that queries a warehouse, the useful context is the facts the tables cannot tell it: which of four reasonable readings of "retention" your company uses, that a purchase flag is wrong for two months of last year, that summing an order total after joining to line items counts each order several times.
The layers of context a data agent reads
The public AI analyst repo, an MIT-licensed analyst that runs inside Claude Code, keeps seven layers. Each is a file under .knowledge/, and how to build an AI data analyst walks through setting them up.
- Instructions
- The rules the agent follows on every question, like reading the other layers before it writes SQL. Without them it skips the rest.
- Schema notes
- Tables, row counts, date ranges and a fingerprint of each table's column names and types. They stop queries against a column that no longer exists.
- Data quirks
- Facts a person learned the hard way, like a flag that is wrong for two months. They stop a valid query on a column that lies.
- Metric contracts
- What each contested metric means: grain, numerator, denominator, window, filters and owner. They stop the agent choosing a definition on its own.
- Verified joins and queries
- SQL a person has checked, and for each join whether one row matches one row or many. Our course's NovaMart store records one that went wrong: summing order totals after joining orders to order_items gave 5.9 million against 3.15 million correct, because each order counted once per line item.
- Corrections log
- Each mistake a person caught, the fix, and the check that should have caught it. It stops the same mistake in next week's analysis.
- Organization glossary
- Business terms and their aliases. It stops the agent reading "activation" as a generic word.
Some of these are measured below: LinkedIn's study covers kinds of context close to the schema notes, verified queries and glossary, and our own runs cover definitions and one quirk. For the instructions and the corrections log, the effect listed is what we expect; we have not measured it.
Which layers change the answer
The clearest public measurement comes from LinkedIn. Their text-to-SQL team scored their system on 133 questions as they removed each kind of context in turn, counting the share of answers rated 4 or 5 out of 5 (Text-to-SQL for Enterprise Data Analytics, Table 2). Read from the top, each row adds one kind of context.
| Context given to the system | Rated 4 or 5 |
|---|---|
| Table schemas only | 9% |
| Add groupings of tables that go together | 11% |
| Add table and column descriptions | 24% |
| Add example queries | 49% |
| Add domain knowledge and jargon | 42% |
| Add usage popularity and join information (full system) | 48% |
The schema alone did almost nothing. Example queries moved the score the most. The one layer of loose prose, domain knowledge and jargon, lowered it. That matches what we have seen in our own builds: Notion pages and meeting notes work better as raw material for written definitions than as context the agent reads directly.
The definitions layer is where we have run the cleanest test ourselves. We asked three engines the same retention question five times each. Without a written definition, every engine chose more than one definition across its runs, and one chose a different definition on every run. With the contract in the metric dictionary and an instruction to use it, every run on every engine chose the same definition. The contract itself, with a template, is on the metric contract page.
A correct note that made the score drop
On Jun 26, 2026 we ran the same 14 questions with reference answers against NovaMart, the synthetic retailer we teach with, on a live Snowflake warehouse. With table names only, 12 of 14 passed. One of the passing questions was "What share of sessions resulted in a purchase?" The analyst counted sessions where the had_purchase flag was true and got 3.33 percent, which matched the reference.
The next run added our team's semantic layer. Its entry for the sessions table carried this caveat, written by a person who had profiled the data:
had_purchase is UNRELIABLE for Nov-Dec 2024:
1,089 sessions have a purchase_complete event AND a
completed order but had_purchase=false (concentrated
around Black Friday). Derive purchase outcome from
events.event_type='purchase_complete' or by joining
orders on session_id; never trust had_purchase for
Q4 2024. The analyst did what it said. Its queries before and after the note, from the run logs with line breaks added:
SELECT ROUND(100.0 * COUNT_IF(HAD_PURCHASE)
/ COUNT(*), 2) AS purchase_session_share_pct
FROM BOOTCAMP_DB.NOVAMART.sessions SELECT COUNT(DISTINCT CASE
WHEN event_type='purchase_complete'
THEN session_id END)
/ COUNT(DISTINCT session_id)
AS purchase_session_share
FROM events The new answer was 3.41 percent. The reference answer had been written on the had_purchase flag, so the case failed and the suite dropped to 11 of 14. On the score alone, the change made the agent worse.
The gap between 3.41 and 3.33 percent, across NovaMart's 1,383,467 sessions, comes to about 1,089 sessions. That matches the count of mislabeled sessions the note names. The note was right. The answer key was the stale piece, computed from the flag the note warns against.
A score only says a case flipped. Someone has to open the case and decide whether the context or the reference is wrong. Reverting the note here would have put a known bug back into every answer about purchases. The fix belongs in the answer key, as its own reviewed change.
How to test a context change before you keep it
The repo's /improve-context skill writes this loop down. Its first line: "Use the evaluation system as the referee. Do not promote a context change because the new answer looks more plausible."
- Name the failure. Read the failing case and what the agent was given. Decide whether information was missing, wrong, out of date or ignored, or whether the problem belongs outside context.
- Write the smallest change as a proposal. One change, with the cases you expect it to fix, the cases it could break, its owner and what counts as success. Nothing goes into trusted context until the proposal is approved.
- Run the baseline and freeze it. Run the target cases and a set of cases that pass today. Freeze the case list, the data snapshot, the model, the grader and the number of trials.
- Apply one change. Make only the approved change and keep the diff. Do not change the grader or the expected answers in the same comparison.
- Rerun and read the flips. Run the same cases again and read each case that changed verdict before looking at the total. Rerun each flipped case, because a small suite moves by a case or two on noise alone.
- Accept, revise or revert. Accept when the target cases improve and nothing that passed regresses. Revert when a passing case regresses for a reason you do not understand. Revise when the runs were not comparable or the target did not improve.
The skill drafts the change and the evidence, and a person decides whether to keep it. The loop needs a set of questions with known answers; the golden set page covers building one.
What goes stale, and how to keep it fresh
Our own team store had a measure named gross_revenue that summed total_amount, which is net of discount and includes the delivery charge. It was marked verified, because verification checked that the SQL ran and the number reproduced. No one had checked the name against what a finance reader means by gross revenue. A pull request merged on Jul 8, 2026 renamed the measure to total revenue and added a real gross revenue on the pre-discount subtotal.
The repo lets you put an owner and a last-reviewed date on each definition and verified query, and it flags an entry stale once its last review is more than 90 days old. A schema check stops the agent from using a definition whose table changed shape. Neither catches a definition that was wrong the day it was written. A second reader who knows the business does.
Start by asking about one metric several times
Shane closed the Aug 5, 2026 workshop with a test anyone can run this week: "take a metric that matters to you… Try to ask it that same question in different sessions, like, 5 to 10 times. And see if any of those answers disagrees." If they disagree, write the definition, put it where the agent reads it, and ask again. The reliability spread page covers reading the answers. After that, write down the quirks your team already knows, and log each correction as you make it.
In the courses
In Agentic Analytics: Build an AI Analyst, week 4 (Context engineering and management) covers where context lives, writing metric contracts and the improvement loop. It comes after week 3 (AI evals for AI analytics), so there is something to measure each change with. See also what agentic analytics is and the free AI evals course.