Definition

Context engineering for data agents

Context engineering is deciding what an AI analyst reads before it works, writing it down, and testing whether each piece changes the answers.

By Shane Butler, Co-founder, AI Analyst Lab. Published Sep 26, 2026, updated Sep 26, 2026.

On Sep 18, 2026 we asked four fresh ChatGPT chats, memory off, how many active customers a grocery delivery database had, and got 99, 99, 2,268 and 4,633. The two chats that returned 99 had found a view in the file that already defined active customers per day and used it. The other two built their own definitions. That view was context, whether anyone meant it that way or not (the full test).

Anthropic's engineering team describes good context engineering as "finding the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome." For an agent that queries a warehouse, the useful context is the facts the tables cannot tell it: which of four reasonable readings of "retention" your company uses, that a purchase flag is wrong for two months of last year, that summing an order total after joining to line items counts each order several times.

The layers of context a data agent reads

The public AI analyst repo, an MIT-licensed analyst that runs inside Claude Code, keeps seven layers. Each is a file under .knowledge/, and how to build an AI data analyst walks through setting them up.

Instructions
The rules the agent follows on every question, like reading the other layers before it writes SQL. Without them it skips the rest.
Schema notes
Tables, row counts, date ranges and a fingerprint of each table's column names and types. They stop queries against a column that no longer exists.
Data quirks
Facts a person learned the hard way, like a flag that is wrong for two months. They stop a valid query on a column that lies.
Metric contracts
What each contested metric means: grain, numerator, denominator, window, filters and owner. They stop the agent choosing a definition on its own.
Verified joins and queries
SQL a person has checked, and for each join whether one row matches one row or many. Our course's NovaMart store records one that went wrong: summing order totals after joining orders to order_items gave 5.9 million against 3.15 million correct, because each order counted once per line item.
Corrections log
Each mistake a person caught, the fix, and the check that should have caught it. It stops the same mistake in next week's analysis.
Organization glossary
Business terms and their aliases. It stops the agent reading "activation" as a generic word.

Some of these are measured below: LinkedIn's study covers kinds of context close to the schema notes, verified queries and glossary, and our own runs cover definitions and one quirk. For the instructions and the corrections log, the effect listed is what we expect; we have not measured it.

Which layers change the answer

The clearest public measurement comes from LinkedIn. Their text-to-SQL team scored their system on 133 questions as they removed each kind of context in turn, counting the share of answers rated 4 or 5 out of 5 (Text-to-SQL for Enterprise Data Analytics, Table 2). Read from the top, each row adds one kind of context.

Context given to the systemRated 4 or 5
Table schemas only9%
Add groupings of tables that go together11%
Add table and column descriptions24%
Add example queries49%
Add domain knowledge and jargon42%
Add usage popularity and join information (full system)48%

The schema alone did almost nothing. Example queries moved the score the most. The one layer of loose prose, domain knowledge and jargon, lowered it. That matches what we have seen in our own builds: Notion pages and meeting notes work better as raw material for written definitions than as context the agent reads directly.

The definitions layer is where we have run the cleanest test ourselves. We asked three engines the same retention question five times each. Without a written definition, every engine chose more than one definition across its runs, and one chose a different definition on every run. With the contract in the metric dictionary and an instruction to use it, every run on every engine chose the same definition. The contract itself, with a template, is on the metric contract page.

A correct note that made the score drop

On Jun 26, 2026 we ran the same 14 questions with reference answers against NovaMart, the synthetic retailer we teach with, on a live Snowflake warehouse. With table names only, 12 of 14 passed. One of the passing questions was "What share of sessions resulted in a purchase?" The analyst counted sessions where the had_purchase flag was true and got 3.33 percent, which matched the reference.

The next run added our team's semantic layer. Its entry for the sessions table carried this caveat, written by a person who had profiled the data:

had_purchase is UNRELIABLE for Nov-Dec 2024:
1,089 sessions have a purchase_complete event AND a
completed order but had_purchase=false (concentrated
around Black Friday). Derive purchase outcome from
events.event_type='purchase_complete' or by joining
orders on session_id; never trust had_purchase for
Q4 2024.

The analyst did what it said. Its queries before and after the note, from the run logs with line breaks added:

SELECT ROUND(100.0 * COUNT_IF(HAD_PURCHASE)
  / COUNT(*), 2) AS purchase_session_share_pct
FROM BOOTCAMP_DB.NOVAMART.sessions
SELECT COUNT(DISTINCT CASE
    WHEN event_type='purchase_complete'
    THEN session_id END)
  / COUNT(DISTINCT session_id)
  AS purchase_session_share
FROM events

The new answer was 3.41 percent. The reference answer had been written on the had_purchase flag, so the case failed and the suite dropped to 11 of 14. On the score alone, the change made the agent worse.

The gap between 3.41 and 3.33 percent, across NovaMart's 1,383,467 sessions, comes to about 1,089 sessions. That matches the count of mislabeled sessions the note names. The note was right. The answer key was the stale piece, computed from the flag the note warns against.

A score only says a case flipped. Someone has to open the case and decide whether the context or the reference is wrong. Reverting the note here would have put a known bug back into every answer about purchases. The fix belongs in the answer key, as its own reviewed change.

How to test a context change before you keep it

The repo's /improve-context skill writes this loop down. Its first line: "Use the evaluation system as the referee. Do not promote a context change because the new answer looks more plausible."

  1. Name the failure. Read the failing case and what the agent was given. Decide whether information was missing, wrong, out of date or ignored, or whether the problem belongs outside context.
  2. Write the smallest change as a proposal. One change, with the cases you expect it to fix, the cases it could break, its owner and what counts as success. Nothing goes into trusted context until the proposal is approved.
  3. Run the baseline and freeze it. Run the target cases and a set of cases that pass today. Freeze the case list, the data snapshot, the model, the grader and the number of trials.
  4. Apply one change. Make only the approved change and keep the diff. Do not change the grader or the expected answers in the same comparison.
  5. Rerun and read the flips. Run the same cases again and read each case that changed verdict before looking at the total. Rerun each flipped case, because a small suite moves by a case or two on noise alone.
  6. Accept, revise or revert. Accept when the target cases improve and nothing that passed regresses. Revert when a passing case regresses for a reason you do not understand. Revise when the runs were not comparable or the target did not improve.

The skill drafts the change and the evidence, and a person decides whether to keep it. The loop needs a set of questions with known answers; the golden set page covers building one.

What goes stale, and how to keep it fresh

Our own team store had a measure named gross_revenue that summed total_amount, which is net of discount and includes the delivery charge. It was marked verified, because verification checked that the SQL ran and the number reproduced. No one had checked the name against what a finance reader means by gross revenue. A pull request merged on Jul 8, 2026 renamed the measure to total revenue and added a real gross revenue on the pre-discount subtotal.

The repo lets you put an owner and a last-reviewed date on each definition and verified query, and it flags an entry stale once its last review is more than 90 days old. A schema check stops the agent from using a definition whose table changed shape. Neither catches a definition that was wrong the day it was written. A second reader who knows the business does.

Start by asking about one metric several times

Shane closed the Aug 5, 2026 workshop with a test anyone can run this week: "take a metric that matters to you… Try to ask it that same question in different sessions, like, 5 to 10 times. And see if any of those answers disagrees." If they disagree, write the definition, put it where the agent reads it, and ask again. The reliability spread page covers reading the answers. After that, write down the quirks your team already knows, and log each correction as you make it.

In the courses

In Agentic Analytics: Build an AI Analyst, week 4 (Context engineering and management) covers where context lives, writing metric contracts and the improvement loop. It comes after week 3 (AI evals for AI analytics), so there is something to measure each change with. See also what agentic analytics is and the free AI evals course.

Questions people ask about context engineering

What is context engineering for data agents?

Deciding what an AI data agent reads before it writes a query, writing those things down in files it loads, and testing whether each piece changes its answers. For a data agent the pieces are concrete: schema notes, data quirks, metric definitions, verified joins and queries, a log of past corrections and the business glossary.

How is it different from prompt engineering?

A prompt is the question you type. Context is everything else the agent reads before it answers. Context files persist between sessions, so a fix to them improves every later question about the same data.

Is a semantic layer the same thing?

A semantic layer is one layer of the context: the map from business words to tables, joins and measures. Context engineering covers that layer plus the ones a semantic layer usually leaves out, such as data quirks, the corrections log and plain-language metric contracts, and the testing that decides what stays. Semantic layer

Who should own the context?

Metric meanings come from the people who own the decisions the metrics feed. The data team usually owns the SQL, the joins and the verified queries. Whoever owns each piece, changes go through review with a name and a date on them, the way code does.

Do I need a vendor product for this?

No. A folder of markdown and YAML files in a git repository, loaded by the agent at the start of each session, is enough to start. Warehouse vendors have their own versions of the same layers, such as verified queries in Snowflake and example SQL queries in Databricks. How to build an AI data analyst

Next cohorts start Oct 19 and Nov 2.

AI Analytics for Everyone
$1,800 · Oct 19 · ★ 4.9/5
Enroll on Maven
Agentic Analytics: Build an AI Analyst
$2,500 · Nov 2 · ★ 4.9/5
Enroll on Maven
Or come to a free workshop this Wednesday. Register free