← All free workshops
Free workshop · Wednesday, August 5, 2026

Build the Context That Makes AI Data Agents Reliable

Ask an AI the same data question five times and you can get five answers. Here is how to fix that.

Live on Maven, Wednesdays at 10 AM Pacific. About 65 minutes.

Transcript

Auto-transcribed from the live session and lightly cleaned. Attendee names are removed; their questions are kept.

Shane Butler: I didn’t record the last one I did, and then I had to re-record it solo, and kind of sucked. So, today, though, today we’re gonna talk a bit about… Context, now that we’ve given you some context about ourselves. And… and specifically, the context kind of required that makes, AI data you know, Agentic systems or agents, reliable. So I think to kick it off, Actually, maybe to get a pulse on the room before we even get… could you drop, like, a yes or no in the chat? Like, drop a yes if you’re using… AI for… data… Analytics, data science work. Today at all, whether it’s your personal thing, or at your job, or a no, if not yet. Just to kind of get a feel of… you know, how people are using it.

Cool, we got a lot of people who have… Yeah, feel free to drop in, like. how you’re using it, even, or what you’re using, Gemini. Oh, hockey stats and stuff, yeah, fun stuff, that’s cool. Nice. We even got an app on there. I’m gonna check that out later. I’m not much of a sports bike guy, but I do like hockey, just purely because it’s just, like, the only sport where it’s like, alright, you guys can fight for it a little bit right now, like, we’re all just gonna pretend that’s all good. All right, let’s get into this, though.

So… I think one of the really important things with, like, why context is so important is, when we think about how we evaluate the sort of answers and output we get from AI data, systems, however you’re using it, right? There’s two, I think, main categories of how we evaluate whether they’re good or bad, or somewhere in between. And sometimes these words can get a bit conflated. But they mean very different things, so capability and reliability are the two things we pay attention to a lot here at the lab. Reliability is, I think, one of the, really important aspects that is actually, like, quite easy to, measure. You can do it without having, like, a whole, like.

Golden dataset or anything like that, or actually knowing the right answer. But it’s not, I feel like, taken as seriously, but it’s also one of these things where it’s, like, it’s really clear that a system’s unusable if it’s unreliable. So when I talk about capability and reliability, capability… this is kind of like, can this thing do the job at all? Can it read my warehouse? Can I write a query? Can it come back with a number? Or can it come back with the correct number at least once? And then reliability is around… and this, this goes… this is just for, like, data. Agents. This is any sort of, like, AI use case.

I… evaluate across these dimensions when I… build systems for evaluation around, like, legal AI, for instance. Reliability is, if I ask the same question tomorrow. or today, am I going to get the same answer I did the first time? If I ask it 100 times, am I getting the same answer every single time? If someone else asks that question on a completely different machine. are they gonna get the same answer that I do? And, like, a lot of the pattern we see, and we see it in our own courses and cohorts in the early weeks when we start building out agents as well, is, like, you know. We will kind of have, like, a sort of, like, demo-ish moment, where we see that this thing is capable.

It can create… answer a question for us. But then, when we kind of, like. push that trial further, and, like, multiple people ask the same question, they all get different numbers, or they… some… some people get the same numbers, some people get different numbers. And a lot of the times, the conclusion at the end of that, the kind of knee-jerk reaction is like, oh, this thing’s, like, this thing doesn’t work. Right? Like, I don’t know, the model is bad. And it’s not necessarily that’s bad, Every one of those people or those runs of those questions might have some very plausible, defensible reason behind why the agent eventually resulted in that output.

It’s capable, but it just isn’t reliable. And unfortunately, an unreliable, I think, means unusable. And so we cannot have, like, Production level, like. Data agents, or decisions made off of, like, data agents that are consistently unreliable. So we’re gonna kind of look a little bit at, like. what’s actually happening there. We’re not gonna go… I’m not gonna open up Cloud Code today, because we’ve given this, kind of demo a few times in our past sessions where we discussed AI evals for Agentic Analytics.

Yeah, I don’t know, Hai if you wanna… or Sravya, if you wanna dig into our past Lightning lessons, or maybe our YouTube channel, and try and drop links to those, but we did do two past less lessons on, on AI evals for… and validation for Agentic Analytics, where we kind of go into how we measure reliability. So we won’t go into that today, but at a high level, basically, you know, we ask it our system, like, hey, a vaguer, a kind of vague question, like, what is our retention rate, for instance?

that’s a deliberately kind of broad question that somebody, a stakeholder could ask, could drop in Slack, you could hear in a meeting where the term retention isn’t actually, like, clearly defined, though many people in the room could think they know exactly what that definition is. And so… when we asked that, and if you watch, like, this, that past less than that high shared, when we asked it against our warehouse on Snowflake, it’s like, we asked it, like, 30 different times. We asked it across 3 different engines, meaning 3 different models, Opus, Fable, and, open source model, GLM 5.2, we run it with a fresh session for every time we ask that, so nothing from one run to another pollutes.

And what ends up happening is we usually get a bunch of different answers. So, if we had just asked it one time, so, like, the first time we asked it to Fable, it comes up with this answer, it says, like. Yeah, retention is 25.3%. And notice that even… we even have, like, an audit trail that… where it can tell us, like, how… how it was measured. Like, the share of customers whose first completed order was followed by a second completed order within 90 days. That sounds like a pretty good answer.

Like, the SQL, we check the SQL behind this, it’s actually valid, it’s ran against Rio Warehouse, the math is correct for the query it wrote, it didn’t really hallucinate any of that… hallucinate any of that actual, like, analytics work. And so it would be actually, if we just were validating it this way, it’d be really easy for us to kind of run this analysis and then share this number out, with someone else. However, on that same Fable 5 model, we did 5 runs against that same engine, on that same question, and actually We got a, a spread of numbers across it, so, It varied between 9.2, and 25.3, so it didn’t answer 25.3 every single time.

You know, same question, same data it’s pairing from, but no definition for what retention was is in front of the model. Which is obviously the context it needs, right? So nothing broke here. if we open up any of these runs, we’d find, like, valid, like, SQL and execution of that SQL, behind it. Like, every one of them is correct in respect to, like, the query it wrote, but they just wrote different queries because they were on their own answering different questions. And so you can imagine, you know, like. We spread this out across, like, many runs, like, 30 different runs, we start running into a bunch of different definitions that, Our agents are proposing.

when they don’t have a definition available to them for what retention means. So in this case, we have an audit trail to identify, like, like we showed up here, what each of those definitions are. And, you know, across, like, all these 30 runs. One decided retention meant the share of customers who came back the next month. One decided it meant people who bought twice within 90 days. One found a membership table and used the share of memberships that were still active, so that’s, like, very different than the other two. Then this all happened because we hadn’t written down what retention actually means, so each run had to go try and settle it on its own.

The question we typed was somewhat ambiguous, and so every session’s gonna attempt to resolve that ambiguity differently. Unfortunately, it doesn’t always tell us the choice that it made. And it gets… so you can see, like, when we… when we zoom out to, like, even more of these other engines, so if we split by engine here, they each go their own way. Opus is, like. pretty consistent. In this one run we did of it, it consistently… but it was consistently on the same wrong fork. So 34.9 is not the answer we’re getting from the other models.

A fable flipped kind of between these two answers, almost 50-50, and then our open model actually gave us a different number, almost every every time we pulled it, it gave us a different number. Yeah, 99.2 would be a… would be awesome, Attendee. But, like, you can imagine this, right? Like, you see, like, in the open model, like, one run said, because of one way to find retention, gave us 99.2%, and another run, it gave us 9.7%. They’re just total… they’re actually both, like, pretty decent definitions, if you look behind the… what they… what they pulled. It kind of makes sense.

But if you report that… if you report 99.2 one day, and then 9.7 the other, or someone else Reports 9.7, like, just looks… it looks stupid, right? It’s just, like, it looks like this thing doesn’t work at all. But, you know, it’s pretty easy to fix this. This is 15 runs with no definitions, so you get 8 distinct answers. But you can obviously fix this with context. So the other 15 runs we did out of these 30 were the same thing, like, all, all different engines, but… All we did was we had a written definition of retention, sitting where the agent can read it in a semantic layer, and when we did that, 15 out of 15, every single model, every single run came back, with the same, same number.

So all the engines, including the Open one, which is, much more affordable, you know, like, it’s, like, a seventh of the price of Opus, and I don’t even know how much of a price of… less price of Fable, probably, like, a 20th. all were able to get that same answer. So, this is, like, a small experiment, but it kind of gets to… to see how, Just adding simple context. Can do, like, huge improvements on the reliability of your models. So, yeah, I think before we go further, like. The other thing to note here is, like, models, like, did get better. As well, so dbt has, like, an open benchmark, their code’s public.

I’ll share some of these, sources, because we’re going to go through some different sources of how, like, dbt, Databricks, and Snowflake all approach, Context and semantic layers and stuff throughout this. I’ll share those in an email, maybe this evening. But, like, the model’s getting better does definitely help the reliability, as well. You can kind of see here, right, like, like Fable performs better with reliability than the open model, even without a definition. There’s some more consistency there. So, models does help, but, like. At some point, you basically get some diminishing returns there, or you can get the same amount of, like.

increases in reliability from, I would say, faster, smaller models than, like, these bigger, robust models, just by increasing Just by having more structured context, layers within your system itself, which is a lot cheaper than, like, the tokens you would spend on a newer, more large, robust, expensive, slower model. That’s still gonna get the answer Wrong some of the time. So here’s kind of, like, the, the mechanism. When you ask for a metric, like. what happens… the reason you get all these different answers… sorry, this, like, finds several… this says finds several readings here.

It’s a little off that there, but when you ask for a metric, nobody’s written down, the agent’s still going to hand you a number, it’s gonna try and settle that definition itself. It’s not like rolling the dice, necessarily. It’s gonna read through your schemas. You can watch everything it’s gonna do in the background, like, it’s going… that’s the way the smarter models are… are more reliable, because they do more of the work to try and find, like, whatever their most logical definition for this is, and they will land on that same, like, most logical definition as they prescribe it. more consistently than, say, like, some of the older models.

So it’ll read through your schema, it’ll find several reasonable readings, it’ll pick one that fits, and it’s going to answer that one. But it doesn’t flag that it made that choice, right? Because from where it’s sitting, the choice was reasonable. That’s the whole… Point of these… Agentic systems to reason through these things for us. But it’s going to start that process over from scratch every time you ask it, and so, in some cases, it will reason down a different fork. I think I liked this… I liked this way that, I think DBT said it in his paper, that, like.

Previously, you know, you’d get an error message, and that’s gonna stop your work, that’s going to stop a meeting from happening, or something from being shared out, but now you have these, like, plausibly wrong numbers that are going into decks, that are going into reports, that are going to meetings that people are acting on. So a lot of what we really just want to do here, in addition to adding context so we have more reliable numbers, is that when we know a number is not going to be reliable. We just want failures to basically show up as failures, so a human can intervene early on, whether it’s, like, in the moment at runtime.

Adding that context, or adding that into the system at that time. So it doesn’t make its way all the way to, you know, a share out. What we’ll talk about today It’s basically the kind of 5 things that we’ve found have fixed it for us. This is all… a lot of this is based on our own kind of experimentation, Within our products, within our companies, and a lot of that is taken from kind of, like, I mentioned, like, many of the, like, enterprise solutions, their approaches as well. A lot of them are kind of, like, circling around the same solutions at the moment. So I’ll go through each one of these at a time. We’ll start with defining a metric.

Which is what I’ve kind of been circling around this entire first session. It’s the very lowest list that thing you can do, and it has a pretty… big increase in your, reliability. So, to do this. I mean, this is something we should always be doing, prior to AI as well, but I would say probably very rarely happens in a very, like, robust manner. I’ve been in many meetings where, We’re talking about a number, and people think it’s wrong, and it’s just because they had an underlying assumption that the definition is different than what someone else defined it as. So, the first thing we want to do is have some sort of layer, context layer, where we’re writing down what a metric means.

So you’re writing the meaning, and at this point, you’re leaving the SQL out of it. This is stuff like, what are we counting? One row per what, so like the granularity.

Attendee: You know where the best ones are?

Shane Butler: I do not know where the bathrooms are, sorry. Well, my bathroom’s down the hall to the right. But, the, the what… what are we counting? One row per what? What’s on top, so the numerator, what’s on the bottom, what’s the denominator? What’s the time window we’re working with? What do we want to filter out? all those kind of, like, more, like, metric definition things without getting into, like, the tables or the SQL or anything, right? This is actually, like, it’s gonna help your AI, whatever sort of, like. you’re leveraging for doing analytics at your company with AI, it’s gonna help it be more reliable.

It’s also just, like, a really good thing for your company as well, regardless of AI, that everyone is aligned on the metric, and you have super clear definitions around it. So, either way, this is just good practice for, kind of product analytics, or any sort of analytics. The reason we leave the SQL out. at this point is because your warehouse changes over time, right? Tables get renamed, columns get split, SQL written today can rot, but the meanings and definitions, probably won’t.

You might change them a bit over time, right, as you’re kind of hashing them out, but Especially if you’re reporting a metric to, like, the street, if you’re a public company, like, that thing’s kind of rubber-stamped, and it’s gonna be that same metric forever, even if the tables that, calculate it beneath it change over time. So that’s why we separate the metric definitions from, say, like, semantic layers that focus more on the tables and the SQL itself. The same definition has to still work if you move warehouses, for instance, or swap the model underneath. Everybody… has some version of this, probably. Like, Snowflake has metrics with a SQL expression and synonyms attached.

dbt compiles them so the agent never writes the SQL at all, actually. Kind of like a lookup. Databricks ranks SQL expressions as the highest value thing you can, put in front of it. Then we put the definition in front of that same 3 engines, and we’ll kind of, like, run the grid again. Yeah, so in terms of your question here, do you define metrics in a separate document and feed it to the Agent? So you can do this, like, many different ways. We have, like, a metrics. YAML file.

In our, repo, you can have it, like, within the repo itself that your agent’s Agentic system has built in, or we also have it in a, kind of, like, communal context repo that multiple people can add to that is separate from the Agentic system itself, but that will point at it every time it runs. a, analytical question. It’ll read in those metric definitions first, and kind of save them in cache. There’s different approaches for all these enterprise solutions as well, so… I wouldn’t say there’s, like, one hard rule way of doing it, but I have basically, yeah, a kind of, like, YAML file that has all these, Things we talked about in terms of, like, the mean grain, numerator window filters.

So, like I said, once we have that definition, read, then I mentioned this earlier. when we run it across, all those engines again, all those 15 runs again, we’re able to get… land on the same number, 25%… 25.6. So that closes that drift that we saw at the start. That being said, it doesn’t do anything about, like. The join or the data, this doesn’t necessarily mean that 25.6 number is correct. So, the second thing we do here, that we’ve found is helpful is, you know. you can have a completely correct definition and still get a wrong number, because the definition says what to count and not which… and not, like, the path to count it along.

So… I think the biggest… error we see here is often when it gets the wrong cardinality. So cardinality just means, like, how many rows sit on each side of a joint, so… one customer has many orders, and one order has, like, many line items. If the agent kind of, like, walks that relationship the wrong way around. a single order can show up 5 times, and you can, like, have a bunch of duplicate rows. That’ll inflate the number without any warning. So even in our synthetic data set that we teach with, for instance, the same measure One of the definitions without adding, like, information around joins and cardinality in a semantic layer.

will come back as 5.9 million versus 3.1 million if it’s just going off the definition itself. So the definition’s, like, fine on both times, but one path duplicates a bunch of rows, because we haven’t told it anything about, the cardinality it should work against. So, you want to declare both, how the table joins. How the tables join, how many rows are on each side. dbt does this with typed relationships, and its compiler will actually refuse a join that would duplicate it rather than building it, and then Databricks and Snowflakes both make you declare, Cardinality up front if you… if you use any of those.

But having semantic layers that define joins and cardinality, in addition to, you know, after the definition, is, Another, to make sure that you’re actually getting reliable answers time after time. I’d say there’s, like… there’s kind of, like, a version of this that goes… goes all the way called a compiled semantic layer, where you define the metric once, and then the SQL… the system generates the SQL for you. So then, in this case, the agent never writes Inside that menu of, like, very specific definitions, but you start… incurring a cost when you are trying to query things outside of that menu. So if you ask a question that the layer doesn’t cover.

those questions are obviously all going to score zero, because they’re just going to get refused. While something more like a plain text-to-SQL that, describes the definition and describes the joints and the cardinality, you’re giving the agent, the freedom to, like, reason through that and try to answer questions that you don’t necessarily have, like, a one-to-one map to answer to. So in that way, you’re kind of, like, increasing determinism, but you’re not giving up the range of questions that you can answer.

you can think of, like, the compiled thing as more like a… just like a if-then mapping, where, the way that we’ve discussed in the past two slides is more of, like, the actual agent reasoning through and trying to figure out, what to write while being as deterministic as possible, which is what AI is meant to do. The third… Solution here that we found that gives you more, reliable answers is, providing worked examples. So, like, a real question paired with, the query somebody already ran, checked, verified, and knows is right. So LinkedIn actually, and there’s a… well, this isn’t a link, this is just, but if you want to type that into your browser, you can. I’ll share this out this evening.

But LinkedIn published something pretty useful here. They have a text-to-SQL system. in production with, you know, hundreds of users every week. And what they did is they did what’s called an ablation test, which means they effectively had a, like, evaluation score on their, queries they were running, and then they, one by one, pulled out each layer of their context, one at a time, and measured what happened to accuracy. This is something a, approach that we teach in our course as well, and something I would recommend every person does at their own… for their own use case, because what works for one company’s data may not work for yours.

The different ways of structuring context and the context layers there may differ from use case to use case. But they pull out each thing at a time and basically see what… see what happens to accuracy. And so, they pulled out, like, the table structures, they pulled out the groupings of which tables go together, they pulled the descriptions of tables and column names. they pulled out, like, the worked examples, which I’m talking about here, so, like, verified queries. They pulled out, like, a crowdsourced layer of domain knowledge and jargon, so think about, like, I don’t know, your Notion or Google Drive or Confluence thing, where you’re just kind of, like.

Dumping a bunch of, like, open-ended context about your company. And they see how… what the impact of removing each one of those things. has on the overall accuracy of their, text-to-SQL answers, and of everything they had, bricked examples moved the accuracy the most, which I think shows, like, how impactful these can be. Though these are, I’d say, like, one of the bigger lifts in terms of, building context. So, for instance, definitions, Yes. These are… somewhat of a big lift, because you have to actually have meetings where people align on the definitions, but ideally, somewhere in your company already, there probably is stuff already defined, say, in the dashboard you have today.

And then, you know, the whole, like. cardinality and joint stuff, you can honestly have agents build this for you, just by going through, your, data models themselves. When you get into, like. building out a, worked examples and verified queries. Not only do you have to, you know, find all of those work examples, you have to verify that they’re actually correct. So, I know there’s a lot of folks that I’ve worked in the past, including myself, who’ve written SQL that’s, like, in the query history of our databases, that’s just, like, has errors in it. And so, if you… are having verified queries or work examples that are incorrect.

that’s obviously going to make the output of your system incorrect, too. So this is probably the heaviest lift one, but it is, I would say, worth doing in terms of, increasing accuracy. So I mentioned, like, you know, so every company, every kind of enterprise solution also has some version of this. You can also make any of this stuff you can make on your own, too. You don’t have to use Snowflake or dbt or whatever Databricks version of this. A lot of the stuff we teach in our course, because we want to teach people, like, how it’s actually built step-by-step, is just building it yourself in your own, kind of centralized, GitHub repo. Where all this stuff is stored in.

typically YAML or JSON files. But everyone has some version of this under a different name, so Snowflake calls them, verified queries, and they store who signed off and when. Databricks calls them example SQL queries. And they’ll have… they have, like, some sort of cap on it, I think. I think Looker used to call them golden queries, and now calls them verified queries as well. There’s a couple different… having this kind of, mapping of verified queries can be helpful for a couple reasons, actually. So, you can use it to steer at runtime. If I am going to… Ask a question right now.

the agent can go look at the verified queries, and it can, one, see if someone’s asked my exact question before, and then it’ll just run that query, or find similar questions, and then look at the queries that ran those, and then based on what it knows about what I’m asking. What it knows about the, tables that we, put into our, like, previous layers here, right? It can reason through what changes needs to be made to that verified query to answer my slightly different question. Or, you can hold back some of those queries, don’t ever use them at runtime, but use them as, like, your kind of, like, answer key, grader set, so scoring the agent against it.

This is really helpful when you’re iterating on improvements to your, Agentic Analytics system, whatever it may be, having some set of verified queries where you can see, if adding or removing different things, makes it better or worse. So this is kind of like the example I said, like, with the… LinkedIn did with ablation testing. They probably had a big list of, kind of verified queries they’re testing against. You just want to make sure you don’t use the same pair for both.

If you use it for both, it’s kind of like, if you’re familiar with machine learning, it’s like, using your training data as your testing data, so you’ll kind of, inflate your accuracy because you’re, you’re cheating a bit. We get into all this in our course, when we talk, we spend, like, you know, a week and have a bunch of async content on evaluation and validation. So everything so far has been about giving the agent more information. But no matter how much you write down, there will always be some ambiguity. Leftover. So… that gap’s… the gap that matters, you want one of two things. Either a guardrail. That makes it stop, and ask you. Hey, like, I don’t have enough information here.

This is a very ambiguous question. I need to get more information for you… from you, or you need a sort of, like, sensible catch-all default for the things that go missing most often. So, for instance, like, Looker… For the… they take, like, the default route. If the user doesn’t give, for instance, a date range. They’ll have, like, a rule that automatically filters to, say. the last 12 months, same way every time. So at least you know if you’re missing that context, because there is some, like, ambiguity there. you know what it will default to. And then, the other, the other angle you can take, like, say, like.

what we do, like, when we run via Snowflake is, the question is missing something it needs, we build skills so it, like, marks it as unclear, and will actually stop the analysis and ask us for most context. Both are kind of fine. kind of depends. You probably want a combination of both of those. You probably want a default fallback. And you also probably want some way, when it’s, like, really unclear, where it will stop and not just, like, make that guess without you. The cool thing is that agents really know when things are vague and ambiguous. So, if you ask a question to an agent, like, what is retention?

And then you ask it directly, whether that def… that question or that definition is ambiguous. it’s actually gonna spot that a good share of the time. So… If you put that into, like, a regular kind of question and answer, setting. the… You can easily build systems where you have a guardrail where, like. If we flag this as ambiguous. don’t proceed out of gate here and make us ask more context. So it’s like, it… it does a good job at knowing when to put the human in the loop, if you tell it to put the human in the loop.

But if you don’t tell it to put the human in the loop, we’ve found that it will just find the most reasonable, Whether it’s a join, or a definition, or what have you, and run with it. It’s also really easy to overdo this as well, like, Like, you don’t wanna… like, if you… if you have it gait too many times, you’re obviously starting to lose some of the, the whole reason we’re using, like, Agentic analytic systems, like, you end up basically just doing most of the work yourself as well. And then also, if you haven’t asked bad clarifying questions, this can actually, do a lot of damage as well. So you want to, like. Not right there. Rule for gating so broadly.

You just want that it’ll, like, add friction every time somebody runs it. You want some sort of, like, narrow trigger. I think, like, the best trigger that we use is that, If we… is in a definition layer, if the metrics are actually gonna fork based on a decision it makes, then we ask it to gate and stop. Okay, so last one, we talked about… Definitions, we talked about adding information around, like, your tables and joins. We talked about, kind of like, showed work, verified queries. We talked about, incorporating gates based on, ambiguous, answers or vague questions, so a… agent can get the human in the loop.

And the last kind of thing we found here is, is having some sort of, like, unstructured context. This is kind of where people honestly often start, but it has, like, the… least positive, impact I’ve found compared to all these more structured approaches. So almost every team probably has some sort of document or document repository, a folder of things that explains the business. Maybe it’s in your Google Drive, maybe it’s like Notion, or Confluence, or Coda, or something else. All that unstructured stuff, domain knowledge, kind of like jargon, how we think about things, meetings that went on, transcripts and stuff like that.

So… that ablation test I was talking about, that LinkedIn ran, where they pulled out each context layer, that score you’re looking at is the share of answers a human expert, like, rated, 4 out of 5 here across 130 questions. So when we gave it… when they gave them just, like, tables and columns. The raw structure, it increased that score by 9%. when it told them, like, table groupings, so which tables tend to go together, it bumped it up to 11. When they started adding, descriptions, kind of what we were talking about, like, connect. that was its first, like, real jump, up to 24.

When we started adding work examples, like I was talking about before, the thing we just talked about, goes up to 49%, so that’s a big jump right there when it had, like, verified SQL. But when they added that kind of, like, crowdsourced domain knowledge and jargon, it actually went down to 42%, so… their best score isn’t one with, like, everything in it, all of this context in there. It’s actually that rung before this, like, more unstructured layer, gets put in there, which I think makes a lot of sense.

When you add, When you add these unstructured layers of context, and this may be totally different in the future, as models get, more and more, robust and can reason through this stuff, I’m certain. But you’re adding a lot of noise, right? Like, if documents that have, like, kind of irrelevant information, or nothing, specific to the question being asked. That’s probably not going to do that much damage, but if it’s starting to read in documents, that are specific to or tangential to the question you’re asking, and they have, you know, opposing views on what a metric is, or how things are joined, or what’s important.

You’re just starting to add a bunch of, like, conflicting information that it’s trying to pull from and reason through. You know, like, half think about, like, all, like, the half-correct kind of explanations that could be in there. Like, if you’re doing something like, oh, read everything in Slack and try and answer this question. So you don’t necessarily want to balloon out the context around a question. I find that instead of being… it’s better to just, like, be direct about how to answer it through those more structured formats. Of concepts? Of Context?

You can still use those kind of, like, the… you can still use AI to go through those more unstructured pieces of context, but I think you use those as a way to create the structured stuff. So, a good example of this would be, if you’re coming up with a definition of a metric. you can pull in all the sources from all these different documents, all these different meetings, all these different slacks, and come up with a few different definitions, that you think may be correct based on those readings, and then Go with your team, align on which one of those is actually the correct one, and then pick that, just that correct one.

And put that in, like, your metric definitions, semantic kind of layer, and then when you run the questions, that’s what the… that’s what your agent looks at at runtime. It doesn’t look at all that unstructured stuff, but you can use that unstructured stuff as the kind of, like, base to build those structured things later on. That way, you’re not conflating all contexts. It’s also cheaper, right? Like… You don’t want it reading all this stuff every time, because that sauce is going to cost you a bunch of tokens. Okay, so we’ve got 12 minutes left, and I think I’m about wrapped up here.

This was, like, obviously pretty much, like, a speedrun of what we found useful in building more reliable Agentic analog systems when it comes to just context, nothing else. this is, like, a very, like, teaser version of, like, week 4 of our 5-week course. Actually, we can spend an entire week on Context Management and engineering. And it sits after a week that we spend on, AI evals. And this is on purpose because once you can actually measure your system, and we have many different ways to measure and validate and evaluate a system, then you can start to prove that a context layer earned its place, instead of just guessing and thinking it looks right.

So think back to that example where LinkedIn pulled out every single context layer to see what How it affected accuracy, you can do that at your own company once you have a strong system of evaluation around your, you know, AI analyst yourself. I’d say, yeah, a couple more things. And then we can get into, like, some questions. I’d say over the last year and a half, like, all the kind of serious vendors in this space. like, they all began drafting context with agents. So, you know, I talked about it just now around how you can use unstructured context to start making more structured definitions. You can use agents to draft. All, like, the schema. stuff is pretty clear.

You can start with agents to start picking out some of those verified queries. By going through your query history, That all being said, all of them also put a human approval step in it. So… Agents and can’t get you, like. A really good head start on building these context layers, but humans still do really need to review and verify it, because if you’re… Pushing in the wrong context to the agent that’s gonna run your data analysis, obviously, you’re just gonna end up with the wrong number. So the work shifts a little, right?

It’s less about writing this stuff from scratch, and it’s more about kind of reading it, aligning with your team, and then having someone’s name, you know, a specific date on there that’s been validated, so people know they can trust it. another… I think there’s a whole interesting kind of conversation that could be had around, like, who even owns that, right? Like, And I don’t really have a clean answer to this yet, but today, like, you know, definitions, A lot of them, like, in terms of, like, the SQL that runs, like, that mostly sits with the data team, even though the definition itself probably originally comes from a stakeholder. Right?

Like, even though the definition sits with the data team, they’re probably not the person in the room when somebody decides what gross revenue means for, like, the board deck, so I could see, like. PMs and operators and, you know.

Attendee: I’m adding…

Shane Butler: Timelines.

Attendee: Part of the issue.

Shane Butler: people in finance, being the people who actually, like, create those, like, definition layers themselves, that’s, like, their, like, human judgment part, and then, like. maybe, the data team is more the people who are verifying or creating those, like, worked example parts. I imagine within a year or two, there’s probably going to be some sort of named function or teams arising around, like, people and teams who are just focused on Context Management and Context engineering itself, probably sitting between You know, the data team. and the stakeholders, or one of those people taking on that role itself.

But kind of just, like, an interesting, like, new role, I feel, that could… come up as it becomes more and more important to make these things actually work. Yeah. A little fun test, I would say, try running a test that I ran at the beginning on your own data, right? Like, what I would do this week, take a metric that matters to you, ask… however, we had a lot of people in the chat say they are using AI for some data analysis. Try to ask it that same question in different sessions, like, 5 to 10 times. And see if any of those answers disagrees.

That’s, like, the first way to start identifying, gaps in your system and figuring out, like, okay, there’s probably some opportunity here to, Improve, like, by adding different levels of context to make this more reliable. If you want to build this properly, we have a 5-week course, called Agentic Analytics, Build an AI Analyst, it’s gonna start. In the beginning of September, week one, we… build a thing, what Agentic Analytics actually is, we go through, like, GitHub, how to build your own AI Analyst repo, we build your filter skill, you design and build your own system, rather than just copying ours.

Week two, you expand it, we go into depth into the architecture of what Agentic analytics systems looks like, we integrate systems connect it to real sources via MCPs, like Snowflake, Notion, Google Docs. Week 3 is evals, as I mentioned a little bit before. We do this in two parts. We talk about trusting a single output. So when you’re running an analysis at that given time, how do you trust it based on reliability, based on provenance, which is kind of like the audit trail of how it got to that answer, based on, like, predefined query checks and triangulation? So triangulation is having, multiple takes at the same question, and seeing if they end in the same, recommendation.

And then we also measure whether the system itself is improving. So this is what I was talking about where you’re building out, like. a ground truth data set, you’re doing some sort of deterministic evaluation, you’re building LLMs as judge, for the subjective work, and you’re seeing if the expansions of your system or the additions of context layers is actually making it better or worse. Week four, we get into Context, so we do, like, a super deep dive into what we did today, right? We’re spending, like, an hour on this. We end up spending, like, 6 hours on it this… that week.

So we get into Context stores, where Context lives, why placement matters, how to fix mislabeled definitions, the improvement loops. Paired with evals, writing… Contracts and checking them against gold cases. And then week 5, we get into a bunch of different models. So, you know, we talked, a lot of our stuff, we talk about using Claude in our, kind of, free workshops, because people are familiar about… with that. But we talk about. different models coming out of Anthropic, and why you don’t necessarily need, like, the most… newest, most expensive model from Anthropic to make stuff work. We work with Codex. We go into open models as well.

We talk about, local versus cloud models, we swap between things, we say… we show a way you can trust an open model as an analyst, and we have some models check other models work. So we’ll have… this will be, like, multiple, live lectures. This is not record… we record everything, you can watch recordings afterwards, but these are not async purely async content, because everything is moving so fast that we are constantly rebuilding this course, so it’s on the bleeding edge of what’s going on with Agentic Analytics. I don’t know any other course that’s doing this.

And then we also have office hours every week, so if you are just watching the recordings, after the live sessions are taken, come to office hours and ask questions as well. And then we have a second course, called AI Annx for Everyone. this is, like, a very different problem. So everything today, assumed you already know what to ask and how to tell whether the answer is any good. That’s the part AI is still, like, really bad at. So, you know, it’ll happily run the analysis, but it won’t tell you whether, like. the question you’re asking is the wrong one for the business, or that the finding doesn’t support the decision you’re about to make.

So that’s what our other course, AI Analytics for Everyone, is about. It’s about filling that gap between what AI can execute and what humans still need to do. So if you’re someone who wants to use AI to… in your decision-making process, but you don’t necessarily know how to Ran the whole, like, end-to-end, like. Analytical thinking process, like, pick the question worth answering, pointing the tools to your data, breaking a metric into its drivers, finding the root cause, establishing cause and effect. Running experiments, turning findings into decisions people can actually act on. That’s what this course is about.

It actually started, two days ago, this current cohort, but, that first session’s recorded. A lot of this content is more async, so you can watch at your own pace, and then we have a couple office hours throughout the week. So we’re actually leaving registration open until the end of this week. Since it’s so easy to catch up on the first week, if you still want to join this one. And for both of these, we’re offering 20% off. You can take your phone out and scan the QR code right now for either of these, and it’ll bring you to the 20% off promo on those courses, or you can go check us out on Maven, and you can use the codes below to get 20% off. We’ll send this out in that email there as well.

Okay. I wanted to have way more time to open for questions or anything, but I know we only have a few minutes left. I can stay over a bit. And then… yeah, we will go through, in terms of, like, If… where is the code? We have an open source, repo if you want to check that out, maybe Hai can drop that into the, the chat as well. A lot of our, courses here, you’ll build this all from scratch, though. So a lot of code is… you’re not… you don’t have to know how to code. The Jorgon and build that within, like, kind of your context. Alright, any questions that came up in the chat, Hai or Sravya throughout the session that I didn’t answer?

Hai Guan: Yeah, I think there’s a few that would be interesting to answer. One is any recommendation… any recommendations besides… Your course, which looks very interesting, on where to learn more about how to build context.

Shane Butler: There’s nowhere on the whole planet except for our course that does this. I don’t know, I would say probably YouTube. I actually haven’t seen any courses that are just about, like. building context, but I guarantee you there’s a bunch of stuff that’s free on… on YouTube. If you look up, like, Context management. I will share… so actually what I’m gonna do, well, I’ll say in two weeks, we’re gonna do another free session. on some Context stuff. So join that. Maybe, Hai, you could drop the link to… to that session in two weeks. It’s like… it’s like, It’s gonna be about, like, building, more into, like, the definitions layer. It’s that first thing we talked about.

And then… I will send an email this evening, and it’s gonna have a bunch of sources to some of the, Articles and research published by, some of those enterprise companies who have been trying to solve the context, problem, so I’d read through those, for sure. That would be a really good starting point. So I’ll share that. I haven’t seen a course, per se, on Context. There’s a whole course on evaluation, on Maven. If you’re interested in that. I don’t think it really goes into… that one doesn’t necessarily go into, like, how do you… Increase… the quality of the output via context. It’s more around how do you measure the quality. Which we’d spend a week on as well. What other questions?

Hai Guan: Attendee asked, how successful has been text-to-SQL? What are the biggest challenges? Where, or what type of use cases slash SQL it has been successful? Are these being used in production?

Shane Butler: Yeah, I mean, I think, I mean, Text2SQL’s been around for a long time, right? It’s gotten a lot better. when LLMs came around. I barely raise the cool anymore. I probably stopped writing SQL. I’m pretty good at SQL and pretty fast, but I still find, like, AI is just faster than me. I think that if it’s… if you can… if you’re really good at SQL and you can read it. it’s pretty awesome, like, you can go pretty fast, especially if you have it, looking at your query history of things you’ve ran before. We find that, like. you know, if you think about it, the majority of the, things that you’re querying are probably within the same realm day-to-day.

I think where it probably doesn’t perform as well is people who do not know how to read or write SQL. who are asking novel questions that haven’t been asked before at a company, and thus there’s not existing work example for. It’s going to, in that case, kind of, like. You know, Reason through what it thinks could be a reasonable answer. But if it doesn’t know how to… anything to go off of, and if the person, kind of leveraging it to answer a question can’t review it themselves, I’d say that’s a pretty, risky use case. But in cases where you can review it yourself, or in cases where, like, similar questions have been asked in the past, I think it does really well.

Hai Guan: Alright, maybe a final one. Where do you see QA teams participating in this? How are the roles evolving?

Shane Butler: Hmm. I don’t know, it’s kind of interesting. A lot of companies don’t have QA teams, right? Like… I think a lot of companies used to have QA teams, and then it’s kind of, like, at some point, like, QA became part of, like. the regular dev’s job, so… I would probably say, like. With everything, there’s probably a bit of, like, you know, that happened before AI. I would say there’s probably a consolidation of, like. roles within a single person’s job, no matter what. I would say anyone I guess how I would think about it is, like, for me and my work, what I can just say is that I’ve gone from doing analysis to building Agentic systems that do analysis for me.

To… and manually reviewing those, to now building systems of evaluation that review the Agentic systems, that do the analysis for me. And I think that probably becomes part of everyone’s job at some point. If anyone is trying to leverage AI to automate their workflows, they also have to know how to evaluate it, validate that it’s right, and know when it’s also wrong. So I’d say, like. QA is gonna be really strong already in that mindset of, like. Knowing how to evaluate something is right or wrong, or somewhere in between.

But… I don’t… they might be taking on more of the other dev work as well, because I don’t really think there’s a centralized team that’s just like, I’m gonna build your evaluation system for your bespoke automated workflow of what you do day-to-day. It’s too slow. But yeah, I don’t have a super good answer for that. I think just in general, like, everyone should be probably doing a lot more QA now. if you’re… if you’re… if you’re using something non-deterministic to run your work for you, you have to QA it. But I think people who have been doing that for, you know, their career can be set up pretty well to do that.

Hai Guan: Oh, Attendee just added a question. How do you generalize the learnings here for other types of…

Shane Butler: Yeah, yeah, for stuff outside of analytics. I think… I think, so, obviously, like, verified queries doesn’t generalize, right? That’s very much of the analytics, relevant thing to do. You could have other types of worked examples, of, like, examples of a workflow someone, goes through. For instance, example might be… when I was build… we’re building Agentic systems for lawyers, like, having the, multiple iterations of a document or a contract that they’re editing to show from start to finish how they got to the end product might be a good work example. I’d say anything evaluation. Yeah, eval stuff is gonna be generalizable.

But yeah, I think context, the type of context, and how you structure it and store it is going to be a little more application-dependent, because even, like, right, metric definitions, that’s pretty clearly, An analytics thing as well. I would say the most generalizable stuff that we teach in our course So, if we go back to… build an AI analyst you can trust, I would say, like. Week 4 is probably the least generalizable, as, like, week 1, 2, 3, We’re building a… Agentic analysts, we’re expanding it, we’re putting evals around it, everything that we teach here could be a totally different use case with nothing to do with analytics at all, and it would generalize well.

Same thing with week 5, as we swap and test different models. But how we’re teaching Context, we’re gonna teach it very specifically into, how do you structure context so it does well for, Agentic analytics use case. Just because, like, I think that’s kind of, like, the name of the game I’m seeing right now, is, like, having specific structures for your use case is where I’m seeing a lot more gain in the accuracies compared to just, like, throwing big blobs of context at it. Okay. I think… That’s all the time we got, but if you have more questions, feel free to ping me on… LinkedIn, And, you know what, I’ll share our Slack group. Real quick, if you’re not in our Slack.

We have a Slack community where you can ask questions as well. Let me just drop the link in the chat here. And then I’ll hang around for 10 seconds. So you don’t lose the link if you want to click on it. But you can contact me there. You can also email us at hello at AIanalyst Lab. Let me type it. Can’t type and talk at the same time. Hello at AIAnalside.ai. Any of those ways to contact us. And, happy to answer any questions in the meantime. But otherwise, thanks for the time. Appreciate you all staying here, staying over a bit. And we have… a session… another free session next week on open source models.

We’re partnering with Moonshot AI, the creators of, if you’ve heard of, Kimi, it’s, like, the new open source models that’s competing at, kind of, like, the, Fable… level of models at a fraction of the price, that’s gonna be a really cool one. So, Yeah, if you check out our LinkedIn or something, you’ll see… you’ll see plenty of links to that floating around. Oh, never mind. Hai already painted above, perfectly. Alright, cool. Thanks, everyone. Hopefully see you next week, or at another one. We’re in our course. Have a good day.

Free, every week

The next one is this Wednesday.

10 AM Pacific, live on Maven. One topic a week. Bring a question from your own work.

WED SEP 30
Ace Analytics Interviews with AI
Register
WED OCT 7
Metrics 101: Define a North Star with AI
Register
WED OCT 14
Build a Semantic Layer So AI Defines Your Metrics
Register
WED OCT 21
Experimentation 101: Run an A/B Test with AI
Register
WED OCT 28
Trust Your AI Analytics: Know When the Number Is Right
Register

Next cohorts start Oct 19 and Nov 2.

AI Analytics for Everyone
$1,800 · Oct 19 · ★ 4.9/5
Enroll on Maven
Agentic Analytics: Build an AI Analyst
$2,500 · Nov 2 · ★ 4.9/5
Enroll on Maven
Or come to a free workshop this Wednesday. Register free