← All free workshops
Free workshop · Wednesday, June 24, 2026

Pressure-Test Any AI Analysis

Live on Maven, Wednesdays at 10 AM Pacific. About 50 minutes.

Transcript

Auto-transcribed from the live session and lightly cleaned. Attendee names are removed; their questions are kept.

Shane Butler: Thanks.

Hai Guan: All right.

Shane Butler: Alright, can you all see this screen?

Hai Guan: Yup.

Shane Butler: Okay. Alright, so, this… In case you signed up for this, lightning lesson, like, a couple months ago, we kind of repurposed this time. We were originally gonna do… Something around, via coding annotations tool. We’re actually gonna go deeper into, AI evals, for analytics. Instead, following up from… basically, this will be, like, a part two to last week. If you guys had attended that. So, same kind of vein around, topic, but I think this one gets a little bit more into… it’ll still be on the topic of ground truths, but we’re not going to, like. vibecode annotation tool today.

Depends on when you… I don’t know why Maven doesn’t update the calendar invites, but it depends on when you signed up for this. You may see something else on the calendar. So, I’m Shane, we can do… we’ll do some quick, quick, intros to start. No… yeah, no annotations today. We’ll do that one eventually. I’m Shane, I’m a co-founder here at the AI Analyst Lab. Been in data science for about 10 years, a little over 10 years, last two years, primarily working in, kind of like, the AI evals and, agentic analytics space. Yeah, pass it off to… joined by Hai and Stravia, other partners in the lab here. Maybe I’ll pass it off to Hai if you want to do a quick intro.

Hai Guan: Sounds good. Hey everyone, good morning. My name is Hai. I am a co-founder at AI Analyst Lab, alongside Shane and Shravia. We met at a previous company called Nextdoor, and we’ve been… I’ve been doing the… in the… in the industry for roughly 20 years. And, yeah, excited to… excited to chat with you all. Off to you, Sharpia.

Sravya Madipalli: Hello everyone, I’m Stravia Maripali. I’m the co-founder, along with Shane and Hai, at the Analyst Lab. Around 14, 15 years of experience, now. Started work with Microsoft, later at Nextdoor, eBay. And most recently, Grammarly or Superhuman. So, very excited to have you all today. And go through this, new Lightning lesson on pressure testing AI analysts.

Shane Butler: Yeah, and then just one of the reasons we were changing this lesson is we kind of are revamping how we do our free workshops, so before, we kind of would Have a bunch of workshop ideas, and then… put them out there, we kind of move dates around. We’re gonna commit to doing a free workshop every single week, and we’ll have different kind of tracks of those workshops, so, Like, they may build on each other, although you don’t have to take one before the other one. You’ll see today, we talk about last week’s a little bit. on top of AI evals, and then we’ll have some fun ones coming up in the next couple weeks around, like, leveraging AI to do analysis in, like, Excel and Google Sheets.

So, we’ll get that, schedule out sometime, probably in the next week or two. But yeah, you can expect, I’d say, probably for the next, you know, 3 to 6 months, we should be having free workshop every week on analytics. Today, you know, if we think about, like, AI analytics just as a stage, and I think many of you have been in these, workshops with us before, so this doesn’t come as a surprise. A lot of the work that used to take us, like. An afternoon, or a couple days, or maybe even, like, a week. Submit analysis, multiple weeks. AI is now able to kind of, like, Bring that time down to… Hours, or in some cases, minutes.

It can, you know, write SQL, obviously, run it, build charts for you, and hand you recommendations based off those charts and analyses. And it does that, though, whether the answer is correct or incorrect. And it’s really hard to tell now which one’s which. So, this hour, like last week, is going to be focused on Given that now a lot of that execution layer is… being solved for by AI, how do we know if we can actually trust the output that it’s giving us? So, maybe just two quick polls first.

The first one, and I know we’ve done this a lot of times in the past, but just so I know who I’m talking to, I want to understand where you’re at with Agentech coding tools, so drop a 1 in there if you’ve never used any sort of agentic coding tool. When I talk… I’m talking about, you know, Cloud Code or Codex, or whatever, cursor. Wind surf anything. Drop a 2 in there if you used them, but not for analysis, and then a 3 if you used Agentic coding tools, and you have used them for analysis.

Hai Guan: Lots of…

Shane Butler: Nice.

Hai Guan: and some ones Pretty well distributed.

Shane Butler: Yeah, okay, nice, this is great. So, we’re all across the board here. And then the second question here is, How do you… figure out today, for those of you who are… so we only have a few threes in here, but maybe for the 3s in the room. How do you figure out whether your AI analysis, like the output, is right today? Not how you wish you did it, but, like, how do you actually kind of do it today? It could be anything, right? Like, do you eyeball it? Do you rerun it? Do you make it show its work? Any… anything else folks do here to basically give a check on, how do I know that this analysis that’s coming out of the AI is correct? Nice, use ground truth. eval with another, SQL query.

Ask for logic, yeah, yeah, definitely. Lean on your experience, compare with other sources. Cool. So, whatever you put… oh, wait, review the code, generates… Yeah, so you’re looking at the actual code. And sample check the numbers. Yeah, these are all really great, ways to do it. Ask to cross, analyze, plan it created, Awesome. So we’ll… we’ll talk a bit about, a couple different, kind of basically paths you can go in terms of pressure testing, your analysis, and, Evaluating whether those are correct. So what makes this hard, though, is that when an AI analysis is incorrect. It doesn’t look incorrect, so you can still… it can still produce a… Probable-looking number and a convincing story.

And sometimes that story can even be far more convincing than, Like, stories that we ourselves create. Oh, I like, I like, Attendee’s answer here. Create a skill that every… after every analysis. That has to create an audit package. That’s awesome. So you get number and story, and yeah, they kind of read the same whether the answer is right or completely off. You do have to do a lot of these kind of, checks that we’ve talked about in the chat. Some of those checks are… manual, kind of, auditing on your own. Other ones, like Attendee and a few other peoples here, are more automated. But when you get it wrong, if you hand someone that incorrect answer. with confidence.

It can cost you, like we said last week, in a couple of ways. One, if you put a… if you put a bad number in front of leadership, even if it’s… Has nothing to do with AI, right? If you do this just because you messed up. your SQL query, your analysis on your own. And somebody catches it later, especially if it’s for an important decision. Now you’ve broken trust. Trust is very easy to lose. So, like, I like to say, trust is really hard to gain, and then really easy to lose. And it’s just easier than ever now to lose trust, because we get answers so quickly. Even if they’re incorrect. And then two. you know, and this one’s worse, right?

Nobody catches it, they trust you, and then they actually act on the incorrect number, and we find out, like, a, you know, a quarter later that a mistake was made, and it has some negative impact on, like, the user experience or the business in some way. And so I… I… Leverage your tech analytics every day. And I still get pretty nervous before I put any of these numbers in front of someone that it outputs. Because I didn’t work them out. myself. I’m sitting there kind of deciding how much to trust an answer I never wrote the sequel for. And so a lot of that work, when we used to do it manually. Before AI.

You kind of had all these… like… implicit, friction points in where you had to investigate and check things along the way. So it took a lot longer because of those friction points, but, It also made the number that came out at the end a bit more reliable, because you’re in the weeks all the time. Yeah, that’s exactly right, balance. Then, if you lose trust, now you have, like, really increased scrutiny. On the numbers the next time. So last week, we ran a free session on, on reliability. If you missed it, there’s a… full hour recorded, maybe hire Stravia, if you want to drop the link to last week’s. In the chat.

But what we basically did is we did one eval, or one approach to evals, reliability. The idea is you ask the same question Multiple times to an agent or a collection of agent. In our case, here’s what happened. We asked, like, what’s our retention rate? And we got… we asked it 5 separate times, 5 separate agents, 5 separate contacts, and we got… 5 separate runs, and it actually came back with 5 different answers. So it came back from anything from, like, 30% to 99%. So same question, same data, but 5 different numbers, and every one of those was actually a real defensible definition of retention.

One run counted anyone who ever came back to the product, one counted, only when there was, like, a second purchase within 90 days. We were looking at retention of, like, an e-commerce synthetic, platform. One counted month-to-month retention. All of them are, like, pretty valid, but they’re just different numbers. So, you could picture handing your… stakeholder, or leader, or the board, that 99… Percent number, and maybe they had some totally other definition that didn’t even come here, where that would have been, like, they’re thinking more in, like, the 9% or 8% range. when that happens, you kind of look like you have no idea what you’re doing.

And this isn’t, like, anything… per se, to do with AI, having a strong metric definition, we could easily all, as humans, just go, if I asked everyone in this room. hey, go write this query, or define what retention is. That’s all the information I’m going to give you. Don’t talk to anyone else. We’d all probably come up with different things and come into a meeting together with different definitions. The difference is that, like. now we’re all talking to each other. We’ll eventually align on the metric definition, if I hadn’t ran this 5 times with AI, it would’ve… I would have just got whatever the first run gave me. And, It doesn’t say, like, oh, maybe it’s this.

There’s no kind of, uncertainty in that answer. So this is actually really easy to fix. Retention was never defined, so when we wrote down, like, the definition that we want… wanted, and we ran it again last week, Then we got the same number, run after run. So… Reliability is kind of a cool eval, because it doesn’t require any sort of answer key at all, or ground truth, as some people were alluding to in the chat earlier. And so, that’s… that’s one eval. Today, we’re gonna kind of look at, like, some other parts of the toolkit here. So, an analysis, the… one of the reasons this is… I think this is a bit different from, like.

doing, like, AI evals on… on other kind of AI products, or on just, like, code in general, is that analysis isn’t one thing. It’s a pipeline. So, you’ll recognize this from last week, too. The question… comes in, it picks the tables and the metric, it writes SQL, It runs it, it reads the result, decides what’s worth saying, writes a recommendation. Every one of those is a place where something can go wrong. They all have an opportunity to fail. And so we need at least one, and in case multiple Evals, or checks? For how trustworthy it is at each one of these points. So we teach… we actually teach how to build every one of these into your own analysts hands-on in our bootcamps.

In our enterprise workshops. Today we’re just gonna go Up one level, because underneath every one of these, there’s kind of the same question. Do we have something to measure the answer against or not? And that kind of question around ground truth. The availability of ground truth splits everything into, basically, two situations. So let me define that thing. To know if an answer is right, you need something to measure it against. A known, trusted answer. Ground truth. It’s basically, like, the answer key to the test. And for a lot of questions you ask, AI, though, there’s no answer key written down, especially novel questions.

So you kind of run into two situations, and you handle them in pretty different ways. So, situation one… is… You can get a reference. And the first thing you reach for is always an answer you can look up. So there’s no single way to get one. It depends on your company. I would say, like, start with the numbers you already trust. So, like, finance reconciled figures, the metrics your board kind of is reported to on whatever a monthly or quarterly basis, a query someone already verified, like, a dozen times. That’s ground truth you already have. Kinda, like, never called it that, but basically.

What we call it here is, like, found truth, where, Looks like we got a… hi, if you wanna… Someone’s posting… what’s… what’s app groups in our chat again? We are not affiliated for that. Just… yeah. So found truth is, like, things you already have. You don’t have to create any new sort of answer key. You kind of have your answer key already lurking in the background through, like, years of historical data, and now you can Yeah, just, like, you can basically set these up kind of like unit tests in a way, where you can rerun those same analysis based off those, existing found truth.

If you don’t have one for the question that matters, like, you can also build ground truths, so you’re gonna eventually ask some questions that you’ve never really asked before. And so, take a handful of questions that actually drive decisions, work out the right answer once, and save it, and that’s kind of like forged truth. So you’re now creating a new sample data set of ground truth yourself. this actually takes… obviously, found truth is a lot less work. You’re basically just, recurating a dataset based off things that already exist, so it’s more like just restructuring. Whereas forged truth, you actually have to be spending time and, running analysis yourself.

getting the answer key, and it depends who you’re working with. This could take You know, hours, days, weeks, depends how much forged kind of truth you want for each answer. A number being important doesn’t always make it correct, so… A metric can also, like, be wrong for years, and everyone still quotes it. So, that’s just something to think about when you’re, like, creating your references off of, kind of the found truth. You have to make sure those are correct references as well. So, for instance, If I have queries that I wrote. a year ago, which have bugs in them that have been solved since.

If I am basically grading, evaluing the AI off of, queries or metrics that were incorrect in the past, then… the AI is gonna inherit those incorrect, Methodologies as well. So now you might be thinking, well, if I can just look up the trusted answers, why do I even need AI. Why not go to wherever those numbers already live, the dashboard, the board deck, metrics doc. And just read them off there. So the trusted numbers are more like building blocks, the metrics, the definitions, figures everyone agrees on. The AI’s job is to do much deeper analysis that you build on top of it. So, you know, basically.

The cuts and the analyses that no one’s ever done before, but based off of those data points that we know are solved. So you usually have, like, a reference for the pieces, but you won’t have one for the kind of whole answer that you’re looking for every time, which means… Situation 2 isn’t, like, there’s no… reference anywhere is that there’s no key for the whole thing, but trusted pieces But there are trusted pieces I can, like, tie it back to, which is how the next type of evals kind of work. So, situation two, This is basically, like, the novel questions you’ve never asked before, where you’re not going to have, ground truth that exists.

And, of course, you can… spend time, like, creating that truth and to, like, forge truth, but as I mentioned, like, that is a lot of work, like, it is very worthwhile work, but you’re just not always going to have the resources to do that. So there are some kind of other things you can do before that. Which are a bit lighter touch, but still can increase the trust of your analysis. So, you can’t prove the answer in these cases, but what you can do is basically evaluate the work. And it sounded like this, Some people have this approach, actually, in the chat as well. And you can run these on, like, any output. So first, reliability, that’s that first thing I said, right?

You can run it again multiple times, does it agree with itself? you saw this last week, if you joined our session last week. This is, like, the cheapest eval. It’s also kind of, like, the weakest eval. A stable answer can also be stably… Incorrect. But if it won’t even hold. At, like, stable, where you’re getting, like, a reliably the same answer every time, every run, for multiple agents. then you’ve got a problem before you can even go any further. So that’s a really good place to start, even when you don’t have ground truth, to understand, hey, are we, basically progressing our agentic analytics system that we’re building?

Two is receipts, so I think a lot of people brought this up in the chat. Basically. make it show you where every number came from, the exact SQL, the table, how to find the metric. You’re not judging the answer here, you’re judging the work that produced it. So, receipt… That looks right. Sitting on a wrong query is still wrong. But you can at least have the receipt where you can read the query yourself. Again, not necessarily giving you the correct answer, but giving you all the tools here to go and make the process as visible as possible, so you can see if the methodology to get to that answer was correct. Three, reconciliation, so you can make the parts add up.

So, segments should sum to a total you already trust, a different cut should land in the same place. this is where those trusted pieces from Situation 1 kind of earn their keep. You can tie, like, the new answer back to numbers you already believe in. I think some people I have said this in the chat, too, where they’ll run an analysis a different way. I’m gonna try and tie out numbers. And then four, it’s kind of similar to, the, receipts in a way, read through the trace, look at the method, not just the number, so did it pick a defensible approach? This is less about This isn’t just about, like, oh, is the query right? But, like, the approach of the actual analysis itself, Is that correct?

So this is where you’re gonna lean on… I think someone mentioned, like, yeah, I have 35 years of experience, so I go and I basically evaluate the methodology based on what I would do. If the method is over… your head? Like… That’s fine, but probably don’t act on it yet then. The principle under all four of these is that none of them proves the answer, but each one basically shrinks the space where a mistake can hide. If you run them together, then there’s a lot less room for you to be wrong.

If you’re… if you have, like, reliability, receipts, reconciliation, you’re reading the trace, you have a pretty good understanding of how you got from the start to the finish, from the input to the output. All those friction points we used to hit. When we would do the analysis manually, we’re basically checking all those friction points retroactively, when we… when we create a sort of, A scenario like this where we can check all these different things. And all these can be automated into hooks or skills, so… they’re reproducible every time, just like I think someone mentioned, they have, like, a… A skill that creates an audit. After each analysis run. you’re effectively bounding your risk.

You’re not verifying the numbers correctly, but you’re bounding the risk. None of them’s perfect, they all have their own weaknesses. It doesn’t mean don’t use them, but, It just means you have to know where each one lets something through. So reliability, an incorrect answer that never moves. still passes it, so it can look completely stable, but still be an incorrect answer. Reconciliation, if the error is in every part of, the segment all the way through, so if there’s, like, a filter applied, an incorrect filter or something applied to everything, then the parts will still add up. So that’ll still pass, even though it’s incorrect? Receipts.

A receipt can be faked or incorrect, and what I mean by faked is that with receipts. if you ask your, if you ask an LLM just to create, basically, a receipt of what happened, it can look through its context, but it’s also going to… it just knows what, like, a good receipt would look like. So you actually want to have hooks or something where you’re… logging the actual things that happened, and then have code generate the receipts for you, rather than the LLM hallucinating receipt. So… You can have a receipt. That’s there, but still, just like… Hallucination. And then traces, yeah, a definition error can read fine, even if you didn’t catch it.

So no single one of these, again, is proof, but they’re just, like, bounding… the risk. And knowing where your evals fall short. is what makes you worth trusting. So, then when you go into that meeting, even if you don’t have ground truth, and you can’t say, hey, we are, like. 100% sure this is correct, or 95% sure based off the, kind of precision of everything we were… of these type of analysis we ran off of, ground truth, you can say, well, we looked at We ran it through 5 different agents, they all got the same answer. We checked through the trace and the receipts and the methodology, and everything.

The approaches all looked sound, and we reconciled it against some other numbers, so we’re confident with the answer that we’re going to report here. In addition to that, like, when these things fail, it gives you really good signal. So this is, like, one of the… the best parts of, Evals is basically the sort of, like. identifying where it’s failing, and doing the error analysis afterwards. This lands you on the actual skill, so once you’ve ran your evals, you can do one of these three things. You can act. Because… everything held. And… Based on how much rigor you’ve put into these, prior four, or into your ground truth, you feel like the rigor fits the stakes of the decision.

You can dig because something flagged, and then you can go find out what’s going wrong. So this brings us back to, say. Last week, when we had the reliability one, the first thing I’m gonna do here is start digging. I’m gonna ask, like, okay, what was the definition for each one of these different numbers that we got from each different agent? And then bring those back to… Our stakeholders, or whomever needs to make the call on what that definition should actually be. Or you can, abstain. Where you just don’t act yet. So, saving’s a real decision. Also, it’s not… doing totally nothing. Basically, saying that you’re… not certain.

based on the outputs of these is going to increase trust with your stakeholders, rather than just acting on it when confidence is low, even if you don’t know why the confidence is low at that time. So, for instance, like, if the method’s over your head, and the stakes are really high, you’ll want to pull in someone else. to kind of review that with you, just like you would with any other regular analysis as a human, right? Like, if the stakes are really high and you’re making a decision, I’m probably running it by a couple people on my team. A couple stakeholders before I present it to anyone else. So one more thing, and then we’ll go into a really quick kind of mini-demo on Cloud Code.

We’re not gonna go super deep into the demo today, but… When you catch an encryptor. It isn’t wasted either, it becomes context, so you can write that down, write the correction down. There’s many different ways we… Manage contacts, there’s different… many different ways we create feedback loops based on… They’re all different based on where you are in the analysis pipeline. They’re also all different based on The rigor, that you want at your company. But anything that’s an incorrect answer is really helpful, because now you’re able to add that as context, correct the system, and create feedback loops so the next time, you don’t run into that same problem again, or similar problems.

So you can put kind of, like, a guardrail into your system on that, correction, and it doesn’t get trusted until it’s verified, so… that keeps context from kind of, like, quietly giving your… your system, basically, like, you know, like a brain rot, or, poisoning your future answers. So we teach that whole loop around AI analytics validation and evals, plus how you manage the context, and then how it feeds back into the system, and you eval again in our Agentic Analytics 201 course, which is actually kicking off this, Saturday. We’ll have some, promo codes and stuff at the end of this for that.

Okay, so I’m gonna run, like, a really quick kind of demoing Cloud Code, and then we can get into… questions, But, let’s say… We have, we’re working with this, e-commerce data set. From this, fictitious company called Novo Mart. You can think of them kind of as, like, an Amazon. And say we have, we get, you know, some revenue by category breakdown, breakdown, or requests from the board, or… some leadership stakeholders are like, hey, I want revenue by category, and we… get some output from our Gentic Analytics system. And, obviously, we didn’t build it. We just asked them, like. What’s the revenue by category? Breakdown. And now we’re like, I don’t know if this is correct or not.

So we’re just gonna show… we have some, validation checks built into… our aging system already. And I’m just basically going to paste in this question… this, this prompt right here. I’m gonna ask it, hey, here’s… here’s the SQL And the breakdown for revenue by category for… Mark, can I even trust this? Sorry, I have it pasted in here. It does. So… The first thing it does, we have, a few different skills, and we have some stuff in our .md file, too, which has, common, basically, like, failures or risks of any analysis. And so, under segmentation, one of the main risks is, like, a fan-out risk, where you can basically end up, duplicating A bunch of, rows when you’re joining across tables.

And so that’s the first thing I looked for, which happens to be, One of the primary issues here. So it says confirmed. Orders has 47,000 rows. out to order goes order items, 75,000 rows, and total amount is an order level column being summed across all line items. So that total amount is in the orders table, but it’s joined to order items. And so it’s going to be duplicating those rows, if it calculates. So it gives us… Its whole kind of, reasoning here, and the multiple problems that it found. Just right off the bat. So I reproduce the query on our live data. and validated it 3 ways. Has 3 independent promises. The first one is that join fanout. This is, like, the biggest one.

So, it sums our total amount, but… Because that lives at a different grain. The joint basically explodes, and you get, multiple… rows per… Total amount, so every order total gets counted more than once. Let’s see if it gives us the actual number difference. Here it is. So, when we do that join. our total comes out as $6.9 million, when actually it should be something more like $3.6 million. So the query overstates the revenue by 87%. That’s, like, a pretty easy… Mistake to make. I’ve seen, like, human analysts make that mistake. I’ve made that mistake myself. Yeah, Attendee. Treat your own.

Attendee: Yes, indeed, thank you. So, Shane, what you did is you took a trace from production, or some test, and… and… Saw… took the query, and the results that they got, and essentially ran as if… essentially, did it produce the right SQL? And from the SQL, did it produce the right, result? Is that correct? Because that… what is important is that from the query to SQL, did we do the right step? Because getting to the… if you didn’t get the right SQL, Definitely, that’s a problem.

Shane Butler: Yep, that’s… that’s… that’s absolutely correct. And so, basically, what we’re focusing on right now with this question is all the way up here, like you’re saying, so…

Attendee: Okay, okay.

Shane Butler: So when we think about this pipeline, Anything… up, like, I think the most important things to get right are higher up in the pipeline, because, like you just said. Like, if the… if the SQL’s wrong, everything downstream is gonna be wrong. Same thing with the question. If we misinterpret the question.

Attendee: Then, you wouldn’t know, you would have to randomly select some traces from the production, unless the user has given a thumbs up or a thumbs down.

Shane Butler: Yeah, you’d have to actually have… so basically what we do is we… Have a hook that records every, query. That is ran, and then it stores it in, like, JSON file for us. In this case, I just… I just pasted, hey, here’s the actual query. Right.

Attendee: Yeah.

Shane Butler: But in practice, you would basically… Force it to save. the query it’s running, so then you can have, basically an audit go back to that exact query. Otherwise, like you said, like, it’s not gonna know what to even run against.

Attendee: Got it, yeah, perfect. I fully understand, and then the good point that you make is that the whole analysis is like a pipeline. And it starts with a trace, but you have to do it as a pipeline and inspect every point in time, but you kind of guess where the failure is going to be, and kind of zoom in on that particular point.

Shane Butler: Yeah. If I had just said… like, I wouldn’t have just said, like, hey… I got this number 6.9 million, is that correct? I needed to add it, give it that additional context of, like, here’s the query it ran, here’s the breakdown it provided, then here’s the total it gave me, is that correct? And you can… you can kind of orchestrate that for… this is, like, a very, kind of, like. baby version of it, but yeah, you can kind of orchestrate that for any point in that pipeline.

Attendee: Yeah, perfect. Thank you.

Shane Butler: Yeah, no problem. So in this case, the first problem it looks for, it has a bunch of different things it checks for when it checks for SQL, and it’s like, this isn’t, like, a SQL syntax check, right? It’s not like… obviously the SQL can run. This is, like, a more, like, logical check of common errors that can happen when anyone does analysis, so… I’ve definitely done analysis in the past, where I’ve… Join to some table, and then, sum to some number from the First table, and had a bunch of duplicates, and ended up with the, overstated, number. So that’s what it’s gonna check for.

The next thing it checks for is, We have, basically a, semantic layer here that, that shows us the… Information around what goes into the… Different amounts. And so, it picked up immediately, like, if I check on those, like, what’s logical to include in total revenue? This one’s gonna be up to… this is one of those definition problems where we don’t necessarily know what the correct definition is, but it flags it for us so we can ask the stakeholder. In this case, there’s no status filters on, say, like, canceled or returned. returned orders. We didn’t say that we wanted completed only, but it’s flagging it for us, just in case, based off of, that semantic layer.

If we were only looking for, completed orders that weren’t canceled and returned, then, we would be overstating by another 15%. And then I think, yeah, looks like its final definition here, again. another definition problem here. Revenue definition… Total amount includes shipping. And is net of any discounts, so we might want to use line total. instead, which is the revenue, net of line discounts, like shipping and, like, promo codes. Like, when you, when you, you know, exit your cart on Amazon or any sort of, like, e-commerce site, and you, At the very end, they spring the shipping on you, and then you can… potentially add some promo code, to reduce things.

So, it’s actually saying here, you don’t want to use the total amount, anyways, you want to use line total against This is, the benefit of having you know, like, semantic… views where we have the definitions around each of these columns, and we have Claude go and, when it audits. Look at those, semantic views. When it’s doing any sort of audits. And then finally, at the end, it gives us, like, given all that, here’s the corrected query you could run. And I might even say for next time. yeah, run the corrected query. Sure, we can run it, actually.

Run the corrected query, and then you could also ask it to basically Update your context in such a way, your definitions in such a way, so in the future, if we got this question, It would run the corrected query the first time around. Oh, and there’s no Q2 2026 data. At least our data stops at 2024, so it’s good to pick up on that. But I think you all get the point. Alright, so back to our slides here, and then we’ll open up for questions, since we’re almost at the top of the hour. So, basically, you just watch eval run on, kind of like, live analysis, when we had a reference, we scored against it and fixed what was wrong.

When we didn’t, we could kind of prove the answer by going through these, basically… When it went through that, like, that fan-out thing, that’s more around looking at the, The methodology itself didn’t necessarily go and compare against an answer key. That’s kind of the whole thing, like, how hard you push on any of this comes down to what your stakes are, in practice, I’d probably, go quite a bit deeper than this if it was a question for the actual board, but if this was a question that, say, like. stakeholder alert, which is curious about. This might be… Just a fine altitude to, get that number to them. So we’ll… we’ll send out a couple… I think someone was asking about the recording.

We will… share the code link. Yeah, we’ll share a couple of links in the, in an email. I basically have a… let’s see… so one thing I have, I have this one-pager. Let’s see, I posted it the other day… I’ll share this out, so I have, like, a kind of one-pager on… Once you use ground truths, what to do when you don’t have it, how deep to go. I’ll share that out in, email after this. I’ll also share out the recording in the next day or two, and then we’ll share out the AI Analyst repo… So, basically, how we ran that code. And we will share out the slides. So, I’ll email most of that stuff out today.

And then if you want to go further into this, we do have, like I said, we have free… free workshops on AI analytics every week. Next week, top of the list here, your next Lightning lesson’s gonna be Analyze Data and Google Sheets with AI. That one’s gonna be next week. And we can drop the registration link for that in the follow-up email as well, if you’re interested in joining us for that one, that should be pretty cool. And then we have a few courses, full courses, that we run. So we have a one-on-one bootcamp. That’s where you build your own AI analyst. Next cohort for that is July 13th to 17th, so just coming up in a few weeks here. And yeah, it’s great.

It’s, like, 4.95 out of 5 stars on Maven. Highly recommend, checking that out if you want to, start leveraging Agentic Analytics. And you haven’t before, so if you’re in that kind of 1 or 2 range from the beginning of this, the 201 is where we go a lot deeper on what we did today. So validation and context management, we also get into a bunch of open source models. We’ve been finding that A lot of the open source models available now are performing just as well on Agentic Analytics as, Sig Flawed Code or Codex. And you don’t have to spend a bunch of money on them. It’s great. And they’re more secure, because they’re open source, and you can literally have it on your machine.

Or in your own safe environment, you’re not sending anything to a third party. So we’ll go into those as well. That actually kicks off on this Saturday. We’re running the first cohort on that, and then we’ll run it again in maybe… Another 5 or 6 weeks after that. A, we have promo codes for all of these, maybe… I’ll drop them in the email, but hi or Shravia, you want to drop those in the chat? as well. And then we also have our 5-week course, AI Analytics for Everyone. This one’s all about, kind of, the end-to-end analytical thinking. You can kind of think of it as, like, a mini master’s in data science.

But it is, rather than learning data science, in executing in R, Python, or SQL, you’re executing in… Cloud Code. Right now, we’re actually in the second week of that, so this cohort’s currently closed. We’ll have another cohort in the beginning of August, and we might even add open source and codecs to that one, by then. You can check us all out at AIAnalysislab.ai as well. We have a ton of free resources there, too. So we have a free AI evals course there that’s, like, self-paced, going through decks. We have a free email course on cloud code analytics, which is really popular.

I just had someone ping me, like, 2 days ago, actually, that they got a job just from going through our cloud code analytics, free email course, and talking about what she did, leveraging that. we have… Obviously our free open source repo of the AI Analyst. And, we just launched another… open source repo called Northstar. If you joined our North… Defining North Star Metrics with AI about a month ago, this is, like, the agentic system that allows us to audit and, kind of pressure test North Star metrics. And then… oh yeah, one more promo. If you take the AI analytics for everyone, you take that 5-week course, we throw in our one-on-one bootcamp for free, or for 10 bucks.

So, Yeah, if you sign up for that, DM us, or email us, or if you’re interested in that. Or join our Slack community, and we can get you set up with the bootcamp, usually $900, just for 10 bucks. Yeah, yeah, for everyone, we rebranded from, for builders. Okay. I think that’s all I got. We got about 10 minutes left. Can open up for any questions folks have, or if there’s any questions in the chat, Savi and hi, that you saw.

Hai Guan: See if the audience has any live questions. Then you cover all the chat questions already.

Shane Butler: Attendee, it’s no… no difference. for the For Everyone, and for builders, I think we’ll probably… expand it a bit beyond, working in Terminal at some point to make it a little bit more user-friendly, but at the moment, no. Some questions I get sometimes, what if I don’t have, like, a data team, or I have really messy data? Can I still… eval, can I still use Agentic Analytics? Can I still eval it? Yeah, there’s definitely, You’re probably who it helps the most, and you can lean on those cheap evals if you have messy data. So, like, reliability and receipts, which sounds like a lot of people did, kind of, the receipts. Thing earlier, as we’re discussing in the chat.

That’s, like, a really good place to start. And then you can save the heavy stuff, like setting up Systems that eval against ground truth. For when you have a few numbers that, like, really matter, that you want to have double checks on, so… A lot of our philosophy is that the rigor in your evaluation of Agentech Analytics output should match the, stakes of the question being asked. So, if it’s a really high-stakes question, You can go pretty rigorous in terms of how much you pressure test that output. But if it’s something someone’s just curious about, or just like, oh, it’s just, like, a number we need to put out for, like, a marketing thing, or we have a kind of, like, in the middle, like.

product decision to make. You can match over here to that. And I think we talked about it already in terms of, like, how do I… Get ground truth without answer key. Before… the found, ground truth is where I always start, so… analyses… And it’s not necessarily, like… I used to say, like, you know, you want to start with analyses people did in the past. Those analyses could be wrong. Like, humans have a lot of assumptions they make. And the analyses that they run as well. So, you actually want to kind of, like, those are an okay source. You can kind of, like, review those, and, like, you want to double-check them and make sure those are even, like, a good ground truth to base off of.

Lots of us run bad analyses. But things that are really important numbers. like, anything that’s, like, a finance kind of reconciled number, anything where it’s like, oh, we report this to, the street, or to the board, or leadership on a regular base… on a very regular basis that you know has already gone through many human reviews, that’s a really good place to start, I think. And then you can get into past analyses. And then you can get into creating your own ground truth. But I always try and find stuff that’s existing. Before investing the time in creating new ground truth, because it is, left.

Hai Guan: Yeah, and for people who have, like, a data engineering team, like, you know, the next level would be investing in semantic layers and stuff like that, where you build those ground troops effectively, and then have a proper maintenance to that. This is a really great head start for that.

Shane Butler: Yep. Any other questions? No worries about.

Hai Guan: Alright, this is not a question.

Shane Butler: No kidding, man, it’d be serious.

Hai Guan: We’ve got a lot of unwelcome visitors.

Shane Butler: Cool. And I think, you know, I kind of touched on these other two already. during… our thing, like, if it’s stable, is it correct? If multiple models agree, is it right? Doesn’t mean it’s correct. But, so stable, just being stable is, like, not necessarily sufficient. But obviously, if you have really unstable answers or unreliable answers that change every time, that’s what you gotta fix first. And so, I think you can learn a lot from that.

It’s a really… that disagreement’s a really useful signal, and it sparks a lot of, conversation and discussion with… between yourself and your team as, like, hey, what do we believe these definitions should be, or how… do we think this context is being incorrect?

It’s kind of interesting, I feel like, with this kind of, everyone starting to leverage AI for analytics, it’s, like, a really strong forcing function for us to Just, shore up all of our data foundations and actually, like, make these calls on definitions that Or kind of… Easy to kick the can on, in the past, because we just didn’t have the bandwidth For enough humans to run the analysis for these things to be as big of an issue as they are now. Alright, I think that’s it. If you have any other questions, feel free to DM us in Slack, or email us at, hello at AIanalystlab.ai. And I’ll send out that email with all the resources. Later today, probably. That’s… it from me.

Thanks for joining, everyone.

Hai Guan: Thanks, everyone.

Free, every week

The next one is this Wednesday.

10 AM Pacific, live on Maven. One topic a week. Bring a question from your own work.

WED SEP 30
Ace Analytics Interviews with AI
Register
WED OCT 7
Metrics 101: Define a North Star with AI
Register
WED OCT 14
Build a Semantic Layer So AI Defines Your Metrics
Register
WED OCT 21
Experimentation 101: Run an A/B Test with AI
Register
WED OCT 28
Trust Your AI Analytics: Know When the Number Is Right
Register

Next cohorts start Oct 19 and Nov 2.

AI Analytics for Everyone
$1,800 · Oct 19 · ★ 4.9/5
Enroll on Maven
Agentic Analytics: Build an AI Analyst
$2,500 · Nov 2 · ★ 4.9/5
Enroll on Maven
Or come to a free workshop this Wednesday. Register free