Shane Butler: Hey, welcome everyone!
Hai Guan: Good morning.
Sravya Madipalli: Welcome, welcome! Good morning, good afternoon, good evening, wherever you are from.
Shane Butler: I am just turning the waiting room off right now. There we go, that’s off. Hey everyone, I’m Shane, Joined here, with Hi and Stravia. We’ll do intros in a quick second. I think a lot of you have been to our… I see some familiar faces in here already, so I think a lot of you have already been to many of our sessions. But just to get it started, maybe drop in the chat where you’re joining from. Want to get an idea of, you know, your city, your country, wherever you are. Are you in the living room? Are you in the office? Are you in your home office? What are your coordinates, exactly? I’m located in South Lake Tahoe. California. Nice.
Hai Guan: Ireland. Boston, Toronto, a lot of East Coast folks. Ukraine.
Sravya Madipalli: We could be biasing the data by keeping it at 8 AM PST.
Shane Butler: Yeah, actually, we moved our… we moved our time zones, because I feel like, we were doing a lot of our… our, our things at, like, noon Pacific time, but I think that misses a lot of time zones.
Sravya Madipalli: But look at that, there are Santa Clara, there is also Seattle, not bad, West Coast. That’s awesome for an ATM meeting.
Shane Butler: Nice, bright and early.
Hai Guan: Attendee’s from Nairobi. Very cool.
Shane Butler: Sweet. While those roll in, we can do some quick intros. I’m Shane, one of the co-founders of the AI Analysts Lab. We do, kind of, research and education around the field of agility analytics. Been in… working in data science for a little over a decade. past couple years more in the, kind of, AI evals, AI analytics space, so that’ll be the topic of today’s, which we’ll get into a little bit. I’ll pass it off to Hai to do a quick intro.
Hai Guan: Hey everyone, my name is Hai. I’ve been in the data science and analytics space for around 20 years. most recently at, consumer tech companies, spent time at Nextdoor, that’s where I met these folks, co-founder of AI Analyst Lab. We provide a lot of these workshops like these, both free and paid, so very excited to have you all here. Over to you, Sravya.
Sravya Madipalli: Hello everyone, I’m Shavya Madipali. I am one of the co-founders, along with Shane and Hai. I have around 15 years of experience, starting with Microsoft. I see some of the Seattle folks, I was there for 6 years. Later moved on to eBay, Nextdoor, and most recently, Grammarly, which was rebranded as Superhuman. Yeah, so very excited to share everything that we are doing in the world of AI, guys. agenting analytics, and, you know, it’s pretty awesome to be part of, in this domain. Yeah. Over to you, Shane.
Shane Butler: Cool! So we’ll get right into it. I know everyone’s got busy… is busy, got a lot of stuff to do. If you can’t stay up for the whole session, we’ll send out a recording. We’ll send out an email afterwards with some of the resources we’re going to go through today. We’ve got some free resources for you, and then we’ll send out the recording, I don’t know, probably, probably in a couple days, by Friday at least. So, where we’re headed today.
So a lot of the… I’m sure a lot of the… if you’ve seen our, kind of worked before, if you’ve been messing around with AI on your own, and you’re working in the data field, you’ve probably seen that AI, like, it can now kind of run an end-to-end analysis, in minutes, really. It can write the SQL, it can run the SQL, It can make a chart, it can reason through and make some recommendations, and it can be really, polished and confident and also sometimes completely wrong. So, getting the answer kind of has stopped being the hard part, and then now, like, knowing whether you can even trust the answer is where a lot of the work lies now. So that’s what today’s about.
I’ll show you, there’s a whole spectrum of ways you can kind of check the work. There’s ways to do it manually, there’s ways to do it. completely automated fashion, and then there’s… there’s space in between. We’re just going to go through one check live today, because we don’t have that much time. We have a whole… we have multiple courses on this in our first… in our 101 bootcamp around building AI analysts. It’s like Building Agentic Analytics 101.
We go through, five of these checks, 3Life, 2 NSYNC, and then we have an entire, advanced 201 Agentic Analytics course where we primarily just focus on, on really AI evals validation, and, context engineering, which happens to be a lot of the ways that you, Increase, the trustworthiness of your output. So, we’ll hop into it. I do want to get a couple reads of the room before we start. So, first one. Drop a letter in the chat here, just to see where folks are. A, if you’ve never used Cloud Code before, we’re gonna be doing a little demo here with Cloud Code. B, if you’ve used it but not for analytics, and C, if you’ve used it, and you’ve used it for analytics. So yeah, go ahead in there.
What do we got going on in the chat? Hi, Answabia.
Hai Guan: Mostly seas, wow. And then some B.
Shane Butler: It’s good for us to get a pulse, I think, also just on, like, where things are. If, you know, we asked this… we asked this question in almost every single one of our lessons, and when we asked it 4 months ago, I’d say everyone was, A or B. But a lot of people are using this for… Analytics now, which is really awesome, and also just makes it so important to be like, hey, now do I trust this thing? How do I set up evals? I’ve kind of done the step one. I used it for analytics now. We probably need a D on here of, like, are you… do you trust when you use it for analytics? Okay, good, that’s awesome. This is who it’s built for. If you’re in A or B, that’s great too.
Next question, also, just to get a kind of spread of the room, what’s your role? And you don’t need to write A, B, C, D here, but you know, what’s your role? Are you a product manager? Are you a founder? Are you in marketing? Are you an engineer? Are you a designer? Are you in data? Chef one. We had the corporate chef the other day, and I’m not… that’s awesome. And then, you know what? I… you know what? I started using Claude for cooking a little bit. I’ve been telling you guys, I’ve been using it for, like, my fitness and nutrition tracking, and then yesterday it was, like.
you had too many fatty foods earlier, you have to have this many grams of carbs, and I was like, alright, tell me how to do it, and then it, like, made me this delicious burrito bowl recipe. Tell me how to cook chicken in a little different way, too. Which is not the topic of this, but if you guys want to know a chicken, Mexican, high-carb burrito bowl recipe, where you can cook the chicken from frozen, doesn’t even need to be thawed, works pretty well, took 30 minutes. Thanks, Claude. A lot of founders in here. Chef turned data engineer, yep. to… back to Chef. Cool. Alright, thanks everyone. Lots of people on the data teams there. good to get a read of the room, who’s in here.
So, to start off, what we’re gonna even… what are we even gonna talk about? So what is this, like, AI analytics thing? Because this is, like, a newer term, and it’s a bit vague. So, when I think about it at its, like. most high level, I would say, like, agentic analytics means something like, you ask a question in, it says English here, in natural language, right? Natural language question, something like. why is checkout dropping? And you have a system that goes and writes the SQL, runs it against your data, builds a chart, hands you back a recommendation. There’s no, like, human in between sitting in the middle of doing it, just kind of does the whole thing. So, input a question.
You add, with that question, you add, you know, your data, whatever context you’ve given it, and then the output is a number and, you know, a little story about what happened. And usually there’s a, here’s what you should do about it. As well. So a year ago, I think, you know, I’ve seen a lot of demos on this stuff. We run a podcast called Data Neighbor Podcast, where those demos got increasingly better and better and better. You know, a year ago, I don’t know, it kind of seemed like a little bit of, like. a party trick, like, what’s really going on in that demo? Is this… is this gonna work on my data?
And now people are really actually shipping real, decisions off the back of it, so I think it’s moved a long way, especially in the past, I’d say, 4 to 6 months. So that’s… that’s the shift, and that’s why, you know, as it scales, and as more people are using it for real decision-making, it’s why that kind of, like, checking the work part of it suddenly matters a lot more, than it used to. So what makes this difference from a person doing the analysis is simple. One that’s wrong? Doesn’t look wrong. So, you know, you get a really clean-looking number. a clean story, like, every single time, whether the answer is right or completely off. There’s no, like, hmm, let me double check that.
And it hands you… it basically hands you that wrong, answer with exactly the same confidence. It hands you the right answer, right? So… and that costs you… that costs you, I think, in a couple ways, a few ways, probably more than this, but the main ways I think it costs you, is that I’ve seen happen a few times already. The first is trust. So, if you put a wrong… Confidently wrong number, and you, like, give it to the board, or a leader, or even just a teammate. you don’t really have a lot of that trust built already, or you don’t have, like, a kind of, like, understanding of we’re in, like, experimental mode right now. When somebody catches it.
now nobody… nobody believes that tool again, right? Nobody believes that approach again, nobody… maybe people don’t, like, trust you. That’s… maybe that’s fair, maybe that’s not… But there’s a real, like, like, trust is, like. highlights to say, like, trust is, like, really hard to build, and then really easy to lose. So, you don’t want to lose trust, that’s… that’s one… one thing that happens when it goes wrong. And then the second, I think is worse, is this, like, oh, people all trusted it, it was wrong, and then we acted on it.
So we, like, ship a thing, that moved the number instead of the thing that actually helped a customer, and you… you don’t find out for, like, a quarter or something, and it just depends… it comes back to this, this system that gave you a wrong answer. So you can make real wrong product and business decisions off it, right? Okay, let’s do another kind… let’s do a little thinking exercise here. So drop in the chat, and it doesn’t have to be one of the things up here either, where do you think… AI analysis goes wrong. Is it, like, the SQL? Is it picking the wrong metric? Is it, like, the recommendation at the end? Is it something else? Feel free to throw a guess in the chat.
Definitions and gotchas. Hallucinating data. Wrong context to interpret analytics, incomplete context. Wrong metric. Everywhere! Context and nuance. Oh, oh, Attendee said it almost happened to him today. Yeah, you… yeah, I mean, like, knowing about the business is a real… that’s a real check. Bad storytelling. Company history’s incomplete. Cleaning analysis? Yeah, these are all really, really good. Probably better to go where it goes right today, since the slices thin. Yeah, I like that. Cool, yeah, so you can see it goes wrong all over the place. It can go wrong at any of those points, which is exactly kind of where I’m headed next here, so… this, you’ll hear this word, AI eval, evals, coming.
So, I’ll def… I’ll give you, like, a pretty, like, I don’t know. my, like, highest level definition of it. Sounds like a little technical one, but it isn’t really. You know, an eval, it’s just a check that tells you how much to trust an answer that you didn’t work out yourself. So, the system handed you a number. Or handed you, as some alluded to in the chat here, a story. And, you know, you didn’t do the math. So, an eval is going to help you figure out how far you can actually lean into that number, how much you can actually trust it.
getting the answer, like I said before, it just keeps getting easier and faster, so knowing how much to trust it, like, that’s actually going to get easier and faster too, but that’s just, like, where the work has moved. So… I think about the shape of this, too. It’s not just a yes or no stamp that says this answer is true, or this answer is false. Like, sometimes we’ll have that, and in cases where we actually have that, then it’s kind of like, you don’t even really need this system, because you’re probably pulling from some… some other place. It’s more about, like, how much on a scale that you trust it.
And you can match that to the stakes of the decision you’re gonna try to make with that number. So, like, you don’t have to go and do some crazy rigorous, eval, pipeline for every single number you do. If it’s a throwaway question that someone’s asking just because they’re curious. And they just want some directional understanding, you can kind of, like. do some pretty quick, barely, barely check that number. A number that you’re gonna, like, report to the street, report to the board, or that, like, a big decision leans on, like, a real decision. you can’t really check it, like, you mean it. So, so trust is, like, I think it’s put as, like, a dial, not a switch.
And then we’ll come back to this at the end… at the end, but like I said before, that… this whole discipline of evals for analytics is what our, Agentic Analytics 201, bootcamp. is all about, and that is coming up in a week and a half, or… I think. It’s a week-long boot camp. a validation. So… If we think back to that picture from the start of the input and output, question goes in, recommendation comes out, it’s easy to think of that picture as the middle, as just, like, one box. that kind of, like, does the thing.
But as we’ve seen just in the chat with everything that you all just said around, like, you know, the different ways this can go wrong, clearly it’s not just, like, one little box in between that, like. turns your input into an output. There’s a whole pipeline in there. It interprets what you asked, it picks which tables, it picks the definition of a metric, it writes the SQL, it runs it, it reads the result, it decides what’s even worth saying, and then it writes a recommendation. So that’s like 7 steps right there, and it’s probably most definitely not exhaustive. Each one of it Each one of those steps has multiple places where they can totally go off the tracks.
And it really matters when you think about this in terms of a pipeline, too, because it helps in terms of, like, prioritization of what you’re going to… where you’re going to want to spend your time in evaluation. An error that’s really early in this pipeline, like, you know, query generation, question intake, analysis, That’s gonna carry through to everything after it. So if I grab the wrong definition of an active user back in step 2, every chart, every confidence sentence after that is just sitting on a bad number, and they’re all still gonna look clean, right? So… Is it right? Isn’t really one question, it’s more like. Many smaller questions, at each of these steps.
So, I just dropped some of the checks. Again, not exhaustive, but, like, some checks we do at each of these steps, so at each one of these steps, there’s something you can actually look at. Like, did it understand the question you asked? Is the SQL correct? Is it using the metric the way you mean it? Did it pick a, like, robust, defensible method, methodology of analysis? Does the chart match the number underneath it? Every step has its own thing to check. And you can keep it pretty manageable, because each of these checks is really just kind of, like, one of four kinds, and I’ll show those in a second.
But the darker ones here, if you, like, look at the shittings, like, the darker ones are kind of, like, where I’d… I’d spend more of my… my time, rather than kind of these. These ones are, like, easier fixes. So there’s a few checks also that don’t sit in any single step. They run across, like, the whole, kind of surface area. Sorry, I’m just reading the chat here. 7 steps, 9 questions, Is slide generator failing? Yeah, probably. Oh yeah, the 9 questions are gonna be, like, all these checks over here. There’s, like, so there’s multiple questions under… under each of these.
So, these, these checks that don’t sit at any steps, they run across the whole thing, so it’s like, do you get the same answer if you run it through more than one model, sort of reliability. Can you trace every number back to the exact query that produced it? Reliability, which is just asking the same question again and again, seeing if it holds running kind of underneath everything here. You put those per-step checks together with these few that are cross-cutting across everything, and that’s kind of our… I don’t know, if you get that, you get a pretty robust system. Not exhaustive, a pretty robust system of evaluating. An agentic analysis pipeline end-to-end.
And then so… so I guess, like, from this slide, the main thing I want to take… take away here is, like, eval is not a single score. It’s, like, many checks at each of these steps. plus some that run through the whole, kind of, gambit. So it’s a lot. So the obvious question… Then, is how do I actually run any of them? Which ones do I run? how do I… to what level of rigor do I run each of these, depending on my specific, kind of use case? So you don’t… you don’t really need a different trick for all nine… of these, checks down here. There’s really kind of just, like, a handful of ways to check any output. So, here’s four of them. One, check it against a known answer.
You’ve probably heard of this as, like, ground truth. Not always gonna have it. But if you happen to have a known answer, that’s one way to check. To grade it against a rubric. Just kind of like, you know, you’d grade an essay, for instance. And humans can use those, rubrics to grade, LLMs can use it. You want to probably start with a human, multiple iterations, to make sure that, rubric is really sound. And then you can all then scale the grading up with, like, an element of the judge. Three, ask it several different ways, and see if the answers line up. And then 4, make sure you just show the receipts.
The last one’s, like, very, low… level lift, and goes a wrong, long way around how much you can trust an answer. So, you know, show the actual SQL where the actual number came from. This is something we should all probably be doing today in our analysis, regardless of AI, right? Yeah, in terms of do models agree, we’re actually gonna do that, today, Attendee. We’re gonna go through reliability today. We’re gonna go just within the same, like, Opus 4.8 model, but we’ll have, like, a reliability check where you run. a few sub-agents on the same question. But… One thing we’ve been doing a lot of is, like, run the same analysis and run it against, like.
different models, like, within, like, the Opus family itself, 4.6, 4.8, 4.7, run it against, codex models, run it against open source models, and see if you can triangulate the same answer. If you can’t, then there’s something going on there, right? There could still be something going on there if they all do agree, if they all make the same mistake. You’d have to have that baked into your system, but if you’re looking at totally different models, It’s a little less, likely that that would happen. And then there’s a few gut checks you can run on any analysis. It’s like, did it even run? Did it use a defensible approach? Does it give the same answer if you ask again?
So maybe, I don’t know, like, if people are already kind of, like… so yeah, like, actually, Attendee… Attendee mentioned it today, like, he kind of had a… Close call with, some… wrong numbers, but he knew the business pretty well. Might have not been the exact known answer. Like, the known answer doesn’t have to be, like. X equals X, it can be, like, I know this space, and if I think about it, like. This number seems way off. But which ones of these do you already do? And you can add, if there’s other ones in here too, that’s fine too, but, like, known answer rubric, asking multiple ways. Showing the receipts. How are people, kind of, you know, even informally, right? Yeah, let’s see.
Good, I mean, good kind of spread against everything. A lot of people are doing four, that’s great. A lot of people are doing one. One, sometimes you don’t have. Four is a really good, low effort. Asked several ways. Nice. Yeah, one’s typical… yeah, anything novel. You haven’t checked before, like, you’re not gonna have one for it. We’re gonna do a, lightning lesson. next week, called Pressure Testing AI Analysis. I’ll drop the… or if Hi or Sravya want to drop the link in there, it’s another free workshop. We’re going to talk about, like, what you do when you… what you do when you don’t have Known answer. And then I’ll have, like, a QR code for that later on in the presentation, too.
Hai Guan: That’s the lesson next week, Shane.
Shane Butler: Yeah, yeah, I think it’s, like, next Wednesday, I think it’s the same time slot. It’s called, like, pressure testing AI analysis or something. Alright, so two caveats. So yeah, go ahead and sign up for that if you’re interested in… Figuring out, like, what you do when you don’t have ground truth. So two caveats here, because I don’t want to oversell all this stuff, right? The first one is about agreement. So, asking several ways and getting the same answer back is reassurance. It’s not proof the answer’s correct. So, if the system is wrong for the same reason every time, it’s just gonna agree with itself every time and tell you the wrong answer. It’s gonna, you know, tell you nothing.
So, actually, I find the more useful signal when, like, we’re building out Agentic analytics systems is… is actually disagreement, because that gives me a direction to, like. dig into and improve, the system, so it tells me, like, what I can fix. And we’ll get a bit into that today, in the demo. And then the second caveat is that, yeah, you don’t run all this on everything. like, match the effort of your evaluation to the stakes, just like you would do with anything in your job. So, if it’s a throwaway question, like, whatever, you know?
You don’t need to go super in-depth on on everything here for a throwaway, kind of, I’m curious question from someone important, we just have to give an answer. And then… Yeah, but never going to… like I said before, like, something that’s, like, you know, publicly reported, or you’re making a big decision off of, or it’s going to the board, like, yeah, you probably want to give, like, a more robust treatment to that. I’m gonna give you a sheet later on that’ll kind of, like, help you triangulate, like, based on, where you’re at with your company and your decision, and what’s available with your ground truth. Where… what… which one of these… how… how verbose… robust do you want to go?
like, a lot of folks that we’ve talked to are on leaner teams, right? Like, they don’t have the resources to go through full campus, so there’s a smaller version of this, like, doing the cheaper checks. That can make… You have a much more better understanding of how much you can trust this, but without having to, like, spend a ton of cycles on it. So, this is our, kind of, current map of what you might check across the pipeline. It’s not exhaustive. just kind of where our thinking is right now. But you can see, like we said before, there’s a lot of checks, whether the SQL’s right, whether it picked a defensible method, whether the number reconciles.
Today, we’re just going to focus on one of them, and that’s reliability. And then, you know, this is a thing we teach hands-on in the bootcamp, the intro bootcamp. You build your own Agentic Analyst, and you build these checks into it. There’s also a whole 201 course built purely on this validation, and then how to improve it, which turns out to be a lot of context engineering. We’ll get back to this later. I’m also gonna send you out email with some PDF worksheets, around the frameworks we talked about today. Okay, so let me… move over to Claude Code. The check we’re doing today is that simplest one.
reliability, and it’s the one you can run no matter how messy your setup is, because it doesn’t need an answer key at all. The whole idea is just ask the same question more than once. See what hold and see what moves. So let me switch over there. And can everyone see… Cloud code should be on my desktop here. If it’s hard for you to see, you can kind of zoom in. Via Zoom. Nice. Seems like… People can see it. Okay, so… alright, actually, before we start… Let’s do, kind of prediction. I think I got a slide on this. Alright, so… Same question, I’m gonna run the same question, I’m gonna run it 5 times. Do you think we’ll get the same answer all five times, yes or no.
I’m gonna ask a couple questions. The first one I’m gonna ask is around… This is… we’re gonna look at, like, a synthetic… e-commerce website, you can kind of think Amazon. I’m gonna ask it, how are we doing on checkout conversions? So, given that question, what’s your thoughts? You can drop in the chat. Lots of no’s. There’s a… there’s pretty low confidence in this system. Pretty low confidence in this system. Yeah, Attendee! Attendee’s like, hell yeah, Gentic Analytics is the future. Same answer. Okay, let’s try it out. Alright, so there’s actually… I’ve created a skill here. And… I don’t think that skill’s pushed yet, but… You are all in luck.
We haven’t actually formally released this yet, but we have two versions of our AI Analyst. One’s AI Analyst, one’s Plus, the plus one’s AI Analyst! open source hasn’t really been developed on since, like, March. And then we have a Plus version, which we’ve been constantly developing on that whole time. That’s actually… Available right now to the public. We’re just gonna move all that stuff into, like, a V3 or B2 of the, I guess, V2 of the AI Analyst repo, if you already have that on your machine, but maybe, Maybe, Hai or Sharabia, you can drop the AI Analyst Plus repo in here. This skill, the reliability one, I don’t think is pushed there.
We’re working… I have it brought in just for the purpose of this demo, but we have a separate repo that I’ll open source that’s gonna be around, like. All the different, like, agentic analytics eval skills. Okay, so… And then I’m using, VS Code here, Attendee. Alright. So, basically, what this skill does… It fires, you ask it a question, you tell it, like, How many, like… Cycles you want to run on that question. It fires that independent runs. It collects all that data and stores it. And then it, compares it. Calculates the mean, variance, etc. And then it does a little report that says, like, hey, is it stable? Is it drift? And how could we improve it? So, we’re going to… Run… what’d I say?
How are we doing on checkout conversion? I think I have this prompt somewhere. Here it is. Okay, so you’re saying, I’m gonna invoke the reliability skill, I’m gonna say, how are we doing on checkout conversion? Run it 5 times. So this should kick off 5 separate processes once it gets going. Yeah, so it spins up 5 independent sub-agents, Measures whether answers agree. Launches those concurrently. Eat those tokens. So you can see it’s all running here. And what it’s actually doing, we have our data in a… Snowflake Data Warehouse here. And so, each of these agents… It’s going to start… Hitting a snowflake. On their own, we actually have a… Something we’ve set up.
So we can always go back and see the work, and, like, audit the work that’s been done. We have a hook here, where anytime there’s, like, a tool call to the Snowflake MCP query. We have it record. They’ll work for us. And I think that’s down in the working. While this runs, I can show it. Oh man, my computer is… So low. Where is working? Just saw it. If I close this down. I’m gonna close some other tabs. Okay, so let’s return those. Oh yeah, here it is, query log. Come on, open up. Well, it gave me the answer here, but it seems lagging like crazy. I wanted to show you guys the query log. Well… You can scroll up here. It’s the wrong thing.
Alright, so it ran, 5 times, and… Man, I don’t know why my computer is going so slow, but you can see it actually came back with, the same answer for this one. every single time. So it came back with, 33.2%, And, You can see it actually lists the source here. And the reason it does this is because we have a… metric dictionary that tells Claude what we mean when we say conversion rate. So, if we go to… Okay, now we’re… now we’re back at, like, regular speed here. So if we go back… if we go into our knowledge base… and we go to Datasets, Nova Mart Metrics Index, there is, you know, this is pretty empty, right? Because… I just added this this morning.
But, there is the… I don’t have this… But, There is a clear numerous definition is. And so if we think about this, it’s actually how Hiding something here, so… This ran… The reason it didn’t have variance was because we had that in the dictionary, but it’s actually hiding, It’s hiding something in terms of, like, does that… is that the definition we actually… needed. So checkout conversion actually probably has, like, several actual readings behind it that could all be defensible, like… you can count from the moment someone starts checkout to when they pay. That’s the one that’s in our dictionary. You could also count every session that ever, ever landed.
You could count, cart to checkout, or checkout to payment step, or you can count people instead of sessions. Those are all, like, real definitions, and they’re all… they’d all end up being totally different numbers on this exact same data. So, our dictionaries, where it pins to one of them. Which is why it came back stable. But, if we were in, like, an actual meeting, like, that definition might not be what our stakeholders were were looking at, or wanted to look at. So sometimes I think it’s more helpful to see the divergence first, because Checkout conversion, in my head, could have meant something totally different than checkout conversion and say, hi, or Strivia’s head.
Reliability just told me the system is consistent with one of those definitions. But it didn’t necessarily say that, like, that was the right definition. So stable… stability is necessary, but it’s not totally sufficient. So we’re gonna do the same setup. with a different question now. And this other question… Is not in our… metric dictionary here. So we should get, variance across runs, I would think. We’re gonna ask… Let’s see… we’re gonna ask about retention rate. So… for unreliability… What’s our retention rate? And so it’s going through the same process again. It’s gonna kick off 5 sub-agents to run the same question. Here’s what I was trying to show you before.
So when it runs SQL, we have it record… that hook has it record the SQL it runs every time. So you can see… It has, like, the timestamp. This is just an ad hoc analysis. If it’s part of, like, a project I was working on, I would tag with that. And then it has the purpose. And then it has the actual, SQL that was ran. So you can see there’s a bunch here on checkout conversion, right? Because you have multiple sub-agents running it at once, hitting Snowflake. This is really nice for later on.
I mean, like, if I got a bunch of different, answers, or if I even got the same answer, one of the main things I would have, which is part of, like, doing, having this validation, is just, like, showing the receipts, right? And so, I could actually have the exact SQL pointed to it, or I could just talk to Claude, too, and so you can see it’s just populating more SQL runs as they come in now, but now for attention. But having that kind of, like. audit, trail is really nice for, just understanding how much you can trust it and know what is going on. You could hit, like, Ctrl-O or something and see what’s running in the background while it’s going to check it as well.
You could… Ask Claude, but, like, I do have concerns about it just, like, hallucinating. What I’ve seen is if I don’t have this recorded somewhere through a hook where it’s actually writing down the exact query it’s running, I could also have it go, like, look at Snowflake query history or something. then it takes a guess at what it did before, looking at its previous context. If you clear out the context, it’s not really gonna know. So, like, having, like, a solid audit trail. Alright, so all 5 runs returned. This time, they’re stattered. It’s gonna record, the results for me. in a, JSON with all the runs, and it’s gonna do the calculations.
And so, you can see this time, like last time, some of the sources were actually looking at a… metric definition catalog. And then this time, it basically just says, like, the source was… it made it up. on the fly. So… These two get pretty close, right? 99.3, 99.2… They’re, like, active in month and month plus one. This is active in month after sign-up. sign-up co-work rain. The first one’s 9.7%, 30-day repeat purchases, second completed order within 30 days. First, remember, we’re asking about retention. This definition, 34.9, repeat purchase rate. Purchasers with 2 plus orders divided by all purchases over the full year.
78% week 1 new users that have more than one session in the first 7 days after sign-up. And then we talked about these ones. These are, like, of the month one cohort, how many were active in the month after the sign-up month. This one’s pretty similar. But these are all pretty solid definitions of retention. Like, if you’ve been in a meeting where you’re trying to pick metrics. People… you could easily have, like. four people in the room kind of arguing these different ones, because they all tell… they all help you in different ways. conversion, even if I hadn’t had the metric definition for, like, cart conversion. That’s a little more cut and dry. I bet we would have got, like.
3 out of 5 matching, at least, but it’s not that much of a surprise to me, where it’s like, man, these are just all over the place. And, like, imagine how this shows up in a… a deck. Imagine if… if you have a leader that says, I want to know what retention… that just asks that, like, very, like, kind of vague, high-level question, but, like, very, like, natural. What’s our retention rate? And… Like, you could tell them 9.7% or 99.3%. So, the way to, to, like. this tells us, like, what am I going to do with this? Well, I’m probably going to go back to my team, I’m going to, like, define retention rate with them. This is actually kind of nice, like. gives me some options for retention.
Maybe they don’t want any of these. But it just tells me where to dig, right? And this is clearly a metric definition problem, and then… to… fix it, if I wanted everything to stabilize, I would get the agreed-on definition, and then I would just, add that to, like, some sort of context somewhere. There’s many ways you can manage context. One way could be something like this, where you have, like. a YAML file where you store all your metrics definitions. But it’s pretty crazy to see the spread. And, like, that would kill trust. If someone… if someone was thinking.
The, repeat purchase rate purchasers with 2 plus orders overall purchases, if a leader was thinking in that in their mind, of, like, what retention is, and then you told them 99.3%, they’d be like, what the hell are you talking about? I’m not… I don’t believe anything Shane says. Okay, so… This was just, like, one of those boxes, if we zoom back out. One check reliability, one box on a big map. So notice what it did and didn’t do. It caught that retention, wasn’t stable. That’s a big warning. It didn’t, on its own, tell me which reading was the right one. I would have to supply that definition. The other boxes on the map cover the rest of it, the receipts, the approach, ground truth.
where you happen to have it, if you have ground truth, and how hard you push on any of this comes from… comes back to, like, your use case and, what the stakes are. If you’re idly curious. One run’s probably pretty. If your number’s driving real decision, then you want to check this stuff, like, more robustly. I just… okay, oh yeah, before… we’ll go over to questions in a bit here. Let me grab you right, I promise. I did just… I posted this right before this session. I think I have it out here, actually. There’s this… I’ll send this in the email, too.
But, this kind of worksheet is, like, a little bit of a… Self-assessment, Helps you figure out which of your analyses actually need this kind of checking, what you’ve realistically got to check them against. you know, do this one first, it tells you where to spend your attention before you spend it. I’ll drop a link to… this post on it, if you want to download the PDF. there, but I’ll also send this, in… an email later today. And then I have another worksheet that I’ll share. Let me see if I can pull it up here… Yes. This one gets more into, kind of, like, the different ways.
We kind of went over these, right, but the different ways To check anything, Gut check-wise, and then… what you’re gonna do based on your, kind of, low-stakes to high-stakes effort. So I’ll share both of these in the email, but you can grab that other one from that link right now if you want to check that out in the meantime. And then we’ll also send a recording out. I think we got a couple of discounts codes, we can… we can show those in email, I think. I’ll go over, kind of, like, what’s coming up. Savia, feel free to drop the… those discounts in the chat, too, but a few things coming up, and then we’ll get into questions. I can stay a little over for questions.
I know we’ve got about 12 minutes left. If you actually want to build this. Here’s what… here’s where we go. Every one of these has a QR code right on the screen, so you can just pick up your phone right now and kind of Click which one you want. Top of the list, on the left, this one’s free. This is what I mentioned earlier. It’s the next lightning lesson, pressure test, any AI analysis on Wednesday, June 24th, 8 a.m. Pacific. That’s gonna tell you, how to tell if an analysis is actually right when there’s not really an answer key to check against. If you scan that, you can register for free right now, just like you did for this one. Then the courses, so we have a 101 bootcamp.
This is where you build your own Agentic Analytics system. That’s July 13th to 17th. We offer it again I think end of August. So if you don’t get this one in July, it’s gonna be another 6 or 7 weeks till we run it again. We just ran it this past weekend. And then 201 is the one built… we do go into… and we also do go into some eval validation stuff in a 101, since it is kind of, like, table-stakes stuff now. And then 201 is built purely on what we did today, like, validation and context. You also go into multimodals, like open source and codecs. That’s gonna be in 2 weeks, or a week and a half? June 27th to July. Fourth, That’s… And then there’s a 5-week one, AI analytics for everyone.
This is, like, the pure… this one is, like, the thinking itself, so that runs right now from June 15th to July 19th. We’ll run again August. That’s about the whole, like, we… we kind of say it’s, like, a mini master’s in data science, but the execution layer is in, quad code. Rather than in SQL and Python R directly. Whether or not you do any of these, you’re gonna… I’ll send you those worksheets and the recording, later today. We’ll get into questions in a second here. One note on the… one promo we are running right now… The 5-week course, AI Analytics for Everyone, the one about the whole analytical thinking, and the judgment behind it, and how to execute, in Cloud Code.
We actually just kicked this off this packed week, so we’re on day 3. It’s a mix of async and live material, so if you want to join this one. You basically have just missed two office hours at this point, and they’re both recorded, so you’re… we’re leaving that registration open until the end of the week. easily can catch up. It’s self-paced, so you take out your own time, and you have… you can have the content for forever. It’s not like you’re… at 5 weeks, it goes away. And then the part that we think is really… makes it worthwhile is if you register for the 5-week course, you add a seat for free, or not free, 10 bucks. in our 101 bootcamp, which is normally $900.
So, you get that thinking course, and the… builder course basically together for the price of the 5-week course. Yeah, you can scan the QR code on the screen. Or Shravia. Or I just dropped a… I just dropped a link in the chat. You can DM us or email us about it, too, if you have got questions, but basically. sign up for AIMs for everyone, and then just DM me, and I’ll give you a code to get into the $900 bootcamp for 10 bucks. Okay, let’s get into questions. Just a few. I’ll run through a few I often get, and then we’ll open it up. I see there’s already some coming into chat. Most common one by far is, like. you know, what if I don’t have a big data team, or my data’s a mess, can I do this?
So, yes. This is actually who it helps a lot. You can take the leaner path to evals. You can lean on the cheaper checks, like reliability, and making sure it’s receipts. You can save that heavy stuff for the handful of numbers that really, really have, like, high stakes behind them. you don’t need a super clean, like, data warehouse to run the same question twice and notice if it disagreed with itself, right? So, there are leaner checks that you can incorporate right now, today, that don’t require a bunch of technical skill or rigor, or really strong data foundations. They’re going to flag, probably, that you don’t have strong data foundations, but that’s great. That’s, like, a good step one.
Next one. I get a lot is, like, how do you get ground truth when you don’t have, the answer key. So, you most certainly, almost certainly have ground truth. You’ve just never labeled it as such. So think about, like, finance, reconciled numbers, queries that you’ve, written, like, dozens or hundreds of times, or someone else has, It’s, like, a nice head start, and it’s a real place to anchor to, so there’s probably ground truth, sitting in the context of your business that’s just not labeled as ground truth, but you can You can kind of, like, Reapply it as ground truth, to test against.
I had mentioned a little bit about this in terms of, like, talking about the sub-agents, but, We also talked about how we do, like, multi-model kind of triangulation. So if multiple models agree, is it right? This is the one you gotta be careful about. So, agreement, like I said, is reassurance. It’s not proof that the member is correct. So, say if models are trained on similar data. They’re gonna share the same blind spots, so they can still be confidently wrong together, Like I said, the more useful signal I find is when they disagree. Because that tells me where to dig and gives me, a lot more context about what’s going on behind the scenes. And then, if it’s stable.
If it gives you the same answer every time, is it right? Kind of similar to the last one, no. A wrong query can be perfectly stable. It’ll hand you the same wrong number all day if you execute that query over and over again. This, you know, just doesn’t even do with AI analytics, that’s just, like, people running queries, but I’ve written plenty of wrong queries. So stable is necessary, but it’s not sufficient. So it tells you the system’s settled, not that it’s settled on the truth. Okay, let’s… get to some questions. I haven’t been looking at these, but hi and Sravya, have you any questions that popped up here? Goodbye.
Hai Guan: Yeah, there’s a…
Shane Butler: We’ve got 10 or 15 minutes after that.
Hai Guan: Yeah, I think that’s… good for questions. Attendee has a question around, two questions. Is there a skill in the repo to help work through metrics definition? And, the second one is, how does this dovetail? with choosing North Star Metrics and the session on Friday. I think that was last Friday. I know that session was more on reasoning, if the metrics chose past a certain rule break, but would it help when the definition of the metric was not completely agreed on within the team?
Shane Butler: Yeah, so we have… I… so, those really good questions. The North Star metrics is probably… we have a few skills around metrics, but I’d say the North Star metric stuff. Which I’m actually pulling out into its own… repo right now that I’ll hopefully get up in the next week or so, but it’s in that plus one, if you all have found that. I think that’s probably our more robust skill around, because it does have multi-modes around auditing metrics, explaining what good metrics are, going through the checks of a strong metric, and then you can have it, it will, like, apply those to your metric definition YAML file in there within the workflow. In terms of, like.
Yeah, we haven’t created anything, but we… maybe we should do this, actually. This would be pretty interesting. We haven’t created anything in terms of, like, hey, how do you… leverage a system here to, like, get alignment. across your team. The North Star metric one, I think, would be really good, because you could have people, like, come up with ideas, and then you could audit them through that system, and it will give you recommendations and show you the weaknesses, and it has that huge knowledge base that cites, like. actual case studies, and it cites a lot of that, like, Amplitude Northstar Metric Framework Playbook, that’s free online. So I think it’s a nice resource to have.
when you’re having that discussion with your team, because it will tell you… you can use, like, the… I think it was, like. the North Star metric skill, and then, like, dash explain, and have a concept. And, you know, that’s… that’s, like, basically, you’re having, like, an expert on this in the room with you, like, talking through your team, but I think we could totally create something where it’s, like. Hey, record this team meeting with everyone talking about the metrics and kind of live give some input, or trying to align, or even, like, rephrase as, like, hey, High is thinking about this metric in this way, Shane’s thinking about it in this way. Where is the overlap?
Why are they diverging? I think we could create something around there. But there’s nothing… Right now, in there, that does that. Really good question, and good idea. You can make it two Attendee, you can make that. You could… you could extend the North Star metric thing to be the teen thing.
Attendee: Thanks for that, I might, yeah, I might give it a shot. And that framing of, like, taking in multiple definitions and running them on the North Star metric, I think is… I think gets that same sort of… Level setting for everyone. Great answer, thank you.
Shane Butler: No, no worries, no worries.
Hai Guan: Alright, let’s see, next question announced at, some companies don’t allow providing full database access to the AI, but allow to use it without giving it any company context. What are your thoughts on organizations like banks and fintech that restrict usage of AI due to data security? Cool session, by the way.
Shane Butler: Just a matter of time. they’ll either rise to the top or burn to the ground. No, I’m just kidding. Yeah, I think there’s, Listen, like, yeah, a big… there’s a couple things you could, do to improve your system. There’s one is, like, building within the system itself, like, agents. skills, how things are structured. The other is frame your input in a more robust way. Another is adding context. I’ve found by far, and especially as the models just progress on their own, in terms of capability, context is the thing that really, improves Them. Something like a metric definition, though. I mean, I don’t think you’re, like, giving any company secrets away. Something about, go access all of my data.
And… store those results somewhere is a different thing, so I feel like you know, kind of some of, like, the… just, like, metric deficits in itself, I don’t… I don’t see that as, like. Kind of, going against, like, those, like, security restrictions around. context. Something that we’ve done at our company, or my company that I worked at before, is anonymized a bunch of data. So we could work with it more reliably that way, so we built a whole system that anonymizes data. And this is pretty highly sensitive data. It was, like, legal, AI company, so you’re… it’s, like, contracts and stuff. And then there are, there are ways to do this more secure.
Like, we also use, when we run things through cloud code, when we do non-anonymized data, we run it in, like, a safe, like, hosted instance, like in SageMaker. We run it with Amazon Bedrock models. We run it where we have, like, zero data retention policies. So I think, you know, obviously, this is a huge question that we get a lot of times, like, how do I deal with security? And a lot of the answer’s gonna be, like. just like with any other piece of software, that’s gonna be, like, legal, security, IT team thing, too. work with.
We can talk through, like, some of the stuff we’ve seen successful before, but, you know, every… every… every AI company, every Frontier Lab is, like, trying to come up with solutions to make this as secure as possible, too, because that’s a huge friction point for them.
Hai Guan: Yeah, and then possible another option, and probably where things are headed, is open source models. Like, you host your own open source locally. So you can even run the AI without internet, and it certainly doesn’t go anywhere. Like, your data doesn’t go anywhere, it lives in your machine, stuff like that. You’ll have to have, or your company will have to have a really powerful machine if you wanted to do the max capability on the analytics side. But that’s… I think that’s a viable option as well, as people are waiting for security clearance and stuff like that.
Shane Butler: Yeah, and we actually do that in our week-long 201 bootcamp, where we talk about Validation, context engineering, multi-models, wellness models, open source models. And I’m pretty bullish on the open source model stuff. I know Hyde’s been doing a lot of, research into it, but I mean, if you think about… I just look at, like, the model evolutions that’s happened recently. Personally, when you have a harness, like, when you build a Gentex system, analytic system, to do analytics around a model. between Opus 4.6, 4.7, 4.8, and Fable, 4.6… does the just as good… actually better than 4.7 and 4.8. when it has a system around it.
If you have no system around it, and you’re just, like, opening up a terminal, and there’s no repo or anything, and you’re asking analytics questions, like, yeah, Fable blew those others out of the water, but it was just, like, totally at parity when you put a system around it. So at some point, like. those open source models will be at the level of Opus 4.6, and… eventually they’ll be cheap, or they’ll be, like, smaller, and be able to be at that level. It’s just a matter of time, so…
Hai Guan: I mean, DeepSeek V4 Pro, the one that I’ve been testing quite a bit, seems pretty promising, but we’ll see.
Shane Butler: Yeah, we’ll do some free session on open source stuff at some time soon, maybe in the next month or two, but… If you really want to get into it, join our 201 course.
Hai Guan: Alright, probably the last question here, let me see… semantic layer explaining definitions, data, metrics take so, so long. Any good advice on how to cut corners responsibly? I can take a crack at this one.
Shane Butler: Yeah.
Hai Guan: Yeah, I mean, Attendee, semantic layers is, so, here’s how I think about it. It is a very powerful way to govern your enterprise-level data across the company. Certainly very painful, because… and the pain is actually pretty interesting. It’s not the execution layer anymore. Like, you can build systems where writing the pipeline building the data models and, actually crafting the… whatever, warehouse you have, or whatever, ETL tools you have, like, that’s not the hard part. That part is pretty simple. The really long, time-consuming, and people… Avoid part is, like, the people and process up front, like.
Agreeing on the definition, agreeing on the specifics, agreeing on the filtering, agreeing on what each component actually means, because they’re responsible for explaining the metric themselves and owning it. The… the thing that I find at my company is that that is the hard part where people are like, oh yeah, just, you know, define something, but then at the end, it’s not actually that simple, because you know, like, we can define something, but if they don’t agree later on, and you’ve built everything, and reporting, and whatever, and, like, they come back to you and be like, hey, I never agreed to this, what’s going on?
That itself is actually worse than doing the upfront work of, like, you know, hey. this is what we agreed on, this is everything that’s aligned, and then everything else is, like, just the click of a button. I mean, I’m exaggerating here, but that’s sort of, like, the thing. And so the top part. just cannot be avoided. It’s just people things.
Shane Butler: Yeah, if you want people to align on metrics, just make metrics that go up and to the right, is what I find. People like those ones, but when they’re going down, people are saying, like, this isn’t the definition I had in mind. Just joking, but… Okay, I think we’re… I have a couple other questions on here. You guys can kind of review the slides afterwards if you want to go… I guess the last one I would leave with here… Is… won’t models get good enough, and we won’t need any of this eval stuff? I think it’s actually the opposite. I just… based on how I’ve seen the models progress, like. the more of the work we hand off, I think the more we need to check it, not less.
So, like, reliability, ground truth, whether it was even the right call to make. None of that goes… away as the models get better. I think they’re gonna get a lot better at being right, but I think they’re also gonna get a lot better at being super convincingly wrong, too. So I would, if you’re thinking about, like, where to invest your time right now in Agent Tech Analytics, I think… Getting some form of evals and understanding that the trustworthiness is a really… High leverage way, use your time. Cool, folks, I’ll send out that email. Thanks, everyone, for joining. Thanks for the engagement, and… All the questions. We’ll see you, hopefully, next Wednesday for that next lightning lesson.
Right.
Hai Guan: Everybody?