Shane Butler: Hello, everyone! Welcome, welcome.
Hai Guan: Ayy.
Shane Butler: Happy Friday!
Hai Guan: How’s it going, everybody?
Shane Butler: Let’s do this as we get started. I’m gonna share my screen. And then… if you can share… see my screen… Where… which screen am I sharing? There it is. If you can see my screen, drop… drop where you’re located in the chat. Let’s do that. Drop where you’re commenting from. I’m in… South Lake Tahoe. Oh, man, France, New York, Greece…
Hai Guan: Louise.
Shane Butler: UK, Brazil? Okay, we haven’t repeated a single country yet. Uruguay? I’ve got another San Francisco, London, Canada, California.
Hai Guan: UA.
Shane Butler: Denver. Duh.
Hai Guan: Copenhagen. Love Copenhagen.
Shane Butler: I’m just gonna get a couple things set up here. Turn this waiting room thing off so everyone can just come in. Okay. Nice, we’ll get… we’ll get warmed up here with just some quick intros. Oh, hey, Shravia’s… hey, Shravia, how’s it going? Shravia’s in the house. make you a co-host also, Soravia. Let’s just get started with some intros while people roll in. And then we’ll go right into the… the, presentation here. So, I’m Shane. One of the co-founders of AI Analyst Lab. Yeah, been working in data science for a little over a decade. Today is actually my last… day at my current job, because I’m going to go into doing an AI analyst lab full-time starting Monday, so that’s pretty fun.
Yeah, mostly in data science for the past 10 years, and the past 2 years, been in more of the AI evals, agentic analytics kind of space, rather than, like, kind of classic product data science. Yeah, hi, you wanna intro yourself?
Hai Guan: Yeah, hey everyone, my name is Hai. I work at a legal tech company. I lead the data team there. I have roughly 20 years of experience in the data science and analytics space, one of the co-founders of AI Analyst Lab, and yeah, excited to… excited to be here today with you all. Shopify one go?
Sravya Madipalli: Yeah, hello everyone, I’m Stravia. Pretty excited to have this session get started with you all, that Shane’s running it. I am co-founder with Hi and Shane for AI Analyst Lab. I have around 15 years of experience in data science, starting with Microsoft, and most recently at Superhuman. Yep, let’s get started. This is, one of the most awaited workshops, I could say.
Shane Butler: Alright, I gotta stop sharing and reshare for a second. Just gotta get my… My screen’s in order. But, yeah, today what we’re gonna talk about is… Agentic experimentation, specifically experiment… agentic experimentation, design and analysis. But, we’ll go through a little bit around, We’ll kind of talk a bit about, like. why this is important. I’ll give you a quick preview of, like, kind of, like, the system that we’ve been playing around with outputs. We’re gonna share a repo with you, maybe towards the end, or while we’re doing it. We’ll definitely send it in the email. It’s an open source repo.
It’s just something we’ve been kind of, like, experimenting with around, agentic experimentation. I’m gonna say experiment a lot today, probably, so this isn’t, like, some, like. totally a production-ready thing, but it is, I think, kind of a way we’ve been learning, and we’re just trying to share stuff as fast as possible as we play with it, and I do think there’s a lot of opportunity in this space. We’ll talk about, like, Where we think, like. AI can, like, help when it comes to experimentation, where It’s probably not gonna help yet, or the human release doesn’t evolve. We’re gonna go through a little tour of, like, how we’re kind of building things out a little bit. We’ll do a quick demo.
The demo to actually run live… I’m gonna run it live, but it’s, like, it kind of takes a while to do all these tests. So… We’ll see how far we get. We might not get to the analysis stage, but we’ll probably get through design, at least. So we’ll get into cloud code to show you that, and then, yeah, obviously we’ll share the repo out with anyone, so you can go try it at your job. And probably make it way better than we have, so… Kind of start with a… some of the challenges of experimentation, like, you know, I’ve worked with a lot of teams We ran experiments, worked with. You know, teams that have, like, you know, like, 40 million weekly active users.
I’ve worked with teams that have, like, a few hundred weekly active users, in terms of, you know, different products, and… The teams I’ve worked with that run tests, well, badly, I would say, it’s not really for lack of, like, tooling, or because they’re not smart enough, or rigorous enough. I think a lot of the time, I see tests Being ran badly is because the human who’s reading the result is the same human who wants a particular result. So, like.
We ship a feature we believe in, we watch our whatever dashboard or notebook, wherever we have this, our… our results, populating, and, you know, a lot of times it’s very easy, if the number looks good, to start building a story around why it’s good. It’s really easy for us to kind of, like, move our threshold, our definition of good, also, depending on where the result landed. It’s really easy to peak early and stop when we’re winning. it’s really easy to explain away, like, guardrails that kind of broke through the thresholds we had tried to predefine. I don’t think it’s, like, kind of like a dishonesty thing.
I think it’s just, like, how our brains work when we see a number on a screen that we care about on a screen. So I’ll show you a little bit before… we’ll get back into that, but just to show you a little bit about, kind of, like, the output of, like. that we’re kind of generating with the system, just so you can kind of get a preview, because there’s going to be a bit of slides here as we get through… go through kind of, like. what it does. So, the output that kind of generates is, like, three parts.
It’s, like, your experiment design brief, so everything before you do… see any sort of data, run any sort of test, And so you can see, you know, it has your… Your whole, your hypothesis, why you think it’s gonna work, the metrics you’re gonna use, the… The… the effects that you can detect given your sample size, what the exposure’s gonna look like, And then it has some stuff around here where we have, like, snapshots and dates and ways of tracking this over time so we can go back and audit it. The other… there’s 3 core documents. The other core document is our… SRM brief, so our, our, our… sample ratio mismatch.
Basically, this is something you want to do, again, before you look at any results that says, like, hey. is the treatment in control, like, evenly, kind of, distributed? And then the last one is going to be your, actual analysis brief here. So, just to give you an idea of kind of, like, what the output looks like, and, what’s in there. So, in this case. The analysis brief here, you know, has, like, your verdict. Your lift, your confidence interval, Every claim cites, like, what file and step it came from. So we’ll watch… we’ll go through how this is all built in an Agentix system, and then we will see how far we get running through the demo. And then we’ll have some time for questions.
Okay, but that’s just kind of the output, so you don’t have to wait to the end to see it. And then Shravia and hi, if there’s any questions that come in the chat. Feel free to… to let me know. I don’t know where I put my chat. Where the heck is it? Oh, there it is. Okay. So… One of the things, like, are kind of, like… thought process around this, too, isn’t that, like. like, genetic systems aren’t solving the math for us. Like, the math is not really that hard. People know how to run A t-test, there’s… there’s Python functions and libraries for all this stuff.
There’s, like, like, dashboards work, but a lot of the time, like, experiments, like, the decision that is made based off the results of the experiments still come… can tend to come out from, like, whoever the loudest person in the room Actually is, and what they want to see out of it. So that’s not, like, a stats problem. That’s a, that’s a who is reading the… result problem, and that’s kind of what we’re trying to solve here. So there’s four kind of things that we see that, and I’m sure you’ve all seen it before in your work, that can really wreck, an experiment. One is, when folks kind of move their decision threshold. To wherever the result landed.
So, you know, we run a test, and then after you after you’ve seen it, like, maybe… maybe we wanted a 2% lift, that’s what we set our minimum detectable effect at. That was our goal from the beginning, but then we see at the end that it’s, like, a 1.6%, and now suddenly We, in our minds, have changed our decision threshold to, like, a 1.5% being the bar to moving forward. It’s really easy to do that. Two… You guys, you all know about this. Peaking early. Stopping when we’re winning, right? So, once we see a lift, it’s really hard to unsee it. It’s really hard to unsee it. It’s really hard to unsee it when your stakeholders have seen it, too.
So, it’s also a lot easier when you’re under pressure to roll stuff out to, like, just roll it out at that lift and tell yourself, hey, like, we’re being… we’re being… Fast or being efficient. Number 3, explaining away guardrails that broke, like, you can explain away anything, like, everything has a story, like, you can spin a story, easy. So, like. explain the guardrails away based on something else, or even… even kind of, like, saying we’ll deal with that later, right? Like, oh, latency went up, but the primary metric moved in the right direction, well, we’ll fix latency in the next sprint, and then, yeah, that never really happens. Or maybe it does.
And then the last one we hear is, like, changing the metric that’s primary. Actually, I’ve seen this a lot. So, like, the primary… the one we said that’s primary is, like, flat. Maybe it’s not negative, it’s flat, but then we see some other, like, secondary metrics that go up. And suddenly, we’re like, oh, these ones look great, those are our primary metrics after the fact. This is really dangerous, of course, when you have, like, dashboards that show every single metric that anyone can see. So every one of these, like, there’s some story being told after the data arrived, and that story being told rewrites the question we were originally asking. And I don’t think it’s, like.
lying or dishonesty, really, I think it’s just, like, what brains do, especially when you put a bunch of hard work into doing something. So, here’s where I think… This stuff actually breaks, so there’s kind of, like. I was, like, 3 kind of stages at a very, very high level. of, experiments, running experiments, right? There’s, like, the design stage, so you write a hypothesis, you pick a metric, you set a threshold, you decide the guardrails, If you’re, like, loosey-goosey with this stuff, you’re basically just, like, sandbagging or hedging, then it’s like you’ve already made up your decision around what it’s gonna be if you’re not strict at this stage.
Instrumentation, that’s, like, more like, you know, like, the engineering work, so, like, the code that assigns variance, that makes sure that events are firing, the days the test, is actually running. It’s pretty straightforward engineering work. It’s high consequence, though, obviously, if you mess it up, but it’s pretty procedural. And then analysis. That’s reading the data, computing the lift, deciding to ship, or… not ship. Writing the readout. And this is where those four habits in the last slide tend to show up, right? After the fact, once you’ve seen the data, where that bias comes in.
So, I think where humans, like, kind of get squeezed, and where humans, Is, is obviously that third one. That’s where I think the bias kind of comes in. The middle one is kind of, like, just engineering. Today, I’m going to talk about the system that we’ve been building that puts guardrails on… more on stages 1 and 3. So, we’re deliberately not touching The middle, like, if we have some agentic system that messes instrumentation up, your data is completely wrong, and no analytical rigor is going to save that.
That’s just, like, not really, like, a risk profile we’re taking in right now, but I think in the experiment design stage, in the experiment analysis stage, There’s, like, actually a lot of guardrails we can apply there. So… Like, platforms that lots of… experimentation platforms exist today, they compute, statistics correctly. They’re gonna give you, like, correct lifts, correct confidence intervals, but they’re gonna leave that judgment entirely to you, or to your team, at exactly the moment that your judgment is most compromised.
You are seeing a result after you spend a bunch of time working on something, and you really want a certain result, and they’re just going to, like, say, like, here’s the numbers, make a… make the call. It’s, like. There’s just so much opportunity right there, too, like… persuade yourself, even subconsciously, into making the wrong call. So those are… those… all these platforms are more like calculators, right? A calculator can’t stop you from asking it the wrong question, though, or, like, tweaking your question a bunch of times to… to get some result. So… the job they take is kind of like the math job.
They… they don’t take, like, this human-in-the-room job, and then that’s kind of where I think, you know, experiments go sideways. And the ways that humans fail experiments is, like, a way LLMs can fail experiments as well, actually. So, both pattern match towards results that we hope for. So, both given access to our hopes and say, like, some, like, borderline number, are gonna find a path to our hopes. that’s just, like, how it’s… how our minds probably work, like… and so in this system, what we’re trying to do, then, is only let each piece do what it’s good at, and then, like. add very strict, kind of, like, separation of jobs, and then lay… of jobs even within each of these layers, too.
So, like, code. Python in this case. their job is to decide, it’s deterministic, so the same input should get the same answer every time. The verdict comes from a function that can’t kind of, like, flinch when the borderline number is in front of you. That’s a thing humans and LMs are going to be bad at when something’s, like, right on the border. So we’re gonna cut them out, and we’re gonna make that a deterministic decision step with code. LLMs, they’re gonna translate, they’re going to hold the conversation.
They ask the questions that pin the experiment down, they’re gonna draft the brief, they’re gonna explain the guardrail, write out the, kind of, like, readout, things that models are really good at. they would flinch if they knew all the context through everything, so they’re not going to be able to see all that context, and they’re also not going to make the call itself. They’re going to read the call from the code, and then do some kind of interpretation around that. And humans judge, so that’s, like, all the work that actually required judgment.
So, like, the design review meeting, where someone catches that we’re testing the wrong thing, the stakeholder conversation about whether this is the right bat, a bet, The team discussions around what metrics we could use, if we’re taking, like. Big enough swings, like, kind of, like, prioritization decisions, these are all things where our brains are definitely the right tools. So, each piece, kind of, like, we want it to live where it can’t, like. flinch or be biased. So what this actually is, we can kind of explain it in, a few pieces here. First, it’s a set of agents, so there’s, like, just, like, 5 agents in here. One is the understander.
It drafts metric definitions, with you as kind of, like, a co-pilot, or you can do that yourself. One is the designer, so that drafts the brief, like, the experiment brief, design brief, and a data plan. One is the critic, so it critiques everything that gets, committed. It doesn’t only critique, like, hey, are we including everything in this, like, design brief, or everything in this analysis brief, but it critiques the way we got there. Then there’s one that does SQL, another one that does, like, an analyst-narrator, so that works around, like, the pros of the readouts. Seconds are skills. The second one here is skills and cloud codes. There’s really two primary skill design. And analyze.
So, they’re those first and third stage of that experiment process. And then there’s kind of, like, three size scales in here that you can check out. One’s, like, to look through, like, the audit trail. Every… basically every single decision that gets made is, cataloged, so later on, someone can audit through and see how, The experiment came to that decision, and there’s one for hooking up data, and another one for, For, running the readouts. And then third are just deterministic stats and, kind of decision tree packages. So this is just Python in terms of, like, hey, on how to run your power analysis, how to run your t-test.
how do you, like, walk through and, like, check all of the, kind of, like, decisions of whether or not you’re going to pass this test. So agents draft, skills, route, and Python decides. I think we can get more into this when I just open up the repo. It’ll be easier than explaining it here, but there’s kind of two layers. within this that are, like, the kind of, like, I’d say knowledge base, output storage. One’s, like, the project layer, so these are definitions that don’t change between experiments. So, like, what a user is, what a session is, what conversion means in your warehouse. Those live in semantic models and metrics. at the project root. They’re just YAML files.
You can read them, you can edit them, you can… you can make them yourself. They’re not a database, they’re just files, but they’re things that just exist at the entire, kind of like, project org level. And then there’s the experiment layer. So, this is the… the, like, kind of formalization of a single test. So your hypothesis, your primary metric, your decision rule, etc, etc. And so each experiment has its own directory with, all of its kind of, like, information that’s going through. The agent reasons through. the second, using the first. So when you say something like, I want to test moving the Buy Now button.
Say you’re working on, like, looking at, like, an e-commerce site, the designer agent’s gonna look at Your metrics directory, that’s a catalog in your… in the root. And it’s going to pick conversion rate as a candidate primary metric, and it’ll ask you to approve that, because it matches the word you used. It doesn’t invent a metric, it picks from what’s already defined there, and if you need a metric that doesn’t exist, you can work with an understander agent to kind of create that metric. So we’ll watch this in a demo in a moment. I do want to kind of burn through these slides, 30 minutes in. I’ll try and get to the demo in the next, maybe 5 minutes. the way I think about it is, like.
like, experiment… this, like, agentic experimentation platform is like a pipeline with a conscience, so, like, the pipeline part is very ordinary, like, ingest data, check it, design a test, collect, analyze, interpret, write it up, and then the conscience part’s like, at fixed points within that pipeline, stop yourself and refuse to continue unless certain conditions hold. So, kind of trying to remove that human bias that we talked about. Like, a critic agent’s gonna look at the drafted brief and flag problems, like loose metrics, ambiguous thresholds, or populations that wouldn’t give us enough power. So, what does a designer do?
When you start a design, this is kind of like… the key… command I think we’ll probably go through today, so design. You start designing your conversation and call it code, the system’s gonna load your warehouse, Doesn’t run any queries yet that’ll reveal what’s already, in the data. It’s gonna profile, what tables exist, what populations are addressable, what events fire, then it’s gonna elicit a question from you. Like, the agent will ask the questions that pin down what experiments you’re actually trying to run. Like, is this the… like, what counts as a guardrail? What’s the… what’s the threshold to ship? It’s gonna draft a brief.
That brief has a hypothesis, a primary metric, guardrails, explicit splits, a duration and a threshold. A different agent is then going to critique that brief. It doesn’t know what you hoped the result would be, like the original designer agent does, so it just checks whether the brief is well-formed, and whether the metric definitions hold up, so you’re trying to, like, remove certain contacts from certain agents. If it passes, it’s gonna actually seal that brief. And then that brief won’t be looked at again until the analyze step. So, the whole design conversation happens before any outcome data gets touched.
On the analysis side, effectively, you’re gonna run… it’s gonna run through, 8 steps. And they’re very… ordered in a very specific way. you have to pass one to get to the next, and to eventually get to your ship or learn or don’t ship, decision. So step one, does assignment split match what the brief said, so, like, it is a 50-50, actually, or was it 57-43, if that invalid, if it’s an invalid SRM, so it doesn’t match, then, like, no lift’s even gonna get computed. It’s not even gonna calculate the lift For you, because you’re gonna cheat if it looks like it’s up, and the experiment is already faulty, so it won’t even continue to the next step. Step two, did any guardrail break?
Again, this is the second step. This is before any lift has been calculated. If your guardrails that were predefined in the design step with very specific thresholds have broke through, it’s not even going to calculate the lift again, because… on the primary metric. Because you already said we’re… you’re not going to move forward. If those guardrails get passed. So it’s trying to remove the opportunity for you to bias yourself. Like, of course, with anything, like, you can… force things through, right? You can go analyze it somewhere else. But it’s going to add friction at every point where there’s opportunity for you to have biased thinking as a human. So, step 3, is the sample big enough?
Again, if it’s not. Then, it’s gonna halt you here. And then only… then and only then, after it passes those… those three, so SRM, Guardrail, and the sample size, is it going to look at your, Primary lift. If there’s no lift at all, it’s not gonna stop. If it… if there is lift, it’s going to see if it moved in the, the predicted direction. And it’s also going to look at the, magnitude of the lift, and if it’s even worth If it’s even worth Pushing out, or if it’s, or if it’s so small that it doesn’t even matter. And then step seven, did it hold across the effect window? So this is, like, novelty effect.
Or are you seeing that on the last, you know, one-third of the exposure period compared to the first two-thirds, there’s, like, more than a I don’t know, 25 or 30% drop. Those are thresholds you can set, within here, in parameters. And, And so, if it’s… if it has knowledge effect, again, it’s gonna… it’s gonna halt. And then based on everything that comes through here, it’s gonna tell you whether you should ship, or not ship, or learn, or whatever. Okay, I feel like we should get to… the demo… this is just kind of saying what I said earlier. Every… you know, like. we think we should give LMs as much context as possible, as much data as context as possible, to, like, make better decisions.
We’re actually going to take a kind of opposite approach here. We’re deliberately giving the judgment agents less context, because the failure mode we’re defending against isn’t, ignorance, it’s motivated reasoning. So we actually want to remove context, as much as possible. Basically create, like, blind judges. So imagine you’re on, like, a team, and you ran your experiment, and then you just had someone totally, with no context on another team, go read all the numbers and tell you whether you’re gonna do it or not. So just a way to remove the bias.
Okay… The output… I showed you a little bit in the beginning, we’ll show it to you again, but it has those three briefs, the design, SRM, analysis, and then it also outputs for different audiences. So you’ll have, like, a very high-level output for, like, an exec. to something that’s, like, extremely detailed with, like, the SQL Esri on for, say, like, an engineer on the team that needs to debug. Okay, let’s go into Claude Code. Because that’s more the fun part. I don’t know if there’s any questions you guys want to answer while I’m switching screens here. Or.
Sravya Madipalli: There’s a couple of questions. There’s one question, I think two questions around Bayesian A-B testing, about people moving towards Bayesian and, like, you know, and how do we see agentic implementation there. So, I think there are a couple of questions on that.
Shane Butler: Yeah, there’s some… there’s some Python, helpers on here for all different types of tests, if you want to run Bayesian approach, you can? I… I haven’t done much of that, to be honest. So, that’s not really what about… this isn’t really a… a, talk about, like, are we gonna move to Bayesian experimentation? It’s more about, like, how are… how do agentic systems, remove bias and kind of speed up the experiment design analysis process. But there is, like. when you look at the repo, you can take a look at some of the Bayesian functions in there.
Sravya Madipalli: Also, there’s a question about, do we have an experimentation course, specifically? And that’s great. Attendee, I think, shared Ron Kohavi’s book, which is great, but we also have AI Analytics for Builders, where we do not fully share about experimentation, but we have almost two solid weeks of content on causal analysis. And like, we’ll share probably more about it later in the workshop as well, but yeah.
Shane Butler: Yeah, we go pretty deep. It’s not like… that it’s not like an end-to-end, you’re gonna be the extreme expert in everything experimentation. It’s more like, here is the kind of, you know. 80% to 90% of how people run experiments in tech today, and then how do you execute it through, Claude Code and Agentic systems? I don’t know if any of the other courses on Maven right now have, here’s how you do it with Agentic Systems, or if they’re more, like. classical. There are some really good courses if you’re looking for foundations on there, but we do spend, yeah, a full week or two, just on that. I think we’ll probably add something, or at least some free content.
Yeah, I mean… I’m trying to kind of burn through this whole thing, but, like. when I was trying to… to talk… think about what I wanted to talk to, I could easily do, like, an 8-hour… walkthrough of this Agent XP thing we’ve built. Okay, so… we’re in Cloud Code here now, So let me just walk through a couple folders, and then we’ll run some stuff. So, two folders, those ones I told you about before, metrics. So there’s, like, 14 metrics I have… I’ve predefined in here for this project. You know, so, like, conversion rate, add to cart rate, revenue per user, page load time, all things you’d actually track, say, like, at an e-commerce website is what I’m thinking here.
Each one is just a YAML file, And on the other side, there’s semantic models, user, session, order, page event, assignment. These are those project-level things that I told you. They don’t change between experiments. They’re kind of like the vocabulary of your org when it comes to metrics, right? So, agents pick from these. they don’t invent them. So we can take a look at one, like, revenue per user, maybe? So you can see here, like, the type, it goes through, like, what’s the type of metric is, it’s typically a mean. That drives which stats functions gets called, right? So if it’s, like, a proportion.
versus a mean, if you’re going to run different types of stat tests, so you want to have that. Around your metric definitions, the numerator is the sum of order value. per… User, so, that gives it some signal as to where it’s gonna pull from semantic models. Denominator is distinct assigned users. So that’s all in… in plain text, right? If we scroll down here a little bit, you’ll see, common roles, Guardrail? So, and then it also gives you this, like, guardrail default number of minus 1%. So that means any experiment that uses this metric automatically inherits a guardrail that catches a 1% drop in average revenue per user.
So, the metric… in itself carries its kind of own discipline, and of course, these files… this is where that human judgment comes in, right? This is where you’re gonna have this YAML file, like, as a discussion with your team, or your stakeholder, or whomever, and decide what comprises, like, how is these metrics actually calculated? What metrics should even be in this catalog? And then, The kind of, like, default like MDEs, or default if it’s a guardrail, decreases or increases you would want to guard against. And when you run the experiment itself, you can change that, you’ll have an opportunity to change that if it’s experiment-specific, but it is nice to have some defaults there.
Okay, so let’s go ahead and run… the design. So I’m just going to type… Design. And it’s gonna be running this scale up here. our design skill… You can read a bit about it. But it says, like, you know, design, pre-registers an experiment, walks from intent to semantic models to metrics to hypothesis to brief to data plane, to your seal brief. dispatching specialists at each step. So, The first thing it does is, it’s going to load that skill and try to pull some Intent for me, so it says, what do you want to test? Describe the intent in plain English, what change, what surface. What… where are you expected to move, so… Let’s say… I’ll type it out. We’re gonna use just the sample data in here.
And what I want to test is… I have a prompt over here. So, I want to test whether moving the Buy Now button above the fold on the product page improves conversion. And so, what should happen now is the designer agent should start reasoning against… Metrics, and should start proposing some primary and guardrail metrics. And give me, some proposals. So you can see… It’s listing out these steps that it says it’s gonna walk through, right? Capture intent. And then there’s gates for out of a lot of these, confirm the intent gate. And then it’s going to ensure the semantic models and metrics exist. Then it’s begin drafting a hypothesis. From that hypothesis, it’s going to draft a brief.
You can’t see these four pending ones, but at that point, it’ll have a critic come in. So… Yes. It’s going to create a new experiment here, so it just created this new experiment ID. And it’s just confirming that this is what I’m trying to do. Yes, proceed. Now it’s gonna go and check through the semantic models to figure out What data is available? And it’s not gonna make metrics up here, right? It’s gonna look at the metrics directory. That catalog we looked at earlier. So, it says here, yes, semantic models and metrics already cover relevant entities, and it’s inferring some stuff here, right?
It’s going to give me an opportunity later on if I want to use different metrics, but there’s not too many to pick from, so it’s pretty easy to figure out here that I want conversion rate as my primary, because I said conversion. in my intent sentence here around what I want to test. If you don’t have a metric within your metric catalog here that it thinks you should test, it’s gonna flag that, and, it’ll give you the opportunity to, create that metric yourself, or work with the understander agent to kind of, like, co-pilot through creating a metric. Once you create that metric with it, it’s gonna check through your… semantic models to see if that even exists in the data available.
If it doesn’t have a metric, it’s not going to run. If the metric doesn’t exist in the semantic models, it’s not going to run. So those are really critical, and it’s a way to also just kind of, like, add some safety nets around the work you’re doing, just having very clear metric definitions and semantic models defined. Otherwise, you know, you could make stuff up. But there’s lots of gates throughout here to make sure it’s going to check that at each step. So, like I said, it picked conversion rate because it matches the word I used. I think it also proposed… Some other metrics up there… Come on, Maria. Yeah, I proposed a secondary, metric here. Buy now, click rate.
It’s probably going to propose some guardrails when it gets into the brief stage. Sometimes they’ll propose them, Beforehand. I could intervene at this point, and I want… if I wanted to, and swap the metric and say, oh, actually, I want to use add to cart rate as the primary, and keep conversion as secondary, if I hit, like, CTRL-C and just interrupted it. So, it’s hitting some errors… While it loops through trying to get the hypothesis. We’ll let it work it out. We can take some questions while this, Ronzo, if there’s anything coming up, Strawberry, and hi.
Hai Guan: Yeah, there’s, quite a bit of questions, so it would be cool as we run this. So, one question is, how do we create the YAML files for the metrics and thresholds?
Shane Butler: Yeah, so there’s a template here. I mean… there’s two ways to create it, right? Like. This is… this is where I think the… the human part is important. So, like, metric definition, hopefully somewhere in your company, you have some sort of metric definition, some sort of metric catalog, some sort of document. Maybe it’s just living in Slack or in someone’s head, or maybe it’s in a SQL query somewhere. Well, like, this is not, like, the automated part of coming up with the definition itself. You can get it into the structure just by working with Claude and explaining, like, hey, here is, Here’s how we define this metric. These are the fields. In this semantic model that we pull from.
Here is the kind of, like, typical MDE, that we want to see. It can infer a lot of this stuff, too, right? If you’re saying, like, revenue, it’s not going to be like, oh, the direction of lower revenue is better. But this is, I think, a conversation with your team as to, like, how metrics are defined. We’re not gonna go into, like, how do you make really strong metrics today. We’re actually gonna do that next week for free also. We’re gonna… we have a session next Friday called North Star Metrics with AI, and we’re gonna talk about how you can, create really strong metrics that kind of, like, align with user value. And then you can kind of, like.
hand that off to Claude to… align to this, template.yaml to make the right metrics. Okay, so this is… What is triggering right now is I ran this… I ran this, right before we started. I tested out this demo, so it’s actually… it caught that and said, hey, this experiment, was already ran, and this experiment ID, is this a fresh one or a replication? We’re gonna say that this is a fresh test, so you can kind of see it go through and draft everything from the start. What other questions?
Hai Guan: Let’s see… how is this repo’s skills? for design and analysis, different than the AI Analyst repo.
Shane Butler: Yeah, the AI Analyst repo, it’s not purely, experiment. focused, and I think with the AI Analyst Repo, like, a lot of that is… I do want a lot of human judgment in there as well. the use case for Agent XP is, like. It’s really… it’s only focused on A-B test. It’s not gonna do root cause analysis for you, it’s not gonna draw you a full deck, It’s just for experiment design and analysis, and the primary goal of it is to remove bias wherever possible, whether that bias is from the human or from the LLM.
Ideally, like, our goal here would be that anyone on your team can, whether you’re a PM, an engineer, a designer, or a data scientist, for a simple experiment, should be able to run this and get a really solid experiment design. Now, that doesn’t mean that, like, someone just runs it. If you don’t have any data resources, like, yeah, this is definitely going to be probably better than what you’re doing today, but you probably want a data scientist to review if you’re not a data scientist. If you are a data scientist. I think this speeds up your design Process, and allows you still to, like, focus more critically on, like, the really complicated, challenging.
Experiments that are more nuanced, rather than, like. the kind of simpler ones. So… Divert work to more complicated Work more for data scientists, and then open up experiment design to teams that don’t have A data scientist working with them.
Hai Guan: Okay.
Shane Butler: What else?
Hai Guan: more questions. Attendee had a couple questions here. One is… Where is randomization defined?
Shane Butler: You’re right.
Hai Guan: Is it possible to randomize?
Shane Butler: Yeah, so you’re thinking around, like, the… There’s that… those three pieces I talked about, right? There’s, the design. There’s, like, the actual, like, running, instrumentation, implementation, and then there’s the, analyze. This isn’t, like, a full, like, it’s gonna look at your user base and assign everyone to treatment and control. Not yet. It’s not something terribly… challenging to do. It will check for randomization after the fact in the analyze stage, but this doesn’t do, like, the assignment and the, like. Implementation of the test itself. But okay, I mean, that’s not something that’s terribly hard to do. But it does require, like.
you know, now you’re kind of plugging into a system where you are turning features on and off for certain users. So, like, you know, there’s whole platforms that are already solved for that. We’re… we’re not really trying to… recreate something that, like, experimentation platforms are already pretty good at. We’re trying to take the part where experimentation platforms aren’t Super good at, where they’re, like, giving you the numbers. Without, like, putting guardrails around, like, did you design that correctly, and are you utilizing them correctly, and making the right judgment from it? So we do have some briefs coming out.
Now, so if we go into experiments… And we go to… here’s our experiment. This is the ID here. It’s starting to create some readouts, so it’s got our, initial intent readout. It’s got a YAML of the hypothesis here, so you can see it extended that intent we had to a pretty detailed hypothesis. It has rationale around, you know, why we believe that hypothesis. it gets into a bit of the, kind of, leading indicators and guardrail metrics. If we go into the brief YAML, it’ll actually assign… start assigning, numbers to these. So, There’s some default, MDE figures in there for each, metric, so it’ll, it’ll leverage those. To start, but… once this outputs to us, we can change it at any time.
So here it says, like, brief drafted with feasible 5% relative MDE. So what it did was it went and it ran a power analysis and, you know, below a 5% MDE, Probably not gonna be able to run this test. It’s not just running… the power analysis for, the primary, it’s running it for all your guardrails as well, which is something that I think a lot of people miss out and forget about. They think, oh, we’re going to make sure we have enough power to measure our primary metric, but then they don’t actually have enough power to know if the guardrail metrics have been triggered. So it’s gonna do that for you as well. And it’ll have those, Those, thresholds that you can’t pass.
It’s going to define… The exposure assignment, it’s not going to… do the randomization for you, like, say, user, you know, XYZ is it, but it’s going to define the unit of measurement for you. And then what’s happening right now is, after the, the draft… The brief is drafted by the designer agent. It’s being hit by that other agent, the critic agent, to kind of review it and go through a review loop cycle. We can keep taking questions. Yeah, so critic, brief consistency, not brief. And… We’ll share this out. And you can kind of read through those, what each of those agents, do, but, like. Now, here’s our critic agent.
The critic agent doesn’t just judge brief consistency, it all Say, like, hey, it’ll… it’ll look at the analysis versus the brief, it’ll look at the verdict versus the analysis, it’ll even check, like, hey, are the stats that you pull during analysis just from the whitelist in agent XP.Stats, or did the LLM go and start… go rogue and create its own functions? So try not as many guardrails as possible.
Hai Guan: One question from the chat is, where do the data get connected? I think at least that’s the flavor of them.
Shane Butler: Yeah, so there’s, like, a connect data skill that we’ve created, and I’ve… Been working with it. Right now, I’m just working on a local .DB database on my machine, some sample data. You can… like, any of this, like, this stuff’s really easy to connect to Snowflake or Databricks or whatever. You can run this skill, and it’ll help you wire it up, but you’re gonna have to create semantic models That tell… the agents where to look. So, like, we have those over here. That tell, like, hey, what source table are you gonna pull from?
What are the dimensions within that table, and then the metrics Over here are going to relate to those semantic models, but you want to have those, like, clearly defined in your… Repo here, so it can have that context about what to pull from where. it’s not just going rogue and guessing stuff from your warehouse. You probably have something already pretty similar in GitHub in order to create those tables. Okay, so it has created the brief. Actually, we’re doing pretty good on time here. We’re not gonna get through the analyze stuff. You guys can run that on your own. We can do another session.
But, it says, like, here’s the brief where it is, do we want to, like, seal it and move on to the next step, or is there anything we want to change? For instance, like, the MDE, I’m gonna say seal it, because we have… 8 minutes left. And then, it should generate a nice little readout for me. What other questions while we wait for that?
Hai Guan: Or… there’s a lot of things going on in chat. For folks who have questions that haven’t been answered, do you want to come off mute and ask?
Shane Butler: We’re gonna put him in, We should put some of these in the Slack channel, too, for afterwards. I can also answer some async, yeah, but…
Hai Guan: Let’s see…
Sravya Madipalli: Maybe we can also tell folks, guys, we have a Slack channel. We are at 899 members right now, so you could be the luckiest 900th member if you join it. I’ll share the Slack link. Yeah. For any questions, you could also, you know, post your questions in the Slack channel, too.
Shane Butler: Yeah, we should pull these out into the Slack, and just answer them all in the Slack. Maybe I’ll pull these out in, like. answer them in a follow-up email, or I’ll create a… A page with the answers to them. Okay, so this… the design… portion has wrapped up. So, like. Yeah, for a lightning lesson, I don’t know if I’d demo that again that way, because it did take 13 minutes end-to-end, which is kind of a long time for a demo. But, in your real life, like, you know… Kick this off, go get a cup of coffee, maybe a glass of water, and a small snack, and you come back, and you got your design ready for you, so… It is a pretty quick first pass at the… Experiment design.
Yeah, let me share the repo right now before I forget. So I’ll drop the repo and, I just pushed some changes to it, and right before this, I did find some weird stuff, so I’m gonna push some more changes probably this evening. This thing will probably be pretty, pretty actively iterated on, and it’s also pretty early… Dazed, but if anyone wants to… Play around with it. Feel free. I think we could probably open it up for a contribution, too. That would be awesome. If anyone wants to contribute to it. Let me… See where our readout is, though. Design brief. Oh, hey, it didn’t give me a nice… Members. You didn’t give me the HTML. or a PDF version of the readout. Just gave me that MD.
So we’ll let it, take another pass at that. I’m gonna switch back to slides. But once the readout’s finished, I’ll, I will… show that. Okay, we only have 5 minutes left here. I mean, I can stay over a bit if we have questions. But I think, I think the main thing here is, like. You know, a data scientist doesn’t have to spend all their time on… the experiments that genuinely… they can spend their time on experiments that genuinely need their judgment. Like, the design review conversation where someone says, you’re testing the wrong thing, that’s still, like, a human meeting, and we can invest a lot more time in that.
I think this cycle, leveraging, like, logistics systems like this, really not only, like, compresses the work and makes it faster, but it does have a real opportunity in this case to add Rigor to it. It’s like… I don’t know. human judgment on making calls around experimentation, we talk a lot about how we need to put guardrails on, like, what agents and LLMs can do and check their work, but, like, there’s a lot of human error. in… that occurs when running experiments, and I think there’s an opportunity here to add some nice, guardrails around it that aren’t, like, crazy rigid through code and have some flexibility around, like, your specific use case leveraging, LLMs.
So… Once that design brief gets generated, I’ll show you guys it here. We do have some upcoming sessions, just… or courses in the next couple weeks, if anyone’s interested, and then we can talk about questions, too. Next Wednesday. For anyone here, if you haven’t… I don’t know if people have… who has and hasn’t used Cloud Code, but if you’ve never used Cloud Code before. We do a super intro workshop next Wednesday. It’s super cheap. We get you to, like, hold your hand going through installing Cloud Code and running analysis in there for the first time with our open source AI analyst repo. It’s, like, 3 hours. it’s pretty beginner.
If you’ve… if you’ve used Cloud Code before, you probably don’t need to do that. But in the following weekend, we have our full Cloud Code Analytics Bootcamp, where we, teach the framework around building an Agentic analytics system yourself. We have you run multiple analyses in there, we have you built on top of our repo, we have you port over stuff that you’ve built from one repo to another, we connect to data warehouses, we connect to MCPs. This is kind of like our foundational bootcamp. Then, later in June, we have our advanced bootcamp, what gets into AI evaluation, validation, and context management, and then multi-models.
So, like, not just Cloud Code, but codecs and open source models. We’ll get a bit into validation, open source, and, context management in the weekend boot camp as well, but it’s gonna be Kind of intro async material. But both of those are live, they’re both, like, 4-5 hour sessions. We’re also gonna run the Cloud Coding Analytics Bootcamp in July again during weekdays for just 2 hours a day. If you’re more of a weekday person, And then… In a couple weeks, we’re kicking off AI Analytics for Builders. That’s that one Shalabi I mentioned, where the sec- where the fourth week is all on experimentation.
This is our kind of, like, think… the end-to-end, sort of, like, how do I become a product data scientist in tech course, what are the skills I need from framing questions, working with stakeholders, developing metrics, setting up the tools, doing root cause analysis, doing trend analysis, doing segmentation, doing experimentation. Running causal inference models, Drafting presentations, doing stakeholder readouts, all of that, but executed in Cloud Code. And I don’t know, maybe we’ll do some execution in codecs this time, too, as that’s, getting… seems like more and more, Robust when it comes to data analysis.
There is a promo going on with Maven right now that you all probably have seen, where they’re top courses. You can get 25% off with this promo code, MAVEN100. This is better than any promo that we run. We usually give 10-20% off. That expires Sunday, June 7th. at 11.59 PM, I bet. So, you got a couple more days if you want 25% off. That’s, like, site-wide off all their top courses, so… I’m sure some of the other ones people dropped in chat around, like, experimentation, too, have that same deal. Let’s see if the… I wonder if it… Finish the experiment brief. I mean, this is the experiment brief I showed you before, so this is, like. what it outfits by.
I want to see the one… That we actually did. Any… any questions while, kind of wrap up here? I don’t know, Sabbi and Hirif, you gotta go, but I can stick around for a bit. Yeah, Attendee.
Attendee: Hey, Shane, hey team. So quick question, perhaps I missed. In this entire course, we were… we were basically targeting the first and the third pillar of the entire process. First one was about designing, and the, third one was about, you know, the post-experiment analysis or readouts, which the LLMs would assist us. what I probably missed, and that’s a question, did you already have, like, some experiment results in the repo? based upon which the third pillar got executed?
Shane Butler: I have some… I mean, I mean, I’m working with, like, a sample… data set, so I have a database locally on my machine… my machine that I’m connected to. It’s not… it’s not in the repo, but, where I have, some experiment results. Well, I have the… I have a table that… those tables that the metrics are comprised from. And then I have a table, like you probably would at your company, where it’s, like, the exposure… the exposure periods at whatever level. So there’s a table that’s basically, like. user… 1, 2, 3, 4, 5, Shane… Is in Experiment ABC… Yeah. Very treatment, this start date, this end date. That’s how I’ve usually… there’s usually some table like that. at your company.
So, the analyze step, which we didn’t get to today, but you can kind of check out how does it… will aggregate and do all of the statistics by writing a SQL query that joins to that table to the, tables that comprise the metrics.
Attendee: Yeah, yeah, yeah, got it, got it. Cool.
Shane Butler: But yeah, this… yeah, this doesn’t, this doesn’t do that middle part, though, of, like, actually, like, you know, building that table and, assigning.
Attendee: Yeah, that…
Shane Butler: experiment.
Attendee: Yeah, that makes sense. I was a little confused where, I mean, I was like, how did the third pillar get executed if we do not have experiment results? But, like you said, it is stored somewhere on.
Shane Butler: Yeah.
Attendee: Yeah.
Shane Butler: And I think you could do that, like, I would love… I would probably work on building it out to… get that middle, stage, but I just think it is, like, a higher risk stage of, like. you know, if I mess something up with the design, it’s like I can go back and fix that before it starts. But… If we, like… Screw up the actual, like, instrumentation and assignment, and then run that test. For a couple weeks, it’s like, oh dang, we just… we not only ran… the data’s useless, but we also could have given a bunch of people the wrong experience or something.
Attendee: Yeah. Cool, awesome, thanks.
Shane Butler: Yeah. What else I think the… design brief is done. I’m just trying to find it on my computer here. Other questions in the… Chats, Robbie, or hi.
Hai Guan: Lots of questions around the Slack channel, so if folks who can’t access Slack Either because you haven’t used it before, or you don’t have an account, or if you have an account, but you still can’t access it, just reach out to one of us, or reach out to me on LinkedIn, and we’ll get you, we’ll get you in.
Shane Butler: Is it saying that you have to have a certain email domain or something? Is that why the…
Attendee: Yeah, that’s what I’m seeing right now.
Shane Butler: Yeah, yeah, I don’t know what is up with that. Let me…
Attendee: Did you guys.
Shane Butler: Try a link again.
Attendee: Yeah, because I have a Slack account, with a personal email. When I click on it, it says it needs…
Shane Butler: Data Neighbor AI Analyst Laboratory.
Attendee: Yeah, they’re saying reach out to the coordinator to create an account, something like this.
Shane Butler: Yeah, we gotta… let me try this…
Sravya Madipalli: Can you see the news Inc.
Shane Butler: Yeah, yeah. I think at some point. The invite links we have, like, revert back to only people with the company, domain can access it, but I think that one I just dropped in, or the one Stravia just dropped in. Hopefully that works. I’m gonna try that. What else?
Hai Guan: The landing page is definitely different with this new one, so that should work.
Shane Butler: Any ques… doesn’t have to be experiment… Agentech experiment-related questions, if there’s anything else. How’s the tool know the right, KPI to… ITUs? Yeah, so there’s… That metric catalog that I showed you, there’s, like, just, like, 14 metrics in there right now, and so it inferred for me that it wanted a conversion rate, because when I, like, told it, like, what I want to run an experiment on, I used the word conversion. if it can’t figure it out, it’s gonna interview you more. If there’s no metric in there that matches, like, what your intent is, it’s gonna flag that.
But, you know, and you could make this more clear, too, I mean, we could add a thing in the Yamify, which I don’t have, where it’s like, questions like these, and have example questions, should use this type of primary metric. In most cases, like, a lot of the experiments are gonna only have, like. if you’re in a specific product area, you’re probably focused on very specific metrics. Like, we’ll go over this next week in our North Star Metric, free session, where, like, you know. each team really has, like, a primary metric that it’s trying to drive, which reflects the user value it’s trying to build. So, there’s something you could definitely map in there, I think, into the YAML file.
Good question, though. And also, like, I think it’s like, I put this in the README on the HNXP thing, but it’s like, this is also just, like, something where, you know, we spent… the past… I’d say since, like, mid-year last year, we’ve been pretty focused more on, like… we’ve been focused on the data analysis side, because we’re more… confident around Agentic systems being able to do, you know, stuff that’s very easily explainable, like segmentation and root cause, like, correlative kind of, behaviors. We haven’t really pushed it into a lot of the, kind of, like.
causal relationship stuff, or we do a little bit in builders, because that gets into a lot of, kind of, judgment around what variables do we create, what transformations do we create on those variables, what methodology do we actually leverage to, try and find this causal relationship? So this is, like, a little bit of a new frontier for us, around experimentation, but it is a lot more cut and dry than, say, like, causal inference modeling, where you’re, like, creating a bunch of variables. Which is why we’re kind of starting there. But as we learn more, we’ll build this out and share it more, too. Nights. Any other questions here in the chat? Any other…
Hai Guan: covered everything, if not async already, so if folks… yeah, it’s too long of a wall of text here to parse. If anyone has any questions, feel free to raise your hand or just come off mute.
Shane Butler: Cool? Well, nowhere’s not. We will, Was there any agent on impact calculation? for… Rollout, yeah. In the analyze… I didn’t… I didn’t get into any of the, kind of, like, analyze phase, but there’s… there’s a bunch of, checks there. There’s, like, the 8… there’s, like, 8, stage checks. It’s not just gonna be like, hey, is it a… above or below your MDE is going to be, like, checking the confidence intervals, it’s gonna be checking, like, is the impact actually large enough that you want to roll it out, or is it, like, statistically significant, but quite small? So, maybe we’ll do another session on that. I think it was a bit too much to try to design an analysis in one session. Cool.
Yeah, so just a reminder, Intro to Cloud Code Analytics Workshop’s next Wednesday. The bootcamp is 13th and 14th. And then Builders is on the 15th, so… Yeah, the… oh, and we do a Builders Bootcamp thing, two for one, if you’re interested in that. Basically, you have to do the bootcamp. and then you decide you want to go on to the 5-week course, we remove whatever you paid for the bootcamp from the price of builders. So that’s a pretty good deal. And then, yeah, the Maven 100 deal, I think, ends on Sunday, so whether it’s our course or another one, I’d recommend checking out their catalog by Sunday. This is probably, like, one of the best deals you’re gonna get on Maven anytime soon.
We’ll sound out an email. With the repo, in case folks missed it in the chat. Though if you look up, like, AI Analyst Lab, on GitHub. you’ll find it in that org, Agent XP. We’ll also send out the recording And, I think that’s it. I know. I’ll send out the PDF or something it generated, too, if you want to check that out. Cool. Well, thanks everyone, appreciate the time. And hopefully see some of you in our courses, or next week in our free sessions. We got two next week. Northstar Metrics, and then we have a pretty cool one from Straub he’s gonna do our ads called, like, what’s it called? Share… What’s the name of it? It’s a good name.
Sravya Madipalli: share, like, basically share analysis any… every… everywhere using AI.
Shane Butler: Yeah, yeah, so it’ll be around, like, MCPs and stuff. It’ll be pretty cool, so… Those are both free. Check those out. Alright, have a great weekend, everyone.
Hai Guan: Bye off.