← All free workshops
Free workshop · Thursday, February 19, 2026

Design Experiments for AI Features

Live on Maven, Wednesdays at 10 AM Pacific. About 60 minutes.

Transcript

Auto-transcribed from the live session and lightly cleaned. Attendee names are removed; their questions are kept.

Shane Butler: Duck. Hello, hello! Can you all… let’s see, make my chat open. Can y’all hear me all right? I’m gonna drop in the chat. I wonder if you can, hear me? Nice. Do you see a screen that says design experiments for AI Features? It’s the next question. All right, W yes. We’ll hang around a little bit, see if more people join. I think I got 180 people signed up, but we’ve got… only got 3 of you in here so far, so it may be a small group. I got some ripples coming in. Let me… let me open up the… Disable the waiting room. Hey, how’s it going, Attendee? Attendee, you’ll guess you’re back for more. Where are y’all located? I guess you’re in Austria, right? You like these late night… Workshops for you.

Attendee: Yeah, that’s right. It’s 9 PM.

Shane Butler: Oh, okay, alright, nice. Well, I appreciate you coming on late, dude. Two days in a row.

Attendee: Yeah.

Shane Butler: Attendee, Attendee, Attendee, where are you guys from? Attendee? Hey, Attendee. I’m located in, California. Santa Barbara.

Attendee: How’s it going, Shane? I’m from, Tulsa, Oklahoma, right in the center of the U.S. over here.

Shane Butler: Nice, awesome. Awesome, nice to meet you. Washington… I lived in Santa Barbara for about 8 years. I went to UCSB for undergrad, moved back later. But it was too good down there, so I had to go somewhere. Or I had to postnat or something, I don’t know. I love Santa Barbara. My whole family lives there. Hey, Attendee. Hey, Attendee. Okay, let’s get started. People join later, they can watch the recording. Okay, so today, I’m Shane. Thanks for joining. I teach AI evaluations and product analytics, AI analytics at Maven. Today, we’re going to talk about how, to design experiments for AI features, so… I’d say, like, a lot of the stuff is gonna be exactly the same.

there’s gonna be some, like, stuff you should do before you go into a live experiment. But, for the most part, if you know how to run a B-test. Treatment control, primary metric, ship decisions. If you’ve done this for traditional features, then you know, like, the… you know, I don’t know, 80… 80-85% of the playbook here. What I’m gonna show you today is just, like, how a few specific decisions, slightly change with AI features, but they can actually, like, really throw off the decision you make when you look at your results. So it’s not the whole, like, playbook kind of thing for A-B testing. I’m kind of assuming folks know some of that.

Some of them will look the same on the surface, but for AI features, it’s, like, beneath the surface, I guess some things change around, like, the distributions of the data you’re working with. But maybe just to get a vibe of the room. quick check in the chat. Maybe drop a 1 if you’ve run A-B tests before, but not on AI features. Like, you’ve ran them, or your team’s random, or your company’s random. Two, if you’ve shipped AI features, but you haven’t really ran, like, A-B tests on them, and then three, you’re, you’re running… A-B test on AI features, I guess. And maybe some things look off sometimes. 3, nice, two, two, one, two… Cool. Alright.

So, we’ll basically get to 3 today, good, everyone’s at least a 1. I guess 0 would be, like, if you’ve never… if you don’t know what an A-B test is. If you don’t know what A-B test is, not gonna get into that today, but there are many people who teach A-B tests, and many great YouTube videos, and you can just ask LLMs today what… how that works. Okay, let’s jump into it, though. We’ll start with a… like… potential use case here, so… Say you have some team they’re building a AI, like, data analyst for their… internal… stakeholders, so for their… maybe for their product org, their, like, product and engineers orgs.

So, this AI analyst, you can feed it some natural language question, like, what’s revenue look like today, or something like that. And outputs… it outputs SQL, It outputs charts, it has a narrative around the data it’s outputting as well. It sits on top of your company’s, like, analytics database, so it has access to data around, like, your users. Events, revenue, sessions, maybe there’s, like, 147 tables in this database. So I’m giving you the details, because they’ll matter in a little bit. let’s say this agent, this, like, SQL agent, there’s, like, two steps in how it retrieves the answer when someone inputs a question.

The first one, It finds the top 20 candidate tables from those 147 tables in your database. That could answer the question, or would be required to join to each other to answer the question. And the second step, it picks, like, the best 5 of those to use As context. So, those tables, plus some example queries, They go into… like… OpenAI, API, For SQL generation. That result gets charted and narrated. And, all along… so you see, like, all along, there’s, like, different steps that have to happen. to get that information just from the input of a question to the output of a user. There’s multiple steps where an AI is interacting with the data and making decisions.

And each of those steps can have, basically, leakage in terms of the users that are actually getting exposed to experience. So in this case. Let’s say we want to run a test, Where we have, like, 1500… daily active users within our product org that uses this, like, internal tool. Like, this tool, like, already exists, but we want to make a V2 of the tool. And it’s gonna have, like, a more sophisticated ranking of, like, what are those top 20 tables it pulls from. It’s gonna expand, like, it’s gonna have, like, more schema context, more context around the data to pull from. And, maybe it’s, like… say, like, we’re back in the day, and we were on, like, GPT 3.5, and it’s… it’s gonna use 4.0 now.

So, they run a standard A-B test, they take the 1500 users, they split them randomly in half, they run it for 2 weeks, and they have some primary metrics around, like, did SQL return the right answer, or return an answer. the results come back, and they say, like, okay, we’ve got a 3% increase, but our p-value, and p-value means, like, is that increase real, or is it due to some randomness? Is that 0.15? But you can imagine maybe this p-value is at, like, 0.5, or something really higher. For this to be a real statistically significant Difference, it has to be… the industry standard is, like, a .05, so say it’s way above that, perhaps.

It’s not so significant, and it’s expensive to run these more complicated models, it’s expensive to have larger context windows, because it… takes more tokens up that you’re hitting the API with. They maybe… they take longer, too. So the team’s like, well, it didn’t really significantly increase whether it returns the right answer, and it’s costing us a lot more money, so we’re just gonna, like, shell that. And so, if you think about it, it’s like, this could be, like. a lot of resources probably went into this, like, could be, like, months of engineering work, maybe, who knows?

And… Decisions can get made around, like, shelving, whether it’s… Brill it out, or… Not seen by anyone else, just based off one number. And that one number can not… can, especially in AI systems, in any… in any system, really, cannot really be an accurate reflection of what’s going on. So, in this case, when we look at, like, The actual people who received the new… the new experience. Of those 743 people we thought were going to receive the V2, we find out, because of a bunch of leakage through all the different, stages that an AI interacts with are… Our query and our data, that only maybe 23% of the people, 167 people. actually receive the experience.

Everyone else got some sort of, maybe, like, fallback to, the V1 experience at different stages. And this is something that can happen without anyone really knowing about it, if, like, the correct guardrails aren’t in place. And this isn’t… this is… this kind of leakage exists for any complex feature, like, not just AI, like, payment systems. search ranking, email delivery, it’s not a new problem, but what’s new is the scale of people who are now building features that have these kind of multi-step opportunities for leakage.

So… like… I don’t know, like, putting AI in a product just isn’t that hard anymore, and there’s a lot of stuff going on under the hood that… it’s not as simple as when teams are like, oh, I’m changing the button from red to blue, and we know 99% of the people who are in our treatment group are going to see That blue button. In this case, it could be that a bunch of the people we lose on the way. So let me show you, like. Some, like, potential scenario where this could happen. So, say you have a funnel. Of, where, like, different users can leak along the way, and this is basically kind of, like, your pipeline of what’s going on from your input to your output. So, in this case.

We have 743 users that are assigned to V2, half of that original 1,500 population. Maybe 156 people never opened the tool during that window. Or maybe they, like, switched to using DirectSQL instead, so that’s an exposure problem right there. now you’re down to 587 users. Maybe 190 of those queries hit a cache, so a lot of the… And this will be with, like, many products, but especially with, like, data querying products, a lot of the query answers are stored in cache, so you’re not having to hit whatever your data warehouse is again. It makes it faster, it makes it cheaper. So maybe the system’s actually serving stale V1 results, but nobody actually knows that.

maybe 129 people hit, like, a timeout on one of the V2 retrieval steps, and so it falls back to the faster V1 model, but it doesn’t inform them of that. So now you’re down to 268, and then maybe there’s, you know, very similar, maybe there’s, like, some, like, syntax errors or permission errors with the V2, and those cause a fallback to V1 as well. So has anyone ever, like, shipped a feature? Or, like, an AI feature. Where you later discovered, like, most users weren’t actually even getting the new experience. Yeah, like, sample size gets, like… pretty screwed up.

Like, this can happen in traditional stuff, too, but it’s just that now, when we’re like, yeah, let’s create this agentic framework, there’s so much stuff that can happen along the way. Sample size is going to be kind of the name of the game in this presentation. There’ll be even more things that affect sample size. So, Say when you account for all that stuff, and you just look at those 167 people who got the new experience. Instead of the whole 750 people, you say, like, oh, actually. The real effect for that small group of people who actually got the experience is, like, 4 to 5 times larger. maybe it’s, like, plus 14% instead of 3%.

And then you might then conclude, like, oh, like, the feature works, we can just look at that. But the catch is, like, like. Attendee, I said in the chat here, it’s, Knowing the real effect is bigger doesn’t make it significant anymore. You have a much lower sample size. You were going into this, you probably… the team probably did, like, a power calculation where they said, like, hey, we need a sample size of 760, and now they only got 167. So, it tells you why the test came back flat.

It tells you that the test came back flat because You have leakage, but it doesn’t necessarily tell you that… you have… a statistically significant result yet, so you have to actually go back and now rerun this test. So… The good thing with this is, like, All of this. can be fixed, like, you can fix the cache, you can handle timeouts, you can… Do whatever you can to get the people who actually are exposed Higher and higher and higher, within that funnel. Honestly, I’m just reading the, the chat here. Ran some A-B tests and never got stats like because the sample included people who never got exposed to the features. Yeah.

The story’s totally different, yeah, when you actually limit to people who are exposed. But if you don’t kind of account for all of that up front, then you… you don’t just get to, like. filter down, and then see what the result is. You kind of have to rerun it all over again. So, definitely something that’s pretty important, I think, early on with AI features, before you run an experiment. Understanding every single possible point of leakage, and… either… Calculating your sample size down lower in the pipeline, where you know leakage has stopped, or doing whatever you can to, like, fix the leakage earlier. The other thing here is the metrics.

So… The second problem, the teacher… the team… Measured this metric, like. SQL success rate, but that could be defined many different ways, and potentially, in this case, like, it’s really easy to measure, like, the wrong thing as well. This goes back to just, like. setting up proper AI eval metrics, a lot of that will happen offline before you even get into a production experiment environment, but all of those eval metrics you create for offline evaluation have to continue through to online evaluation as well. You also want to evaluate user value stuff too, but it’s, something like, if you measure, like, SQL success rate. that’s not going to, like, does the SQL compile and run without error.

If users actually care about something else, like, does it return the right answer? then that’s not actually helpful. So in this case, like, say they had 200 reference questions from known correct answers maintained by the finance team or something, and when they checked V2 against those grand truth answers, like, 5% of them were incorrect. They have, like, a confidently wrong answer, where the SQL executes, the chart renders, the narrative sounds authoritative, but the numbers are wrong by, you know, more than 10%, and the user can’t really tell without going through and looking at the SQL themselves.

So… This is a failure mode that’s genuinely different about AI, It looks like success when it isn’t. And, you know, you can end up in situations where this is out in the wild, and then it’s, like, two weeks later, after a full ramp of this new V2, like, a VP or something presents some quarterly revenue to the board, and it’s, like, totally the wrong number. I don’t know, has anyone ran into this? I guess this is more like a analytics thing than an experimentation thing, but, like, where someone’s had a metric that looked fine on the surface. When something was, like, totally wrong underneath that got reported out.

Like, this can happen in regular… Dashboards, or just queries that your analyst pulls for you. That same problem’s gonna happen… with AI, no matter what product you’re doing, so all those offline eval metrics have to be pulled through to your online experiment as well. So, I mean, for a traditional… for a traditional product, like, this team probably ran, like, a pretty well-designed experiment. They may have been fine, they probably had, like, 50-50 split, reasonable sample size, they had some metric, they had, like, a… Two-week window, which is probably long enough for that.

But for AI features, There’s kind of these different changes within each part of these decisions along the way of… Of both. designing your experiment and then reading the results later on. So, like, they didn’t get this wrong because they’re… they’re sloppy or anything in experiment design, they just got it wrong because they weren’t checking off a few additional boxes when they’re working with AI features. And a lot of that actually just has to do with Sample size, and also, like, manually Kind of reading the results, or having some… Way of automating up, scaling up manual reading of results.

So what I’ll do it right now, I’ll go through these kind of, like, four… part piece of decision that you’ll want to watch out for. It’s like, size it, measure it, guard it, and read the results. For each one, I’ll show you kind of what they assumed. What’s actually true, what to do differently. So, in this case, like, if we look at, Size it, we talked about this a lot already. When you change abutting color. Right, everyone sees the same button. When you change the AI model, everyone gets a different output. So, as we said before, like. In that funnel, when everything’s leaking out, And you have to… account for that?

It means you mean, like, the… the fraction… that gets the treatment requires more and more users you’ll need to put in your sample size if you don’t fix those parts of your funnel. So, like, step one, like, say that that gets diluted. Only a quarter of users get the treatment, so you measure it… your measured effect shrinks to, like, a quarter of the real effect then. The issue here is that detecting smaller effects doesn’t… just need proportionately more users, and use disproportionately more. So, think of it like trying to, like, hear a quiet sound in a noisy room. If the sound’s half as loud. you don’t just need to listen twice as hard, you have to listen, like, four times as hard.

That’s kind of how power works as well when you’re getting your sample size, so if you don’t fix those leakage effects. You just have to have so many more people at the top of your funnel. There’s many ways the sample gets diluted. Leakage is one of them. But there’s multipliers, Of other ways it gets diluted, too. So, one, the AI side is just… noisier. So, when I say noisier, you have… Much more variance in the experience that the users are getting, so… When you create a test that’s, like. red-blue button, you know everyone’s getting the blue button, but when you have a AI feature output.

I could ask for the same query as someone else in treatment, and we could get totally different responses due to the probabilistic nature. So you have much wider variance In your actual experience that you’re getting, which just kind of dilutes the result of what you’re getting as well. So… The other part of this, too, is, you have to, like, segment everything, like, the AI features could perform really well in one aspect of your data, or use cases, or users, but really poorly in another. So, an example might be, on really simple queries. The V2 does great, but on complex ones. It does terribly, which means now you have to, like, segment down your data.

And for each of those segments, you have to have, like. A larger sample size to compare the treatment segment of simple queries to the control. segment of Simple Queries. And all of these things, like, multiply together. They’re not just, like, additives, so you can end up in situations where you need, like. You know, like, 10 times as many users as you think you do. So measure it. when you test an AI… when you test an AI feature, your primary question is. Did the AI give the right answer, which isn’t directly observable? We talked about… this team using, like, SQL success rate. But they probably had these other levers, levels they should have looked at, like, is it syntactically valid?

Did it execute? Did it return non-empty results? Did it get the actual correct answers versus some dataset we have? Yes, I will… sorry, just again in chat. I’ll, share a recording of this either today or tomorrow. It’ll auto-email out to you. So, 4 different metrics, defining the metric precisely is step zero. You’re gonna have a lot more success metrics and guardrail metrics going into this than you would a normal experiment. Once you define it. measure it, it has noise, like, like, so, so define… so… trying to see how I want to phrase this. One of the hardest things about this is, like, the returning a correct answer, thing at the bottom. So, like. When you define it.

and you measure it, that measurement in itself has noise. Like, if you use, like, say, like, an LM as a judge, a second model that grades the first. To understand if your answer is correct. A lot of people will use these because they scale. It’s a lot easier to have an LM as a judge judge a bunch of experiments’ results than a human. But say it only, like, is correct 80% of the time, so you have, like, some human annotators, and the element’s a judge, and about 80% of the time, the element judge agrees with the human annotators. Like, that sounds really good.

At, kind of, face value, but the problem is, imagine you’re grading, like, exams with a pen that randomly marks some answers right and some answers wrong. Like, it just randomly Mark’s right answers wrong 80% of the time. Both the treatment group and the control group are getting graded with this same kind of noisy pen, and so you’re not only getting, like, this noise in one group that you’re comparing, the mistakes really wash out the real difference between the groups. So now you have variance not only in The experience that people are getting, but the… yeah, variance And some, like, untrustworthiness in the actual metric itself.

So, the fix here… Is whenever possible, like, my rule of thumb is, like, use automatic checks and code alongside judges wherever possible. So, like, these first three, like, syntic tactically valid, executes without error, non-empty result, those are all things that could be checked in code. I’ve seen people… have LM as a judge, like, run those things, because they’re in the mindset of, like, oh, the LM as a judge can check if the sequel returns the correct answer. Let’s also set them up, the LM as a judge, to see if it executes without an error, or if it’s, like, tactically valid.

Stay away from that, like, any sort of metric that you can write in code and have actual deterministic checks around, that’s what you’re going to want to form metrics around. Do you already have… non-deterministic outputs you’re dealing with, and you’re gonna have at least one non-deterministic metric for an LLM as a judge. So try to reduce the others as much as possible. And then the signal is kind of like, if you see those moving in different directions, that’s a bit of a red flag for you to do, like, more manual look into what’s going on here.

I think one of, like, the kind of cool things I’ve seen Most when we’ve ran experiments, though, is, The… the pairing of offline and online evals, so… there’s kind of, like, a couple, I don’t know, school of thoughts right now. There’s a lot of people who are really, like, skip straight to experiments, skip to offline stuff, because you won’t know… what impact the AI is having on your users unless you get in front of users.

And I, like, definitely believe that a lot as well, but I think Before going into that, setting up your offline eval metrics can really set you up for success in creating great online eval metrics, and it can also eventually Help you predict what those online eval metrics will do Without having to… get kind of, I don’t know, risky, untested product right in front of them. So… What I mean by that is, So offline testing, you know, that’s what can tell you… a lot of that tells you AI quality, like, was the SQL correct? Offline testing is where, like, a lot of the annotation happened. Experiments are the only place where you can get, like, user value metrics, stuff like task completion.

or a time to insight might be one for this, data analyst example. Adoption. They’re fundamentally different. He had, like, perfect SQL accuracy. But you can actually have terrible user outcomes still, because the user asked the wrong question, or the chart was confusing, or the narrative was buried in the insight. The strategic… one of the greatest, I think, strategic payoffs of running online experiments and also having a strong offline evaluation culture that I don’t see a lot of people doing yet.

is over many experiments, you can start to see how those offline quality variables relate to real user metrics, so you can effectively Identify those relationships over time and build a sort of calibration, so you know, hey, when our offline correctness improves by 5 points, we can predict that if we get this in front of users, their tax completion may increase by 1 to 3 points. it’s a bit of a flywheel, takes, like, 5, 10, 15 experiments, kind of depends, like. how successful those experiments are as well in capturing the variance of user behavior.

But once you have it, you can start to move much faster because you have a lot more confidence in what your offline eval metrics will mean in an online environment. You still have to run the experiment eventually, but, I think you can spend a little more time that way iterating offline, which is a lot faster, making changes. And just running your offline eval suite, you know, that runs in minutes, or seconds, or a day. Rather than, like, putting something out in front of users and waiting 2 to 4 weeks for feedback. Guard It, the third one here, so, traditionally.

there’s probably some guardrails around, like, latency, so how fast something’s going, or, like, a crash rate of a site or something. There’s usually, like, 2 or 3 guardrail metrics. In AI features, there’s a lot more guardrail metrics going on. So I just read in the chat. How do I define task completion? That’d be, like, a… more of, like, a plot… that’s gonna be, like, use case specific to whatever your… product is, so it’s… it’s basically your in-product representation of what the user is trying to accomplish outside the product.

for… if I was trying to buy something on Amazon, they’d probably have task completion be something like he added… Shane added something to his cart, or Shane checked out his cart. Like, I got onto the Amazon site. I went around, and I, like, looked at a bunch of stuff, And, then I put it in a cart, so I didn’t just, like, browse and leave, and then I checked out. Task completion, I’d be… it’s, like, outcome achieved, but I think you could have, like, sub-outcomes to, like, a broader outcome, but… so it could be steps, but I think it’s, like, a step that means something with respect to user value. Like, if I go on Amazon.

And I’m just, like, sitting on some page, looking at something that’s not really, like, my task completed, even though I did get into a product. Even putting in my… Carts, probably not really a task completions by checkout. If that makes sense. So, guardrails, yeah. There’s a lot more guardrails with AI features, so… Latency is definitely a big one. Your cost per queries. So, your cost per can go up with how many tokens you’re using, could also be going up with, like, the cost per token of a given model. There could be, like, safety violations, PIA violations, like hallucination rate, confidently wrong rate, data leakage rate, it could be a guardrail, even.

So, there’s gonna be, like, you know, several more metrics across quality, cost, safety, and operations. A lot of those will be identified when you’re building your offline eval suite, when you’re creating failure modes, but I think Sometimes people overlook, like, the… cost-latency aspect of this. So, you could have… Basically, when you have more guardrails, it just means there’s, like, more chances of blocking the ability to ship a product confidently, since there’s a trade-off and a catch with each guardrail and your success metrics. So, quality of improvements you know, they can break other card rails.

Like, if you could improve the quality of your output, you could improve the correctness of your SQL, but that better retrieval would mean more API calls, means higher cost, better model can mean more latency. So you’re testing a bundle where improvements on one dimension cause regressions on another. And… This, too, you know, A degree happens in, traditional software as well, but I think it’s, like, a way bigger, impact here with AI, especially which, like, it’s like, you have to spend a lot of time even testing not only the AI feature that you’re building yourself, but which of the frontier models you’re using?

Like, what version of Gemini, or Sonnet, or Opus, or whatever you’re leveraging, and mapping out those trade-offs between cost and Latency quality, so… Speed, cost, and quality. is a much broader spectrum of, more… a lot more opportunities for one canceling out the others with AI features when you’re experimenting. And a lot of this stuff, you’re not gonna necessarily see it.

the true… Trade-off in your offline evals, because you’re not… your offline evals aren’t necessarily capturing how the user’s actually going to leverage the product, so you could be testing on a ground truth suite in offline emails, and it’s like, yeah, it’s pretty fast on this, and it’s pretty cheap, and the quality’s high, let’s ship it out. And then a user could be using it in a way that you had no idea they were going to, like, maybe some huge context they input, or really, really long chats, if it’s like a chat product, that just takes your, your, costs through the roof.

I’ll guardrail, and then what I don’t know about guardrails is just, making sure… you’ll get these probably when you’re building your offline email suite, but, like, making sure you have, like, actual real metrics measurable around them, so… You know, if a team has a guardrail around some not having some percentage of confidently wrong answers. Now you need to build some mechanism to Flag if an answer is confidently wrong. Just spot checking it when it’s… offline is okay.

Like, you don’t have a big sample size, but once you’re in a production, an online experimentation environment, you have to figure out some way whether you’re doing LM as a judge, or you have, some of those code checks I took before, like, it’s… Something like a data analyst thing is… A bit easier, because you can do things like, hey, let’s, like. run these numbers different ways, and see if they tie out, and if, like, the one they said doesn’t tie out, but it was completely wrong about tying out, but it’s gonna be different for each product.

Alright, and then Decision 4 is just, like, actually… Spending more time reading the results and seeing what’s going on in the output beyond just, like, the single number average you look at. Looking at the entire distribution of results. And segmenting it is extremely important in AI features, because you can have thing… you can have a feature that rolls out that works for 95% of use cases, but doesn’t for, you know. some 5%, and that 5% is, like, really important. Maybe it’s, like, the most complex things that people even use AI for in the first place, so they don’t even care about the other 95%. Maybe it’s things that really break trust if they’re wrong.

So, like, if it’s… really good at… I don’t know, some internal product query for a specific team, but it’s really bad at, like, finance queries, where, like, members are gonna be reported out to investors and the board. That… that’s probably not gonna ship, right? So, in that case, it’s like… You could have an average where, like. The average says this thing improved by 6%. But if it fails on a certain really important, cohort of users, or a cohort or segment of use cases, the right decision isn’t ship it, it’s like. ship it for simple queries, but hold it for complex queries, and having some definition around that.

So, this, again, does happen, I think, sometimes in traditional, traditional product, A-B testing, but a lot of times, people just roll things out for everyone. I think with AI features, it is really a lot more of, like, targeted rollouts, where, like, okay, we’re able to lock it in for this segment of use cases. We can roll out for that, but we still need to go back to the drawing board. for this remaining… maybe it’s a big… maybe it’s, like, it can only roll out for 30% of use cases, and 70%, a lot of work still needs to be done on it. So I think there’s, like, a lot more… Targeted and, subsequent, like, incremental rollouts, rather than shipping out to everyone.

At one time after an A-B test. The real practical implication for this, though, is to decide what are those segments before you run the experiment. So… If you were to run your experiment, and then just, like. slice and dice by every possible thing, and try and see where it gets wrong, and it doesn’t get it wrong, and then just roll out for what it got right, and not for what it got wrong. You’re gonna have that sample size problem again, and it’s also an issue of p-hacking, which means, like, there’s gonna be, you know, 5% of the time. Even if your test says you have a statistically significant result, it’s not actually. So, if you define those segments and use cases, beforehand.

Then you can set up your experiment in a way where you’ll have a large enough sample size to address them. Yeah, I feel like it could be quite extensive, too, Attendee, in terms of the dashboard. I feel like this probably… this could be one of those things where it’s more like. less of a dashboard and more of, like, someone has to do some deep dive analysis on really big experiments. But, yeah, it could be pretty extensive, for sure. I think y’all… I think people will… I think, like, as you go through the experiments over time, too, and also with offline eval, but especially with experiments. Like, there’ll just be, like… you’re not gonna get all the right segments in the beginning.

Like, some segments we choose are gonna be… oh, we didn’t even need a segment by this, like, it works pretty much all the time, these are more similar than we thought. Some segments we’re gonna miss, and so I think over time, there will also be, like. Probably, like, key segments that come out that are probably more prioritized in whatever your, kind of, like, experiment result dashboard are that you always look at, and then some sub-ones that you only look at if it’s specific to, like. some very specific change you’re gonna make where you’re like, oh, I think this might not work on, I don’t know, this set of users. Yeah, it’s mainly, like, custom to the feature at this point.

Like, right now, so the company I worked at, or worked at. we don’t have a big, like, experimentation platform. We run a lot of this stuff in, like. Jupyter Notebooks, experiment by experiment. The other thing that I think is, like. A little non-intuitive, but it makes sense when she’s thinking about it, is running experiments for longer than you think. And it’s not for probably the reason you think? So… You’ve probably heard that The industry is still figuring out how to experiment. Yeah, I think the industry’s definitely still trying to figure out with AI evaluation. Yeah, we do a lot of causal inference, Ahead of time, rather than experiments right away for, like, big changes.

So, the product I work is in the Legal AI sector, and so… Our users are, like, prosumers, like, they’re lawyers, they know their shit, and if we… give them bad output from AI, like, they’re gonna see it, like, very clearly, and we’re gonna lose trust with them. There’s also a lot of, like. you know, privacy and security implications and confidentially… confidential… confidentiality implications around, like, that product. So a lot of the… Testing and analysis we do is actually through causal inference, because we are… timid about breaking that trust.

We also do a lot of, like, beta testing with a small group of a small group of, like, hand raisers who, like, know they’re in a kind of experimental mode, but I’d say for sure people are still figuring it out. Okay, what was I talking about? Oh yeah, running experiments longer than you think. Okay, so… you’ve probably heard of, like, novelty effects, so… This is where a user explores, they get the new treatment, and they kind of explore the new experience of the product, and they test its limits, and you probably see, like, an engagement spike, but then it kind of settles down after they’re like, whatever, like. This isn’t that much different than the last thing.

So, that’s, like, a very real thing that happens in most products. You’ll have, like, a big spike in your impact. And then it’ll kind of taper down a little bit. So that’s why folks usually want to run their experiments for a few weeks. At least, like, two is usually, like, the rule of thumb for a lot of people. Depends on the seasonality of, like, how people use your product. The other reason you want to run long experiments, right, is to get higher sample size, so the longer you run it, the more observations you have. The new thing with AI, I think, why you want to run them longer, is there’s a deeper issue around users will actually change how they interact with the feature.

Like, in the case of a data AI analyst, they’ll change what they’re actually asking, the inputs, not just how often. So it’s not just that, like. We run it for 2 weeks, and we capture all this production data of, like, these are all the input types of questions people ask. Hey, it worked great. After those two weeks, the inputs people put in there could evolve to a totally different set of inputs that our AI analyst has never seen before, and we don’t know how it’s gonna react until it sees those inputs, since it’s all non-dimeristic. So, say, this V2 nails simple queries. And users learn to trust it. So what do they do next? Like, they’re gonna start with simple stuff, right?

they’re probably gonna start asking harder and harder questions, things that require more joins, more complex date logic, more educations, more, like, analytical thinking. So… Maybe you before had a mix of, like, 80% of queries people ask are… are simple and 20% are difficult. Maybe that totally shifts to, like, they’re 50-50 now. So the AI is actually… As it progresses and being out in the product, it’s being evaluated on harder and harder problems, or just different problems. than it was at launch.

So… I think, like, a two-week test catches kind of, like, the short-term effect, but misses how users will naturally adapt to your AI, feature that’s being rolled out, and then how that feature is going to respond to those new adaptations. So, I don’t… I don’t really have a good rule of thumb for, like, how long to run an experiment for, but I’m definitely in the camp of Man, we should probably have, like, 5% holdout group for, like, a couple months at all time with any change, just in case.

And just to understand, I mean, it might be the first experiment you run, it might be, like, run it for… a couple months, or, like, have that hold out, and try and get a vibe around, like, when did you see the change in whatever your success metric from, like, AI quality or user success, kind of taper out. I think it’s gonna be different from product to product. Maybe there’ll be some products where it’s like, oh yeah. they just… users behave the same with it, no matter how long it’s out, but I think for… at least for this case, with, like, a data analyst. I could easily see it.

I’m, like, playing with one right now, and my… I started with very simple questions, and then after it got those right, I’m like, let’s see how far I can go. I think humans are naturally like that, want to offload as much of our… this work to this thing as possible. Okay. We only have 10 minutes left, and I want to answer some questions, so I’m gonna skip this, like. exercise. But… I’ll do this summary, at least. So, for experiment design decisions, I think that changed slightly with AI. Size it for real math, not everyone gets the treatment.

The AI side is a lot noisier in terms of the, The output that’s gonna happen, and you’re gonna need to segment it to understand how it affects your users across many more parts of distribution, the average doesn’t really work anymore. All that stuff multiplies, so your sample size is probably going to be larger. Measurement matters, so… Define your metrics all up front with offline. evals, they’re gonna be unique to your use case, and then carry those through to your online evals. that’s something that’s a bit different than historic… historically traditional A-B testing, where you probably have, like, your same user value metric you measure across many experiments.

These are going to change feature to feature, or change to change. in that future. Build guardrails, there’s gonna be… More of these, so try and account for those as much as possible. And then the fourth… yeah, check segments, and then run longer. The averages are gonna hide everything. you know, plan your segments before the experiment, and I think the experiments are probably going to want to run for, like, 4 weeks or longer, or at least have some holdout. I don’t have, like, a rule of thumb about that yet, but… That’s what I’m seeing, at least in terms of, like, how I see users behaving differently over time with… when rolling out AI products.

If you want to learn more, I have, like, a free email course, like a 30-day email course. It’s, like, the bite-sized version of our Live in-person course, so… You can check that out at bit.ly… slash free slash AI eval. I’ll actually put our… Website in the chat right now. If you go to… AIanalystlab.ai. Most of these links will be in there. So… If you want to learn more about emails, but not looking to get into a full course, I would check out that, 38 email list. The first… the full course is on AIEval.ai, so it’s a… Six-week course. We cover, like, the entire end-to-end pipeline. You can check that at AIEval.ai. The whole syllabus will be there.

If you have questions, feel free to DM me on LinkedIn, or email me at sean at AIeval.ai. I can tell you more about it. And, yeah, 25% off. If you put in this promo code, you bounce 25. That’s it. We’ve got a little less than 10 minutes. Anyone have questions?

Attendee: Hey, Shane, thanks for running through all of this. I appreciate how you’re, making me think. another level deeper on this for running the test. It’s been really interesting.

Shane Butler: Nice, yeah, no problem.

Attendee: I’m curious, If you do a, like, a one-month delay in terms of, like, running your split test and it’s going out. Traditional split tests, like, that makes sense. You know, it’s gonna take time. Anytime you run a split test, there’s some time element to see. Does it make sense to go general audience? For some features. If you do, like, a 3-month delay before going out to everybody, you could also just lose product-market fit. Which is something that, like, Lovable’s talking about, like, every 3 months they have to reinvent themselves, and they’re the fastest growing company in the.

Shane Butler: Yeah.

Attendee: So, I’m just curious how you’re thinking about this trade-off between Wait until it works for everybody, and then release in, like, the traditional model, versus, if we don’t get it out there, we could just completely… Even regress product-market fit in some use cases, but at least we could, you know, be losing a potential adoption.

Shane Butler: No, yeah, it’s a really good question, because, yeah. Like you said, the pace of everything is just changing so rapidly right now. I think it depends a lot on the product, and a lot on the users, and their kind of, feelings about leveraging AI. So, for me, in my day job, like I mentioned, like, we’re working with lawyers, and… Depends on the lawyer, but there’s just, like, lawyers, I’d say, like, banking, finance, some stuff in, like, health, like, where there’s… there’s a lot more skepticism in AI, because there’s a lot more risk of things getting wrong, like.

If we, like, screw up some… medical record, or if we screw up, like, some contract for a lawyer, like, that totally breaks trust with that customer, and they’re gone. So, like. That probably weighs a little heavier than product-market fit for them, but if there’s a user base that’s, like. you know, like, AI evangelists, like… like a product like Lovable, where it’s like, hey, people are… they love it, like, they’re in here because they just want to, like, see what this thing can do, then I think there’s a lot more, like. wiggle room. I think in that case, I would… I’d probably, like, just have… so, so… so not… I don’t have, like, a scale or something of, like, oh, based on this risk.

Should you, like, rent it for this long based on this risk? Should you rent it for that long? I’d probably… my recommendation would be, like, yeah, if you’re low risk, you can run it shorter, but maybe, like, have some holdout group, like I mentioned, just to see… just to understand with the users if there is some sort of, like. weird, long-term quality… Regression. So, because if you do see that, then, like, the competitors are probably going to be experiencing that, too. And so, if you, you know, you could roll it out to 50-50, For 2 weeks, and then you could pull it back to, like, 95.5, and then see how that 5% runs for another 2 months.

If there’s no weird regression in the quality metrics, then, then yeah, then you could just keep running two 2-week tests instead. But if there is, I don’t know, my angle might be, like. publicizing that kind of, like, hey, like, our competitors are doing XYZ, like, how do you think about it? After two months, we tried that, and we see that after 2 months, quality degrades as, like, things hit some complexity. That might be one angle to go, but I honestly haven’t, like. thought about that as much for my day job, just because, like, we’re more trying to build, like, for us, like, the trust space, I think, is a lot more important. Does that answer your question? Kind of.

Attendee: I think… I think I’m hearing you say, first you’re… you’re looking at product-solution fit, because it actually has to do the job. And then, with that getting out, looking at a higher risk industry, working with lawyers, health, etc, You need to confirm that it keeps working. And to the extent that it’s working, the best case is you work with a small pilot group first, and then it can go out. It at least saves, like, tighter feedback loops. And then… For that extended window before you keep going wider. Are you… are you, like… if something breaks, then that’s, like, what you’re looking for. Like, how long do we want to wait to see if something breaks? If nothing breaks, then we release it?

Or it’s like… We identified all of these issues, and then it’s just, how long does it take for us to fix them? I’m sure there’s a little bit of both, but…

Shane Butler: Yeah, I think it was kind of, like, the former more, like, having some… Having some small held-out group that’s not released a new thing, just to see, like, for your specific product. Do you… is there… evidence that users change their behavior over time after the rollout to the point where something will break, that you need to account for later on. If there is. Then, next time… When you’re developing the feature offline, you might be able to, like, identify, like, hey, we know Last time we rolled out this feature, there was these, like.

long time… longer-term native impacts when people stress test it at the edges of the capability, so we need to, like, actually stress test the edges of the capability a lot more when we’re doing the initial kind of, like, offline tests. That might be one way to protect it against it, but… I think early on when you’re experimenting, just trying to, like, have some… holdout group, so you can, like, make that measurement. is pretty helpful.

Otherwise, because otherwise, it’s really hard to, like… tease out after you’ve had something out for 2 months, if you see, like, quality metrics going down, it’s like, oh, is it because of this, or it’s something, you know, the other dozen things that got released in the past 2 months?

Attendee: Cool, I appreciate it.

Shane Butler: Yeah, no, no problem. What else? We got a couple more minutes. Attendee, I like those guitars in the background. Those are sweet.

Attendee: Never fails to get a response.

Shane Butler: Yeah, I bet… that’s why they’re there, huh? Conversation starter.

Attendee: Not really, it’s a side effect.

Shane Butler: Haha, thanks. I’ll see in the chat, if you would do it all over again, how would you approach setting up experiment systems frameworks from zero? Yeah, I think, something that we set up… That we didn’t set up initially, actually. This is kind of interesting. So… We ran a bunch of offline evals, we were pretty confident in them. And then we’re releasing a product to our users. But, like, as we’ve kind of talked throughout this, like. Going offline to online can be, like, drastically different. Experiences with, like, what happens with quality. And so there’s still a lot of… there’s still a lot of, like, fear from us of, like, do we want to really, like, release this to a bunch of people?

Like, even when we beta tested it with, like, some hand raisers, it was like. it’s hard to really get the vibe from, like, a few people, especially people who are, like, raising their hand to give feedback. Maybe they’re, like, very positive people, or maybe they’re raising their hand because they want to get negative feedback or something. So that’s just, like, biased feedback. That’s hard to… parse out. What we ended up doing… was we, like, created another thing we called, like, shadow mode, or, like, it’s, like, background mode, so basically. You can run your current, like… output from whatever your current V1 is.

And all the users see that on production data, and then in tandem, in real time, be running the, running the V2 in the background, but, like, don’t actually service any of that to the user. But basically mimic whatever’s happening in product, as if… with the V1 model, as if it was going to be V2, and then send all of that data And store it to your data warehouse or whatever, and you can do, like. basically your experiment on that. It’s not gonna be… As good as, like… you’re not going to get, like, the user value task success stuff from that, right?

Because they’re not actually going to see it, so it’s not going to affect their task success, but it will get you, like, a much more realistic view on all these, like, online product metrics around quality or some of those, like, code-based ones. For us, I say, well, what would I do if I, like, started over with that is because we kind of didn’t think about that until we were, like, against… until we were, like, ready to roll it out, and they were like, oh, like, okay, let’s build this thing out, and then that took us, you know, a few more weeks or something to build out. So, that’s something I’d think about building out early. Cool.

Alright, I gotta jump to another meeting, but, Yeah, no problem. I’ll send out the… Recording and the deck. tonight or tomorrow, I’ll email it to everyone here. And then we’ll put it on our YouTube and stuff, too. And… yeah, feel free to connect with me on LinkedIn. check out that website, AIanalystlab.ai, if you want to learn more about some of the other free workshops we’re putting on, some other stuff we’re working on. We’re working on a pretty cool open source project right now, this, like, AI Analyst. project, in Cloud Code. We’re probably gonna share that out next week for anyone else to play with, so, yeah, I don’t know.

You stay up to date with that, follow me on LinkedIn or something. But thanks for the questions, and thanks for staying… hanging out for an hour. Talk to y’all later, hopefully.

Free, every week

The next one is this Wednesday.

10 AM Pacific, live on Maven. One topic a week. Bring a question from your own work.

WED SEP 30
Ace Analytics Interviews with AI
Register
WED OCT 7
Metrics 101: Define a North Star with AI
Register
WED OCT 14
Build a Semantic Layer So AI Defines Your Metrics
Register
WED OCT 21
Experimentation 101: Run an A/B Test with AI
Register
WED OCT 28
Trust Your AI Analytics: Know When the Number Is Right
Register

Next cohorts start Oct 19 and Nov 2.

AI Analytics for Everyone
$1,800 · Oct 19 · ★ 4.9/5
Enroll on Maven
Agentic Analytics: Build an AI Analyst
$2,500 · Nov 2 · ★ 4.9/5
Enroll on Maven
Or come to a free workshop this Wednesday. Register free