← All free workshops
Free workshop · Friday, March 20, 2026

Run and Analyze Experiments with Claude Code

Live on Maven, Wednesdays at 10 AM Pacific. About 70 minutes.

Transcript

Auto-transcribed from the live session and lightly cleaned. Attendee names are removed; their questions are kept.

Sravya Madipalli: Hello? Hello, everyone! I see most of you all joining now from the waiting room. Hello! Hi! It’s great to see you all. I think most, most of the people will start joining as well. We could probably start with introductions, from the start. I am Stravio Maripali, I lead Growth Data Science Team at Superhuman. It was previously Grammarly, it’s now… we are now Superhuman. I have experience at Microsoft. and a bunch of places, like at eBay and then next door. That’s where Shane, Hi, and I met, actually. So, Shane and Hi are part of my team, and, Shane and Hi, do you want to share a bit of intro about yourself, too?

Hai Guan: Yeah, yeah. So, hey everyone, my name is Hai, and I’m the head of data at a legal tech AI company called ENTRE, and previously, worked at big tech in Nextdoor, LinkedIn, Pinterest, Meta, so on and so forth, leading data science teams. So, really happy to be here and, speaking to you all.

Shane Butler: Hey everyone, I’m Shane. I’m a principal data scientist. a lot of the work I do is around AI evaluation, AI analytics, yeah, been in data science for about 10 years. One thing I like to do in the beginning of these things… I’m located in South Lake Tahoe, California. Where is everyone coming in from? Because usually we have a pretty… International, drop it in the chat, wait for me, Northern Utah, India.

Sravya Madipalli: Awesome, love it!

Shane Butler: Cool, like, that’s where everyone’s from.

Sravya Madipalli: Oh my god, I can imagine from India, it’s so late there! Great you guys are interested in joining us from India as well, and a bunch of other places, pretty cool places.

Hai Guan: Attendee, I’m probably 5 minutes away from you. You’re in Pacifica Bay Area.

Sravya Madipalli: Nice, nice! So exciting, so exciting. Okay, we have a pretty packed agenda today, guys, so… as soon as… I think we’re getting most of the crew in, and probably maybe one last question about… before we get started. What is something that you’re looking forward from this session? Any… anything interesting that you’d like to share in chat? While the new people who are coming in could share about where they’re coming from. Okay, okay. Good? Yeah. Cool. Sure, that’s fine. I think, I… I’ll probably get started. Let me start sharing my screen. Yes, we’re gonna cover all of that, guys. Like, I love all the great responses. So… Can you all see my screen? Okay, I see a yes from Attendee. Cool.

So what are we going to talk about today? We’re going to talk about running and analyzing experiments with Cloud Code. Yes, I talk about Cloud Code here, but guess what? We also do Cloud Web Cloud Web UI, because I know that most of you could easily access Cloud Web, and we have some interesting things to share with you at the end as well, that you could literally take out of attending this free lesson. and use it in your day-to-day work, okay? So, let’s get started.

By the way, I’ll be presenting today, and any questions that you have about what I’m presenting, and any questions about other things that we do at AIAnalystlab.ai, please, you know, push through your questions in the chat, and we’ll have Shane and Hai reply to you as well. Cool. So, let’s just get started. Okay, let’s start with a quick poll. Can you answer to this question? You could type A, B, C, or D. No judgment at all. Based on what you all feel like. Give me a minute, I’m trying to… I can’t access the chat, so, Shane and hi, if you see something, you could, like, you know, respond to them, but…

Shane Butler: We got a lot of… A lot of C’s, a lot of high, performing experimentation.

Sravya Madipalli: Oh, that’s pretty cool!

Hai Guan: policies will…

Sravya Madipalli: Oh, nice! That’s so cool to know, actually. We generally try to cater our presentations to everyone, because what we try to do here is something that get people who are not even… not as experienced, you know, people who are stakeholders, not even data scientists, how do they get on board into understanding experimentation, and how could they, you know, get all the technical details and help from Cloud Code? So, it’s absolutely okay, even if you are, like, an A or B, so, or even D, you know? So, there’s no judgment at all. But glad to know that most of you are Cs. But let’s get started. Today, we’ll probably try to do the C that most of you all do, but how do you do it with AI? Cool.

So, let’s… I’m doing it the other way this time around. I want to show you what is the final output that Claude Cord gave me. And let’s then dive into the theory behind experimentation, because I wanted to share with you all cool things at least it was pretty cool when I tried to test it out. Let’s see if you all can see this. Okay. So, this is the prompt that we gave. So, this is our repo that we work off, pretty, like, you know, detailed. We have a bunch of agents, like, we have a bunch of things going on. And in the repo that I have my data file, this is what I gave as a prompt. Analyze experiment. And this is the CSV file that I used, which contains streamflow data, of an A-B experiment.

I’m using the experiment analyzer agent. This is the agent that we’ve built, and these are the inputs I gave it, and I gave it some context around what… business context around, like, you know, what is… like, what is the average lifetime value, and, you know, things like that. And it also has access to a bunch of, like, questions and analysis that you have to do when you are given, like, an experimentation question, right? And then, I ask it to, you know, do a final readout in a Google Doc. I literally paste this thing, I don’t do anything else. You could actually also talk to it with Whisperflow if you want to. But let’s say I pasted this, right?

I’m going to go through the entire demo with you if we have time at the end, but I wanted to show you the end result. So, I’m sharing with you a window, let me stop sharing and share with you what the output looked like. Va-da-da, there am I. Okay, so do you see my Google Doc? This is the Google Doc that it generated Streamflow Regional Playlist Experiment. This is like a demo output, as you see. It has the executive summary, it has question after question, like a table with the details. and, like, what happened to these metrics? What’s its reliability? How does it look like for deep dive segments?

And, you know, I don’t want to get through the details because there’s some interesting stuff that I want to share in the demo directly, but I wanted to share with you that this was the output it generated with one prompt. One single prompt in our repo, and that’s what we got. Isn’t that cool? What do you all think? Would love to know. What do you all think. And let me get started with my presentation back again. Cool. So, now that you saw the end result, I wanted to share with you how we got there. What are the multiple pieces that go into this, like, brain of how did it come up with that, output, what went into it, and details? So, before that, I want to share with you something.

So, if you want to build this reusable system, like one prompt in and an output out, obviously you’ll have to do some feedback and checking and a bunch of things, but if you want to understand these… how we build these agents and skills, we have an AI analyst bootcamp coming. It’s a two-day, weekend-intensive program. We have way more than experiments covered there. We also talk about funnels, we talk about root cause analysis, storytelling, like, how do you actually get this to present in a… not just a doc, Google Doc, guys, but also a deck, like a presentation. I… I have a demo of a presentation that I did as well, I’ll share with you at the end. But we cover all of that.

And for the people who are attending this lightning lesson, both live and also who have signed up, we have a code for you, Claude20. Please try to use that, and leverage this to get $120 off. This is our first attempt at doing this bootcamp. I know a bunch of my friends and all the, like, you know, people I know, they’re like, you’re giving this for way, you know, shorter amount, smaller amount than usual, but we wanted to share with you all everything that we are learning, and pretty excited to share that. The other thing that I want to share is we have an Operations repo. We have this, in GitHub, with, like, AI Analyst Lab. slash AI analyst.

We… Shane, Hi, and myself share a bunch about this in our LinkedIn posts as well. Please try to, you know, get this open source repo, clone it, and see, all the things it could do. Share us your feedback. What we’ll teach in the Analyst Bootcamp has a lot more things than what we have in the repo. We are actively building what we share for the bootcamp. Another interesting thing as well, we have a Slack community. We currently have 2 or 3 plus members, people growing by the day. We… people come in, ask questions, you could free join… free to join anytime, and we are there to help, like, you know, navigate through any challenges that you’re going through in this space.

This is a QR code to get into the bootcamp, guys. Yeah. Cool. Now, let’s get started again about the problem. So, here’s a common pattern we see with experiments. We turn it on, we wait, read the dashboard, green means ship, and red means kill. I know a bunch of you are already doing great things, when you mentioned that you’d do the entire analysis, but… I’ll be honest, from my experience, not everyone does the entire workflow. Or a life cycle of experiments need to be done. Especially when we want to ship things very fast and move fast, that’s when we see this happen a bunch as well.

So, when you do, based on just PR, green means ship, red means kill, it’s almost like traffic light watching, right? So, the problem is that treating experiments as a single step instead of a life cycle. So today, I’m going to show you the entire lifecycle, the different stages, what are the common ones that kind of get skipped, and then we’ll demo what happens when the obvious Ship 8 answer turns out to be wrong. Cool. What’s experiment lifecycle? I think a bunch of you who are data scientists that work in this space probably already know it, but for the people who are stakeholders and are newer to tech. I’d like to, like, you know, share this in deeper detail, right? So, it starts with why.

Like, why are we experimenting in the first place? And the design, how do we want to design this experiment, right? And how… what is the power? Like, how long do you want to run it, and how many users? And how do we want to, like, once you start running it, what are some things you need to check for? And then, once we reach our, you know, the goal of running it for a certain amount of time, how do we analyze it? And how do we… what do we decide, on? Do we ramp it? Kill it? Iterate it? Iterate on top of it? And then, how do we all ramp it? So, we’ll cover all of this. Let me know if I’m going too fast, guys.

Hi, Roshan, please interrupt me if there’s something in the messages that you want me to stop. I’m currently directly diving in, because there’s so much content I’d like to share.

Shane Butler: You’re good, we’re managing the chat, but I’ll interrupt you if we need to answer something directly.

Sravya Madipalli: Okay. Providers, like…

Hai Guan: the Attendee already registered for the bootcamp.

Sravya Madipalli: Oh, that’s awesome! Okay, cool. I’m kind of feeling like I’m giving a big monologue, guys, so please, you can interrupt me or share anything. So, okay, let’s get, back to this. So, why do we experiment, right? So, products go through this life cycle. We first build. And we want to measure what we build, and then we learn, right? So, experiments, basically, so this isn’t new. It’s the way the lean startup loop works. So, experiments, or the most critical ones that we are discussing today, are in the measure step. Without them.

your PM probably would say, hey, I think user wants dark mode, and your designer would say, oh, onboarding is way too long, and maybe your CEO would come and say, just change the pricing, guys, you know, this is, like, there’s a lot we could do here. So, all of those are just opinions until we go test them. So, how do we go test them? We first start with design, right? So, step one is basically design the experiment before you build anything. Start with a hypothesis. And if you notice the structure, we have something like… we have this, like, we believe this change will impact this metric because of this reason. The because forces you to articulate why you think this works.

That’s the mechanism. Then, pick one North Star metric and a guardrail. The North Star is what you’re optimizing for from this experiment, and the guardrail is what you refuse to break. And I want to give a quick plug-in about… we did a free Lightning lesson on how do you define metrics on this. Hi actually ran it a few weeks ago, so if you’re interested, please drop in the chat and, you know, Shane and Hai could respond. Also, you could go to AIAnalystlab.ai. There, you could see a bunch of free lessons that we’ve, you know, dealt, we’ve shared in this area, and a bunch of things that are coming in future as well. Okay, now coming back.

Here, in this lightning lesson today, we’re going to have a running example of a streamflow dataset. So, what is the experiment that we’re going to have? That we’ll… that will run through this entire Lightning lesson. We believe showing regional playlists will increase weekly streams because users engage more with music that reflects their local culture. So, that’s the hypothesis. See, if you look at it, we have a because, right? We think because of that, it reflects the local culture, and there’ll be more people who’ll engage with it. So, this is our statement, and this is our experiment that we’ll run with.

And the North Star that we are aiming to improve is streams per user, and the guardrail is churn rate. We don’t want people to churn, right? We want them to engage, so we want to ensure that that guardrail is met. Cool. Now, now that we have the statement, the hypothesis, the because, the reason, we need to understand how long do we want to run this experiment for, right? And how many users do we want this experiment to be, you know, exposed to? So. What is the effect size? So, is 10% lift easy to detect? So, if you look at it, 10-person lift probably is slightly easier to detect when compared to the 1-person lift. 1-person lift in, like, control to treatment would need way more data, right?

Specifically, for a significant level, it’s usually around 5%, that’s an industry norm, and power usually is around 80%, that’s another norm that we deal with. So, for this analysis, after the power analysis, what we came up with is 6 weeks and 25K users. So, what happens, let’s say if we skip this. Most of good teams don’t skip it, but there are a bunch of teams that I’ve seen skip it, where they think they have this knowledge already from the past experiments, they have the domain hypothesis, and then they want to move fast. Let’s say they’re blocked on the data scientist, they skip these things, and what happens when you skip them is you run too short, and you miss the real effects.

And there’s a bunch of peaking, also, that happens. Let’s say on day 3, someone sees a significance, and it’s such big… it’s such big of an effect, they’re like, oh, let… why do we wait on getting the impact on this experiment? So let’s ship it. And then they ship noise, because it’s just day 3, right? and… And another thing also happens when you don’t do the power analysis and all of this the right way. You basically can’t tell if a null result or a null means actually there’s no effect. Or, it doesn’t mean you have… you don’t have enough data to talk about it. Basically, there’s not enough power to say you know something, you learned something, right?

And, like, probably the data science folk in this meeting would know that peaking is, like, the most common way good experiments get ruined. Okay, so this is a small sneak peek about what our existing free repo already has. So this is an experiment design pipeline that we have. It basically, you know, comes up with a hypothesis, it, you could… the power calculation, and what exactly happens in the design, and what are the guardrails, and you know, running happens right after this. So, every good experiment follows this structure. The system auto-generates all of it, once you give it enough context. Hypothesis, power analysis, randomization of design, guardrails.

And the full pipeline, including the distant rules, also is part of, like, the next stages, right? So today, we are picking up after, once the design is done, how would the experiment run? So that would be part of my demo, but I want to share you a sneak peek of what would happen what would be if you run it, the design skill that we have in the design repo today. So these are literally the charts that it created for the experiment that we are going to talk about today. So, on the left, you have a feasibility curve. What does it show? It basically shows you how many weeks you need to detect different effect sizes. A 5% lift needs around 6 weeks. A 3% lift, we need 16 weeks.

Wow, that’s in the patient zone, right? And on the right, you have a power sensitivity matrix. Each cell shows you… shows, basically, your statistical power at different combinations of FX size and sample size. The green boundary is where your test becomes reliable. So, 80% are power or above, basically. So, notice that at 12.5K users per arm and a 5% effect, we hit exactly 80%. That’s why the streamflow experiment, if you remember that I gave you this number before in the design, basically talks about, having the experiment run for, 25K users over 6 weeks. And all these, to mention it again, these charts are auto-generated by the system that we have. Cool.

Yeah, so now let’s talk about running and deciding. So, what do we want to do once, like, you know, we basically are condensing two steps in the stages here to explain you the theory. We’ll go through this in detail if we have time in the demo. So, we are condensing two stages here. While the experiment runs, check for the SRM, right? Sample ratio mismatch. Ease the randomization clean? Or not. You probably need to watch out for the novelty effects. Early lift, because of, you know, it’s an exciting new feature, probably, right? So, a bunch of clicks come from that, but not all of that stays the same. So, early lift. fades if it’s not real.

And when it’s time to decide, you ship, kill, or iterate it, right? And then, remember that the statistical significance does not always mean practical significance. Basically, a tiny lift that’s significant for us, it feels like experiment is a win, but it not… it might not be worth shipping, because it adds complexity. This is something I deal with day in and day out in my current job, in my previous… in all my experience, I’m sure a bunch of you who are from data science also deal with it a bunch. That what is… you know, good enough to ship, even though if it’s, like, a VIN.

And when you do ship, not every experiment has to go through this 5%, 25%, 50%, 100%, but the ones that affect your ecosystem of product, if it changes a bunch, you probably go through that. But another important thing that I see a bunch of, you know, teams miss is having a holdout. So… ensuring that you ramp gradually, especially in case of important, like, you know, UI changes, or… something that affects users. You basically try to ramp gradually, and you always keep a holdout. So, the experiment doesn’t end when you just say, ship it, it doesn’t mean that you’re done with it. Yeah, so this is what covers what I just mentioned.

So you monitor it at each stage, especially for the important changes. You basically ramp gradually, and you monitor, right? And you understand, and you also add a holdout to understand the long-term effects. There are, like, a bunch of experiments that actually show great results, probably at a later stage. For example, I have experience in working with engagement of the product. Let’s say top of funnel is engagement, and bottom of funnel is revenue that you get from these users. Unless you have, like, metrics like lifetime value, which predict the revenue that comes out of these users, it’s pretty hard for you to see engagement show an impacted revenue in the experiment timeframe.

But guess what? If you add all these engagement-driven experiments that are wins to your holdout. What you’d see, you know, over 6 months of time is that you actually see a great improvement in revenue coming through. And imagine if you have a holdout, for, like, a year or so, let’s say in, like, a… in, like, a scenario of you have it for 5 years, for some, and to understand, I’m… the long-term effects of improving engagement is going to be so much higher on your revenue, but you cannot understand that in your 2-week, 3-week, or maybe 4- or 5-week experiment standpoint. And that’s the reason it’s very important to have a holdout.

Shane Butler: Hey, Stravia, sorry, can I interrupt real quick? Is the, so I know the design… the design experiment skill is in the repo. Is the, full… Experiment workflow in the repo as well, or is that something you’ve been iterating on your… on your own?

Sravya Madipalli: So, the full experiment thing is something I’m iterating on. I’m planning to add it for the bootcamp. So currently we have the design one, though. So, you could actually play with it yourself and see what else, you know, needs to be done to it. So, the repo itself is pretty smart and intelligent. It has a bunch of things in itself. So, Yeah, I’ll probably talk about it more when I go through the entire, like, demo. Hopefully we have time. I’m trying to go super fast to get to the demo. We’ll see. Awesome. Yeah. Thanks, Shane. Yeah, please interrupt me like this. Cool.

Okay, so these are, like, a bunch of… these are, like, 8 questions that I see, like, you know, separate, like, a good analysis from great. So, not every experiment needs all aid, I’ll be honest, but knowing what to ask means you won’t miss those things that matter, right? So, the… one thing is basically understanding if the experiment’s set up correctly. Did the treatment move the metric? Like, what is the reliability? What are the… so one thing that I see people miss when they’re, you know, in a hurry to ship something before the quarter ends, because they want, like, wins out, is what are the learnings that you have from the experiment across segments?

These are something that you wouldn’t know unless you dig deeper. into understanding, oh, why… what happened for probably users that are newer when compared to the mature users, you know? Like, asking these questions to look through multiple segments would really help you, not just to Not just for that experiment, related, you know, decision making, but also all the new iterations that you want from your experiments, like, that you have in the quarter, that are in the backlog that you want to run with your product team. Product dev and engineering team. Okay, so, today, we basically are going through all of these through the Streamflow dataset.

I’ll actually give you exact prompts to do it yourself in Cloud Web UI. And you could… we’ll have, like, a one-stop shop prompt, and we’ll have, like, multiple prompts for every… each of these stages as well. Okay, time for a live demo. Cool. So, let me just stop sharing here and share my screen and get started. Where am I? Okay. Give me one minute… If you’re trying to… okay. Let’s go back to the demo. Let’s see… we have our Cloud UI right here, okay? And what you’re trying to do here is… I have one that I already ran it. If for some reason we’re gonna have issues with my internet, or, I don’t know, Claude usage, we can look back at what I ran it. But let’s do this live, guys.

We have some time. So, I have the dataset. This is the dataset we’ve generated for this. And I’m gonna give this prompt. So… Okay, so what am I doing here? I basically gave it a prompt that I’ve uploaded a dataset, I gave it the context of what the dataset is, I gave it what the columns are, what each column’s… almost like a data dictionary, right? What… each columns are, and I tell it, like, you know, what should be the step one be? Like, what do you look for? Like, sample review of mismatch, the pre-experiment balance, and the duration check. What’s the date range, how many weeks you want to run it, and stuff. Let’s see what it comes up with.

One thing I want to let you all know is this is something that I’m facing right now in my company when I’m trying to do things. the column names, so I didn’t give it what are the re… what are, like, each of the columns mean here, right? for example. if you don’t give, like, you don’t have a data dictionary of what each column means, what Claude or, like, maybe any other, any other, like, you know, LLM does is it understands based on what a column name looks like. So, we basically, had this issue where let’s say there’s a column name called Anonymous.

It takes that into context of what it knows, and just takes it as, like, comes up with its own assumption of what a column name could mean, and goes and runs with the data. So be very vigilant when you give column names without what the column names mean. In this example, it’s pretty obvious, I kind of created the dataset for the simplicity. I made it so that each column names actually mean what the name, could, you know, what you could assume based on the name. But if you’re trying to do this at your work. Ensure that you also give what each column actually means. Because not all column names are pretty descriptive. Now, let’s see. What did it give me? So, it went through the step one.

The sample ratio mismatch, it did, it basically looked at the split, and it gave me a pass, the verdict. It’s pretty cool. No evidence of a randomization or a logging bug. And then it basically looked for continuous covariates. It came up with what are the categorical covariates, and, like, proportions are virtually identical across arms. It made sure that what we are looking at is a pretty balanced dataset. And then, it came up with a duration check. The experiment spans for 6-week cohorts, and it gave me from when to when, and, like, it also identified a small concern about the weeks 5 and 6.

And… basically, this is what it says, like, it’s something that, this is, like, worth investigating in Step 2, so that’s what we’ll go ahead and do in the Step 2. How much time? Okay, we have 30 more minutes. I’ll try to do the Claude web UI, like, demo for the next 5-10 minutes, and I’ll run in another demo and go through how my, you know, workflow looks, what is something that I’ve built for this Lightning lesson for experimentation, in the Cloud code after this. Cool. So it’s currently thinking through what, basically my prompt is that I’m asking it to do the treatment effect and also the reliability. I’m telling it what test to run.

what is the experimentation t-test I want to run, and, you know, how I want it to be displayed, and I also talk about the effect size, can you give me the effect size? And… look, and also I talk about, like, what is the post hoc power? Like, do you give the observed effect size and sample sizes? What is the power we achieved? I also talk about minimum detectable effect. What’s the smallest lift we could have detected at 80% power? And I also ask it to give me a one-sentence summary. look at what it’s giving me. Exactly what, I mentioned to you guys, that I asked for. And yeah, so look at the one-sentence summary, the treatment increased by 14%, right? Which is statistically significant.

So, guess what? Looks like we have a winner here, guys. We have our… not… like, we have the metric increased by 14%, which is amazing. So, there is something that you would see that most of the people would probably go… there are a good chunk of people that would probably tell and announce that, hey, we have a winner, because we ran it for the amount of time we wanted to run it, so they did their diligence, they didn’t peak, probably. And they got a winner. Let’s look at the step 3, and do the segment analysis, and think through… And look through if what we have is actually a winner or not.

So, this is a longer one, because I’m asking it to look through and break results by all my important dimensions. So, the dimensions here are going to be user type, if it’s an existing user or a new user, and I give it 5 values for the… so the region has 5 values. And we also have device, mobile, mobile, or desktop. And I ask it, for each segment, do these same tests, do… can you check for Simpson’s paradox, can you check, do the guardrail checks, and for any degraded segment, look for our guardrail, which is the churn, or not. Let’s see what it comes up with. It probably might take a bunch of time, because this, I would say, is, like, the more detailed… oh, this is pretty fast. Nice.

So, look at what we have. So, the lift is brought across most segments. But, guess what? It found the guardrail to be degraded for new users. And the trade… but the… the trade-off is that existing user churn dropped from them, but not as much. But for the regions, it’s pretty fine. So… The churn treatment users are essentially the same as churn control, so we’re not loosely… losing, like, heavy or lightweight streamers, right? So… and the bottom line is that the regional playlist feature is a clear win for existing users. But is a retention risk for new users. The exact population streamflow needs to grow.

So, if… imagine we went to step 2, and we saw this win, we… and people didn’t bother to look through segments of, you know, like, the critical user types, we wouldn’t have understood that this is what the… this is what was the issue, right? So, for the new users, it was a retention risk. So, now let’s dive deep into what… now that we know this, what do we want to do? We basically, if you remember, going back to the theory that I talked to you about, it’s not just statistical significance, but also practical significance. We need to understand the ROI. What is the impact of running this experiment, and what is the, you know, return that we’re going to get with it? So, let’s do the step 4.

How much time do we have? Okay. I know I’m running fast. Anything, Shane and hi you want to share based on what’s happening in the chat while this runs?

Hai Guan: I think you’re good to go. There’s a lot of questions, they’re very desperate. We’ll get to the things that we can’t answer in chat in a bit in the Q&A.

Sravya Madipalli: Oh, makes sense. Awesome. Sorry, guys, I’m trying to run as fast as I can, so that I could go through two demos with you, the Cloud Web UI demo, and also the Cloud Code demo. I want to share all those details, and also, if possible, have at least 15 minutes of Q&A at the end, which means we have 10, 11 minutes now. Cool. If this is going to take time, I actually have this run for us. How does this look like? Step 4. Let’s see if this is still running.

Yeah, and another important thing that I wanted to share with you all is that if there are audience, like, who are listening to us today, if you do not understand the details of this, because I saw a bunch of non-tech folk, people who are in, like, you know, our stakeholders, from design, from product, also join. We have, another course called AI Analytics for Builders. That’s basically a course where we’ll go through… it’s not a weekend bootcamp, it’s going to be a 6-week long course, where we talk through every detail of how do we set up experiments, how… what are the metrics, and all of that with Claude Code.

So, almost like an enhanced version of the bootcamp, but that covers 6 weeks long. Yeah, you could look at all of that in AIAnalistLab.ai as well. Okay, so… it’s actually creating the chart. So, when I ran it to make sure that I have something for you guys, I didn’t do the charts here. So, for the demo, I was thinking it’ll be cool to get a chart, and we did get a chart, that’s pretty awesome. So, this is something, literally, I’m doing right now with you guys. Because I was like, let’s have a thing with the chart. So… I think this is something you can copy-paste, yeah, copy image, how cool is that? Okay. So, what does this say?

So… It basically says that the churn feature reduced churn among existing users, saving, like. And then, 53,700 users, worth $11.8 million in LTB, but if we shipped to new users, the chance spike would burn, like. So many. So yeah, this is what we would understand, what happens when we look at things by segments, right? And also understand the lifetime value and ROI here. So the count… so the conservative scenario talks about deliberately zeroes out the churn benefit, and still, like, still returns 21X the infrastructure cost. So, even under the most cautious assumptions, this is a strong share for existing users. Yep. Okay, so let’s go to the step 5, which is the final prompt that we have.

So, I’m doing it step after step, so that you go through this process with me. in for multiple questions, but we could have one single, one-stop prompt as well, that does this entire thing for you, and gives you the docs and, you know, give you the deck and the doc as well. So, when we share these results with you over an email, we will basically share with you The… every step-by-step prompt, and we’ll also share with you the detailed, like, you know, one-stop shop prompt as well. Yeah. Look at this! It basically, talks about, per segment trip decisions, what do we… what do we ship, what do we not ship, what do we iterate on, what do we ship?

I… I’ll be honest, generally, like, in real time, it wouldn’t be as feasible for you to ship, like, specific areas, for example, the Pacific region, right? So you want to do something at the cost of not maintaining it, so this is where I would say You would use this… agentic system to get to you to a place where it does 80%, probably 85-90% of your work, but there is definitely 10, 15, 20% of the work that you still need to do, because not Unless and until you give it the entire context, which probably the more you work with it, the more context it gets, and the better it gets at giving you these decisions.

But there could be some reasons that, hey, shipping it to non-Pacific regions is something not possible for infrastructure. We possibly want to, you know, do the entire ship. and then iterate on it, right? Because we don’t… we might not want, like, super custom product features as well. So, that is something that you, as a, you know, a data person or a product person or, like, you know, stakeholder would make a call, based on all the information that this provides. So yeah, look at the net annual impact, the retention LTV saved, and the net ROI. So this is the executive summary.

Streamflow’s regional playlist experiment is a clear win, with one critical caveat, and it gives, like, all the details that I shared in the earlier prompts. So, in this prompt, I actually did not share it to give me a Google Doc, or, like, a doc, but I did in my previous prompt. So, look at this. I actually asked it to give me a doc, and Claude Webb generated a doc, guys. It’s not a… I thought it was pretty cool. Look at this! It’s actually a very… and it’s something that you could download, it’s in a docs that you can edit and share. So, I will share this other prompt as well.

So, I wanted to share a more detailed prompt, and hence I did this for table structure, but it could also give you an actual doc that is, you know, that looks actually pretty good, directly ship-worthy. I’m sure you have to do a bunch of changes. Based on what you think is the right context, but it’s in a pretty good place already. Okay, so let me stop sharing this and go for Claude Code. prompt. called Cloud Code Demo, okay? Anything quickly, hi or Shane, that you’d want to tell? Before I jump to the Cloud Code…

Hai Guan: I’m gonna paste in some upcoming, lessons as well that might be interesting to this group, as you’re pulling up Claude Code, Sravya. Yeah. It would be all free, and they happen over the next, week or two, and you know, it’s also about Cloud Code, and just drop it in there. Definitely register, and you’ll get the recording afterwards.

Sravya Madipalli: Awesome. I’m generally not seeing any messages, guys, so sorry if you have questions for me, but I quickly saw one message before I’m sharing my screen about, did Claude have context about LTV? Yes, it did. If I didn’t give it the context about the ROI, like, you know, the ROI or average LTVs, it would, you know, take a bunch of things itself. and create scenarios. If your average delivery was so much, this is what you would get. If it was so much, this is what you would get. But I, for this demo, I gave it the context. Okay, now let’s go to Cloud Cool. Cool. So, for the people who were at the start of my, of the lightning lesson today, you might have seen me, given this context.

to this, to this cloud code, I… sorry, give this prompt to the cloud code, and we have a bunch of reports, like, you know, all our bootcamp repos here, all our AI analytics for Builders, the course that we have, we have all of this, like, in-depth knowledge that it could use to get this information. So, let’s say this is the analysis… this is the prompt, I’m running… I’m running it again, so that… Oh, give me a minute, I actually want to run it and ask it to give me something. So, it has the prompt to give me the entire, and I want to ask it, can you… Share with me a plan. Before you execute. Let’s see what this is gonna give me. So… It’s basically going to look through all of this.

So I just gave it a permission. You could also give it, like, you know, go to settings and permissions and change that, guys, to make sure it doesn’t ask you for permission every time, but I tried to do it for some things, because I want to be a gatekeeper for some permissions. look at the prompt, look at its plan that it’s generating. So, everything that I showed you, in the Cloud Web UI, it has all of this In its, like, you know, in its bank on how to generate this analysis. I’ll go through this, but I wanted you to check out the prompt that I gave. So, analyze experiment in Streamflow experiment.csv using the experiment analyzer agent.

So, this is the agent I created to showcase with you all how do we do experiment analysis in Cloud Code. and the inputs, and all of this, and I also talked to it about using another agent at the end, then export the final readout as a Google Doc using Google Doc Creator agent. And also, I talked to it about, after the 8-question analysis. it completes, pass the results to experiment readout agent. So I have a bunch of agents and skills that are inbuilt in my system that it uses to come up with an answer to this question I asked. And this is the business context that I gave, guys. Like, someone asked this question, I briefly saw. Okay, so it actually goes through this entire prompt.

This is the plan that it created, this is the data overview, this is the stage one, like, what it does, like, all the steps, the questions it goes through. it actually creates an output for each of these in this, so if you’re interested, you could go to this working direct… you could go to the directory, like, I have my Lightning Lesson 5 here, so it puts all of my stuff here. And Stage 2, it does the experiment readout, it does this, then it does the Google… in Stage 3, it does the Google Doc export. Let’s do something to make it easy for you guys to understand. I’ll ask it… Let me actually talk to it. Can you give me an ASCII diagram of your entire workflow?

So, this is what I do most of the times, I literally talk to it, not even type. with VisperFlow. So… This is what… It is doing. It takes all this… Knowledge that we built into the experiment analyzer agent. goes through all these questions that I put as part of the agent, when to do what. for each combo, like, what do you do, when do you pass, when do you block or halt, and… like, because you don’t want it to run through all of this when you have issues with the SRM itself, right? So, there’s… this is… the stage two. Once we have all of these, how do we do the readout? So, ensure that you generate visualizations. If you see, we have a bunch of helper files that we have.

A good chunk of this, I would say 60% of this, or 50% of this, is already part of our free repo. So please check out the free repo, it’s already in, like, something that we… that Shane and I are sharing, and also it’s an AIanalystlab.ai as well. But I did create a bunch of agents for this lesson. Specifically, I worked over the last weekend for this, to basically understand how do you get these questions, but also do the readout in a way that you could literally share it with your stakeholder, right? So it reads and synthesizes, it’s designed as the storyboard. How do you take the context? What do you… what is the tension? And what’s the resolution?

This… this narrative is something that we built into our skills and agents. And we have different charts, and we write the narrative, we build a ramp plan, and the final one is the Google Doc. as simple as we think, Google Doc is not… I mean, once we have an agent in scale, it’s probably easy, but it’s not, it was not super easy the first time you did it, because it does not have as, easy access to where an image sits and where, like, the text starts, so you need to work with it to ensure the images and text have, you know, a gap with each other and the formatting is right. So, this is what it is, and let me share with you the final output once again. Yep. What am I sharing?

Oh, wait, I’m sharing the wrong one. Let me share the right one. Yeah. Okay. Cool. So, this was the doc that it shared. If you look at what we’ve done in the Cloud Web UI, It’s literally the same thing in a Google Doc format. I did Google Docs because most of the tech companies I know in Silicon Valley use Google, like, you know, Google Workspace. So, it basically did this, it did all the formatting, the colors, the, you know, the original one was not that pretty. It made it pretty, because we wiped with it to make it pretty. The charts, like, what exactly happened with new users, and what’s in Pacific region. And there’s a Simpsons Paradox, that’s what it talks about.

And these are the charts for the guardrail alert. And was the experiment running long enough, and what is the ROI impact? And it talks about the existing users shipping, and the new users do not ship. And another cool thing is, it actually created the deck, too. I didn’t get a, good enough time to work on work. So this, we also have a Google Slide deck. export, agent, I’m also a reviewer agent, but I don’t have as, as much, as many interesting things on top of it yet. That’s something I’m gonna work. Over this next week. But yeah, this is already what it generated, not too bad.

it basically talks about what exp… like, all validity checks being passed, what’s the primary metric, and how does the segment reveal. So look at the segment reveal. So it just did, like, it had some narrative built in. It talks about Simpson’s Paradox, like, you know, what it gave. And lift by region, churn guardrail, like, weekly lift trend, And finally, the business impact. And it also gives recommendation. Again. All of this is nothing. I didn’t do any code, I didn’t do anything. In fact, I spoke to Claude Code. Because we have the system built in, the agents and skills, that’s the reason why, like, we got to this place. Okay. I think that’s where my demo ends, guys.

We have around 11 minutes, we have a hard stop at 1, unfortunately, so 11 minutes to take questions.

Hai Guan: Hey, Attendee. Well, good to see you, Attendee.

Attendee: Hello. Quick question about, like. accuracy? You seem to be, you know, iterating through this very quickly and just trusting it. like, if, let’s say, you were responsible to, you know, give a report to the CEO, and your job was on the line, how would you change the workflow?

Shane Butler: Second one.

Attendee: If you weren’t giving a demo.

Shane Butler: I can take it. I can take it, Savia. Yeah. So, here’s my philosophy around anyone using cloud code for analysis. You’re accountable. The human is accountable for the results of AI. No one gets to go to the CEO and say, oh, the AI screwed it up. just like any analyst, like, if I gave my results to my manager, and he reported that out without taking a look at it, like. that’s on him, too. So there’s many ways to validate this. One thing I like to do is go through and run it on a bunch of backtests of analyses I’ve already ran in the past, see… you can just even, like, one-off do this yourself.

I have it, basically also take any of the Python and SQL, and I have it put it into a Notion document with the results, so I can validate that there too. You can also build, like, extensive AI evaluation, kind of metrics around the validity of this, too. So if you have, like. 300 ground truth experiments you’ve ran in the past. whatever, year or so, or two years. You can run this on there and do some sort of, like, auto-research thing to optimize the, like. performance of that in terms of, like, how accurate the results of this were versus, like, the human ones back there. So definitely, like, there’s ways to have structured validation, like any AI product.

But at the end of the day, I think, like, you have a very good point, like, don’t just trust it and give it to the CEO, like, because it is your job, and you gotta, like, be accountable for that at this point. Yeah. Does that make sense?

Sravya Madipalli: One thing I’d like to add to what Shane said, Attendee, is that what’s the intent? Like, what’s the goal of this lightning lesson, right? The lightning lesson goal is not that hey, take this AI agenting system, and, you know, you don’t have to work anymore, or, you know, or just you let it work, and you give it directly the output, not at all. What, the goal is It is pretty smart and intelligent, and we could make it to half, or way more than half of the work, that you basically let it do the work, and you decide on what you like, or what you have problems with it, and work with it, give it the feedback, and let it run. I agree, Attendee, this was a demo.

I know it… I probably did give it, like, something that I could share with you all, but… That does not mean that this is always going to do some great things, right? There is a bunch of things… I’m running it in my company right now, and there are so many things that it… also gets wrong, but guess what? The more I work with it, the more I get the skills and agents to a good place, the better it gets in the next iteration. And I’ll be honest, I’m banking on the newer model that’s going to come out in, like, maybe 2-3 months, and I don’t probably need to work with it as much as I worked with it now to get to a good place that I could present my findings to CEO directly, right? Yeah. So Attendee?

Sorry, one thing. Attendee, anything else you’d like to ask, or are we good? Or we could move?

Attendee: That’s good for now, thank you.

Sravya Madipalli: Okay, awesome. Yeah.

Attendee: Yeah, Shabi, I had a quick question. So, could you maybe share which agent is doing what work in the different steps of the experimentation? Because I saw you had, like, multiple agents, right? So, if we could get a quick walkthrough of which is tasked with doing what in the entire cloud code process?

Sravya Madipalli: Yeah, we only have 6 minutes, so I’m wondering if I start doing that, it’ll probably take the entire time. What I can do is, probably in a follow-up, after this lesson, I can try to give some, you know, summary, so that you could learn from it. And, I’ll also be honest, I’m not sure if it could cover all of that in a summary in, like, even in the… even if I take 6 minutes. So that’s the reason we plan to do all of that in a lot of detail in the bootcamp, and also the 6-week long course, where we get into a lot of depth. So, if you’re interested, please check them out, but we’ll also try to give you stuff, for free, or in links as much as possible, okay? Awesome.

Hai Guan: Travia, there’s a really good question from Attendee on what’s your perspective on someone junior who hasn’t yet built the knowledge or skills to fully validate outputs themselves? Do you think the level of reliance on Claude should be limited, or more… should more emphasis be placed on learning through feedback from others?

Sravya Madipalli: Oh, I love it. It’s… I think this is all the problems that we are dealing with as data leaders in this space. Like, this is something that I’m thinking through literally myself. I have a team of, like, you know, 12 to 15 people, like, and… I am… it’s the problem of the hour that we are trying to solve, with evolving AI. To me. I think you’ve had it in your question itself.

I would have… accountability on area owners, or metric owners, or, like, the tech leads of the space, to ensure that the skills and agents and the context that’s built into it is reliable, is up-to-date, and is something that people could, you know, check And review with the lead before they output something, especially the younger, the newer ones. Until we get it to a place where, you know, the… let’s say we get it to a place where 95-99% of the reviews always are passing with the tech leads, or with the area leads of that area, of that metric, then yes. Maybe we are getting into the newer system or a new model, where it’s almost getting it right most of the time with the context we provide.

But, that said, it is also very interesting that the… recent… the newer grads, or the recent engineers, have way more, flexibility and adaptability to work with AI, so that’s, I think, a great, thing for students and people who are passing out, or who are, like, newly in this field, that they could leverage all these systems To understand and work with how to work with it to give feedback, while relying on the subject matter experts, and the area tech leads of data, or, you know, engineers, or whatever, to ensure that whatever they produce goes through their review. Attendee?

Attendee: Hey guys, could you touch again on how you have Claude, learn from any mistakes you find when you’re double-checking its work?

Sravya Madipalli: Yeah! you talk to it, is, I would say, the simplest thing. But, I… I mean, there are multiple things that you, understand when you ask the same question, Attendee, to it. You basically tell it, this is a clear discrepancy. I would not expect you to make this mistake. What do you need to be corrected in your system to make this happen, right? So, there is a, plot itself tells you, oh, I basically don’t know this gap, and this is why I made this mistake. It actually tells us where it goes wrong, as well. Along with you yourself identify, right? Because we are the area experts whenever we ask the question.

So, I would thoroughly encourage people, when you work with these systems, to work on the areas that you entirely know the nitty-gritty details about. Because only when you know all the nitty-gritty details, you could find these flaws that Claude itself probably can’t find, or ignores, or jumps between loops, and tries to… assumed things, right? So. I would say half of it that Claude itself, once you point out that, hey, this is wrong. Why did you get this wrong? It gives, like, a, you know, reasoning for why it made that decision, and you understand. why it made the… in that reasoning, and a part of it is what you would identify as part of when you look through, what did you do to get here?

And when you look through that workload, you can understand that, oh, this is probably the reason why it got wrong, and you talk to it and ask it as well, yeah.

Shane Butler: There’s also, I’ll send you this article, Attendee. There’s this thing that people are doing from Carpathy called, Auto Research, which is really cool to, like, automate some of this, where it’ll basically… It only works if you have a bunch of ground truth to test on, but like I was saying, if you have, like, a bunch of historic experiments.

basically, like, it’ll train on the experiments, and then it’ll try… it’s an attempt at the experiment, and then it’ll look at, like, the historic thing of, like, oh, what’s all the stuff I did wrong, and actually do its own error analysis that’s, like, by different dimensions of what it got wrong, and then it’ll feed that back into the agent and say, like, rebuild your system. Like, whether it’s your knowledge base, skills, or agents. And then now try it again on another test experiment, and you have to create some metric, like, you could have these some, like, similarity metric, or some sort of metric of accuracy, and it just, like, loops like that in a, like.

forever process until it optimizes that metric. I’ve been doing that for a few different systems I’ve been building out, and… That’s, like, a really cool way to get it to bake in all of its learning back into the system. I’ll send that to you, it’s, like, a pretty new thing.

Attendee: Awesome, yeah, I think I’m… I was even thinking about just a maybe simpler tactical technique, like how to… how to have it correct its own mistakes, so it’s the equivalent of, like, taking your dog, your puppy, and rubbing its nose and something. You know, correct it to reward it and replenish it when it… so it doesn’t do the same mistake twice.

Hai Guan: Yeah, a quick answer to that is, you ask it to think about why it made the mistake, and don’t do it again, and then it’s generally pretty good, and then you do it incrementally.

Attendee: Cool.

Sravya Madipalli: Okay, we are almost at time, guys. Unfortunately, we have a hard…

Hai Guan: Yeah, I think we’re out of time. So, for people who are looking to ask more follow-up questions, I think there’s a bunch of engagement here, thank you all so much. Definitely hop into our Slack channel. We’ll have, you know, like, if you post those questions, we’ll be able to answer them over there. Otherwise, we’ll send out all the links and stuff.

Sravya Madipalli: Awesome. It was amazing, guys. Thank you so much for joining, joining our… our lightning lesson. Any questions, we are always… we are… all three are active on LinkedIn, so, you know, message us, comment on our posts, what else you need, and check out… check out more on AIanalystlab.ai. Okay?

Attendee: Very good.

Sravya Madipalli: Bye!

Attendee: Thank you.

Free, every week

The next one is this Wednesday.

10 AM Pacific, live on Maven. One topic a week. Bring a question from your own work.

WED SEP 30
Ace Analytics Interviews with AI
Register
WED OCT 7
Metrics 101: Define a North Star with AI
Register
WED OCT 14
Build a Semantic Layer So AI Defines Your Metrics
Register
WED OCT 21
Experimentation 101: Run an A/B Test with AI
Register
WED OCT 28
Trust Your AI Analytics: Know When the Number Is Right
Register

Next cohorts start Oct 19 and Nov 2.

AI Analytics for Everyone
$1,800 · Oct 19 · ★ 4.9/5
Enroll on Maven
Agentic Analytics: Build an AI Analyst
$2,500 · Nov 2 · ★ 4.9/5
Enroll on Maven
Or come to a free workshop this Wednesday. Register free