← All free workshops
Free workshop · Friday, May 15, 2026

Analyze Metric Trade-offs in Claude Code

Live on Maven, Wednesdays at 10 AM Pacific. About 60 minutes.

Transcript

Auto-transcribed from the live session and lightly cleaned. Attendee names are removed; their questions are kept.

Shane Butler: Admitting… Hey, everyone!

Hai Guan: Hello.

Shane Butler: Welcome, welcome.

Hai Guan: How’s it going, everybody?

Shane Butler: what, let’s see, usually we ask where people are from. As we warm up the chat. But, does anyone have any… Weekend plans, given it’s.

Attendee: On… on to your Cloud Code Analytics… Intro to Cloud Code Analytics 5. That’s… that’s, what, 3 hour in the morning? That’s… That’s a big plan for tomorrow.

Shane Butler: Nice, lovely weekend plan. Joining us for another 3 hours. That’s awesome, Attendee. We’ll see you there. I’ll be there as per my weekend as well. And.

Attendee: Isn’t that early? Where are you based out of?

Shane Butler: Yeah, 7… it’s 7 a.m. to 10 a.m. I’m based out of South Lake Tahoe, California.

Attendee: Oh, yeah.

Shane Butler: And roughly. Yeah. Yeah. Where are you? Where we’re at, Attendee.

Attendee: I am in Vancouver.

Shane Butler: Oh, nice, nice, nice.

Hai Guan: Yeah, so it’s…

Attendee: It’s pretty.

Hai Guan: early. Early crew.

Attendee: Early crew, yeah, and for you, for me, for both of us.

Shane Butler: Anyone else got fun weekend plans? Feel free to drop in the chat. Trying to think what I’m gonna do.

Hai Guan: Anyone else joining us tomorrow? Does anyone know what we’re talking about? Nope. Okay.

Shane Butler: We got Attendee joining us tomorrow.

Attendee: Yeah, you mean for tomorrow? Yeah, I’m joining you guys somewhere.

Shane Butler: Oh, nice, nice. Attendee says this is their rot weekend. Yeah, before the… oh, yeah, yeah. And then this summer, just, I’m sure you just got back-to-back stuff to do. Oh, I think I’m going to, like, a craft… A craft thing, actually, at a friend’s house. My wife signed me up for… to do some crafts with some friends. I don’t know what they are exactly. Excuse me.

Hai Guan: That’ll be fun. Nice, very cool. We’ll give maybe another minute, and then we’ll get started. Is that cool?

Shane Butler: Yep. Attendee is doing, Cloud Code Analytics with us tomorrow. Yep. 7 to 10 a.m, 3 hours, nice! We’ll see you there tomorrow, should be fun.

Attendee: True.

Shane Butler: Saravia, do you have any plans this weekend?

Sravya Madipalli: Yeah, I’m traveling for a health retreat in North Carolina, so… This is my first time doing something like this, so pretty excited about it.

Hai Guan: That’s so cool.

Shane Butler: That’s cool.

Sravya Madipalli: Nice. Are we… are we doing where people are coming from right now?

Shane Butler: We’re trying to get people’s weekend plans, but no one wants… no… everyone’s doing very secretive things this weekend, I guess. No one wants to share what they’re doing with us.

Hai Guan: Maybe back to the original prompt.

Attendee: It’s a long weekend here in Canada, so there’s some plan with taking my child to a water park.

Sravya Madipalli: Nice.

Shane Butler: Oh, nice, water park.

Attendee: Yeah, he, he looks… It’s summer, everybody is out in Canada in summers. Nice. Everything is super crowded.

Hai Guan: Very cool.

Shane Butler: Attendee’s going to the river. That sounds nice.

Attendee: First of the season, so hopefully, my very pale legs will be okay on Monday, we’ll see.

Shane Butler: Do you… do you, like, float… float a river?

Attendee: Yeah, we live 15 minutes from the river, and so we’re out probably every week.

Sravya Madipalli: Weekend, weekend.

Attendee: And it’s supposed to be 84 here on Saturday, so I’m gonna take that opportunity to get out and start building some color on the very, very winterized legs.

Shane Butler: Nice. That sounds lovely.

Hai Guan: Wow. Walt sounds like amazing plants for a lot of people. Certainly, I think the… the best plan is to… is the… workshop tomorrow.

Shane Butler: He likes workshop. It’s dead on Saturday.

Hai Guan: Yeah, cool. Should we get started? All right, well, welcome again, everybody. Thank you for joining us. Today’s lesson is going to be about analyzing metric trade-offs using Cloud Code. And hopefully… I’m not sure if you guys, have seen some of our content. We’ve run something like 15 Lightning lessons by now, so we’ve got… we’ve done quite a bit of this in the past, and I think many of you are familiar with us. If not, just a quick introduction. My name is Hai. I run the data team at a legal tech company called ENTRA, and previously spent my entire career at consumer tech companies like Nextdoor, LinkedIn, Pinterest, Meta, and so on and so forth. So, I’ve been doing this since, I guess, earlier this year, and the three of us have a lot of fun doing it. I’m gonna pass it over to Sean.

Shane Butler: Cool. Hey everyone, I’m Sean. Yeah, one of the leaders of the AI Analyst Lab, and I’m also a principal data scientist at the same company as Hai. a legal tech company called ENTRE. Yeah, I’ve been in data science for about 10 years, and I’d say, like, the past 2 years have been focused mostly on, like, AI evals and agility analytics, so… Yeah. Really happy to have you all here join us today, and… talk a little bit about the capability of Cloud Co. as it pertains to analyzing metric trade-offs. I’ll pass it over to Shravia.

Sravya Madipalli: Hello, everyone! So excited to see you all here. I am Shavya. I’m in the data science field for the past, probably, 14, 15 years now. Been with Microsoft, then eBay, then Nextdoor. That’s where Sean High and I met, and most recently at Superhuman. So, yeah, I’ve been part of this journey with Hai and Sean. For quite some time. Very excited to share all our learnings with you.

Hai Guan: Cool, yeah. And, the lab, we call it, that we’re all working together on is called AI Analyst Lab, and if you’re ever interested in some past lessons, past workshops, or upcoming courses, or workshops, take a look at their… over there. We’ve got just really quick We also run different workshops and bootcamps, and we have one tomorrow on Introduction to Cloud Code Analytics, and then we have another one bootcamp just a weekend after, which is 2 days. We’ll talk a little bit more about what those are, later in the lesson. Okay, so what are we here to talk about today? We’re here to talk about guardrails. So, like, understanding how things move against each other, and then how you, as either a data professional or a business owner or a decision maker, come up with plans to sort of mitigate, the downside that may not be visible to you. So, so in this slide, assume that, we have you know, like a product team, optimizing for, something like a mobile checkout, funnel. So when you, you know, like, when you go buy something, let’s, let’s, let’s think, like, Amazon, there’s the, you know, if you shop on your phone, you probably have, go through a bunch of payment steps, go through a bunch of add to carts, and, and finally input your credit cards and stuff like that. At the end to, to, to actually check out. So, you know, like, I think it’s probably not super surprising that, Things could go, you know, like… one metric could go up, and another metric could go down, and as a rigorous sort of analytics practitioner, or data practitioner, our job is to make sure that we cover all the blind spots as early as we can, and plan ahead of time as much as we can. So, this example here, very simple, it’s like, hey, assume if Or let’s say, like, if something, caused or some sort of feature that, that the team, launched on a checkout, flow. led to conversion rate that goes up 15%, perhaps, but then two weeks later, actual net revenue was down. this is probably not that big of a, let’s call it, like, right now it’s still hypothetical, but probably more common than many people sort of, like, have the, have the discipline to, to, to, to plan ahead of it. And so this entire lesson, what we’re gonna do is to really just think through how to Use a very simple framework to, to make sure these traps aren’t being, aren’t being stepped on. Who, actually would love to get a poll here, like, who, who, who’s… Encounter something similar at your workplace, in your role, or have seen this play out from time to time, or ever? Is this a common thing?

Shane Butler: I saw this a lot, working with, emails and notifications. Lots of, trade-offs. In terms of optimizing those.

Hai Guan: Nice, yeah.

Shane Butler: Attendee’s got their hand up.

Attendee: No, I’m saying that I saw.

Shane Butler: Oh, you saw, you signed it. Oh, cool, excellent.

Attendee: this happened in several of my clients. This is a concern, especially in the e-commerce industry. small shops based out of Texas will see this. not only 15% is a very small number for them. Sometimes they say it’s double or triple in a very… in a very short span of time, and then they fall down. Pretty, pretty significantly.

Hai Guan: Yeah, well, yeah, I mean, trade-offs, decisions, happens all the time, and so, this is… what we’re gonna talk about is really just, you know, a tool to help, make it, I guess, more streamlined. I mean, tough decisions is still gonna need to be made at the end of the day, but, you know, as much of a structure we can put around it, the less taxing it is every time that it feels like an ad hoc thing over and over again. Cool, okay. So, before we get into the actual framework and a demo, with Cloud Code, just wanted to sort of, like, let this group know, if you follow us on LinkedIn or in previous lessons, we’ve been pretty, like, we’ve been pretty open about, we’ve developed something called an AI Analyst, which is. almost like a senior product data scientist clone in, Claude Code. So, we’ve got an agentic system where, it… it bakes in all the best practices for how to be a really effective product data scientist. And so, you know, this is free to clone. Everybody can go and, and, and… and take a look, or use it for your own work, or your own interest. It’s all cool, open source, completely free, tons of skills, tons of agents, that we encoded the, you know, our collective 50 years of knowledge in this field. And, you know, from… in this lesson, what we’re gonna be talking about is gonna… we’re gonna grab, like, a very specific piece. from the AI analysts, System to just walk through, the guardrail component of how you should think about success metrics, but then also, like, what we call shadow, in, in this lesson here. Cool. And yeah, just quick plug for, for folks who are interested in, you know, maybe you’ve seen the AI Analyst Repo, maybe you’ve seen us taught a few lessons in the past, and you’ve heard it from this, from the audience from this work… from this lesson as well, earlier today. We’re running a couple things, tomorrow we have a workshop. This is hands-on. We’re gonna help you walk through how to install Cloud Code, how to clone the repo, run the real analysis, and all three of us will be there for 3 hours to make sure everybody is totally set up. That’s sort of like the shape for the workshop tomorrow. And it’s 25 bucks. We just want to make sure… we wanted to make it free, but we also want to make sure that people actually do show up. So a small… a nominal amount to hopefully get you onto the journey of, really up your productivity in the whole data slash analytics space. And then the weekend right after is the bootcamp, which is a two-day thing, 4 hours each day, where we’re gonna walk you through how to build your own AI analyst system. So, all the thinkings behind, you know, like, what goes into spinning up a skill, spinning up an agent, how do you wire them together, all that kind of stuff. That’s gonna be the, that’s gonna be in a week and a little more. We’ll come back to the specifics on these at the end. So right now, let’s go back to the content here. Cool, okay. So, the main idea of this lesson really is, is this. So, every success metric, has a shadow. So, the shadow, it’s also called guardrail, is the metric that gets worse when you optimize the success metric, you know, too much. So, it could be, like. I think the, you know, we’ll get into the examples later in a later slide, but a lot of teams don’t actually think about these, shadow-slash-guardrail metrics up front. So, when the success metric moves, they would immediately say, oh, yeah, that’s a win, like, we’ll just, you know, like, that looks great, we’re done. Or, or they celebrate too early. But if you don’t check these shadow metrics or guardrail metrics, then you really don’t have a complete picture of the mechanics of how things are working. So, you know. In addition to making sure you make the best product decision or business decision with a holistic picture. you also need to… it also helps you teach, kind of like, hey, when this thing moves, this thing moves like this. And so that muscle memory, is actually, I would argue, more important than the on-the-spot sort of one metric, or one instance, or one feature that you are, or your company is shipping. So, hopefully that makes sense. Cool. Okay, so here are some examples of, shadow pairs, so things like, what left… on the left side is success, on the right side is the guardrail. So, first one here is, is, you know, like. very commonly, if you’re in the e-commerce business. Conversion rate, for example, so purchase… purchase conversion rate, or, like, checkout conversion rate, something like that. And then the guardrail for that could be, what is the average order value? So if you, for example, if you make things so easy to check out, like, to buy, does it… How does it affect how much you’re buying? And these always, you know, like, these would move, these would move in interesting patterns, typically. The second example here is, signups. So, converting into a, for example, like, signing up for an account on some website, for example, the… the, the guardrail for that would be, are they retaining? Like, are these signups just, you know, like, very low-intent people who are just, you know, having a very easy time to get an account, or are they actually serious about, sticking around, which is the day 30 retention in the guardrail component? So, you know, like, that’s kind of like, you know. It’s hard to… When you start naming these. pairs, it’s gonna be hard to think… how having visibility into this, would, you know, make it easy to game the system. For example, you know, in the case of signups, like, the easiest, most more A very classic example is, like, hey, I can always make the sign-up button bigger, right? And people would convert, and so on and so forth. But, like, you know, probably, most likely, they’re not going to stick around. The people who sign up because of it. And then… whoops. And then we have, engagement, versus NPS, so net promoter scores or, complaints. So, you know, like, you can have very… what do you call it, like, attention-grabbing sort of, stuff that gets people to be very hooked to your products, but then, you know, like, your, your complaints might go up, or your NPS might go down, so it could be, you know, like, you could send more emails, but then people might, might just get pissed off, because of it. And then the final example here is, support speed, so, like, if you run a, custom, or I guess a call center of some, or, like, what do you call those, like. Customer experience, teams where, you know, the job is to help people, resolve their concerns or whatever. You can imagine, hey, if they’re gold on just support speed, like, how fast do you close the tickets, then it could be that the guardrail for that is how much of those tickets, when you are so aggressive in closing, would actually become open again. Like, you know, the actual issue may not have been resolved, but then because the metric itself is to measure against the time to which the ticket closes, the guardrail would be, okay, it would suck if the same issue is open again. Does that make sense? Hopefully, this is pretty straightforward. Cool.

Attendee: Yeah, this is pretty straightforward. Thanks.

Hai Guan: Great, okay. So, let’s see. So, the actual framework here that, that, that is, is, helpful, and we walked through two of the, sort of, like, three things here. So, really, there’s only three decisions that you really want to think about in terms of, pairing guardrails to your success metrics. And, we obviously we talked about two of them. The one… the first one is success metric, like, you know, like, no surprise. What are you actually trying to improve? And, what are you trying to actually measure to proxy that improvement? So, just some very concrete examples here. A bad success metric would just be, you know, something called engagement, for example. Like, what does that mean? It could be defined a billion different ways. A good one would be very specific. So, for example, in the whole e-commerce, space, you know, could be… purchase rates, for new users, or something like that. So, very specific, very intentional about the thing that you are actually trying to move, with a feature or with, whatever, product improvements that you may have. And then the second piece is the guardrail, right? Like, the guardrail itself. Like, what is the shadow here? It is the thing that catches the trade-off, so I think a lot of teams that I’ve seen in the past might be like, oh, let’s just track everything, or let’s just look at whatever we have. That… that we don’t deem success metric, and see what that looks like. The really… a really good practice is actually to predefine it up front. So, like, be intentional about Hey, if we think this is the thing that we’re trying to improve. what is something that we’re not willing to compromise? And so… and also be very specific on that. So here it’s, you know, if you track everything, that’s a… that’s a bad statement, generally. A good one is if you list out, sort of, like, the possible things that this could go wrong that you think would be, would be detrimental for the business, that you know would be detrimental for the business, and then just make sure those are explicitly named. Making it up front is actually… One of the most important aspects of this, because when you already launch the feature, for example, and then you then start to hunt for data or metrics to either support or… or make the decision on the fly, that’s actually really, that’s actually not very healthy, because then you’re… you’re basically just trying to, trying to use data to kind of, like, You know, like, like, almost like data theater, to, to prove a point, if you will. Okay, and then the third one here is actually, what do you call that? Like, actually put a threshold on the guardrails, that… that you’ve picked. So… it’s one thing to be like, hey, here’s my guardrail, or a set of guardrail metrics that we care about, that they shouldn’t, you know, move, or whatever. In reality, if you drive up, for example, a success metric. most likely, the guardrail would also move in a different direction. That’s why it’s called a guardrail anyway, because, you know, like, there’s a good chance that there’s some… there’s adverse effects on that metric, but you’re hoping that the gain of the success metric outpaces that by a lot, for example, or that, you know, in the best case scenario, it doesn’t actually move. Or, you know, even improve. So, pre-aligning and pre-defining on what that threshold or the willingness for you to be comfortable with defining, let’s call it, like, a feature to be successful, if it moves success metrics by this much, and it doesn’t degrade the guardrail by more than this much, then you have a really powerful combination of a clear, criteria that you don’t even have to think about when you actually launch the feature or, or make the product improvement. And then, on the slide here, there’s a couple kind of, like, ways to think about it. One is, like, alert threshold. So, for example, if it only moves by a little bit, and then, like, you know, in that area, based on your prior context or your business domain knowledge about how your products or how your features move, you know, maybe 3-5% is not a big deal. So, anything below that, you’re good to go. And then there’s another one, which is like, hey, an absolute kill threshold, which is, if it goes worse than this threshold, then we’re absolutely killing the thing. Or, you know, like, we’ll absolutely pause it and then go figure out what’s going on. So… Always having these, kind of, predefined is going to be helpful so that you’re not scrambling after the fact to be like, oh, like, what does this mean? Like, is this good? Is this bad? And then, like, make decisions on the spot. Does that make sense for people? Like, in terms of, you know, why this is important? Like, doing it up front versus doing it afterwards?

Shane Butler: I think something else, and you might get, you might get to this later, hi, is, like, it’s not just about, like, you as an individual. making these decisions around the alert and kills. It’s also about just, like, being able to align teams, so… You want to take as much of the decision-making process as possible, and just do it before any data is even created or looked at, so everyone’s aligned, and basically, like, their opinions aren’t biased or polluted after the fact based on what happens. So, like. I mean, if you, like, have, like, a finance team that carries about one metric, and a sales team that carries around one metric, and a product team that cares about another metric, if you’re all aligned in the beginning, then it makes it really easy. after everything’s said and done, to move forward with whatever your decision is. If you don’t have this alignment around thresholds before, then those conversations, can end up, like, with a lot of conflict, they can get, like, even, like, political. And all this just unnecessary stuff that’s basically, like. You can avoid by just setting some numbers ahead of time in a document that you can point back to later on. Because everyone’s out… out there trying to do what’s right for the business, but everyone has also, like, their own priorities, right?

Hai Guan: Yeah, the more complicated, the more tricky, the more sensitive a feature or an experiment is, the more important it is to pre-align stuff, so that you don’t get into Sean’s described scenario, where it may even get political. Because when people see that, especially people who are like, hey, it affects my area positively, then I’ll go… I’ll probably go crazy to try to, you know, like, not care about the other areas that get affected by this. All right, cool. So, we’ll get into the cloud code component pretty quickly, but let’s do a very quick exercise here. For each of these success metrics, maybe drop it in chat, what guardrail would you put in? So, let’s… App Store rating. If this is the success metric that people want to optimize for, what is, what is a shadow? Shadow for cart says cart abandoned… yeah, that’s a good one. Bot response, app store reading daily active users… yes, Attendee, you got it. That’s awesome. Yeah, some sort of engagement or feature usage. So, the idea is, like, oh, you know, app store rating, may or may not be reflective of the actual, you know, like, almost like a qualitative versus quantitative aspect, and both can go… can go in different directions. Maybe we’ll do one more, and then we’ll get into cloud code. Push notification, what is the… Bing. What is the shadow here? What is the guardrail? NPS, push disables, unsubscribes, yes. Great, yeah. Unsubscribes or, or, anything that indicates, just being annoyed. So NPS is, is fine, unsubscribe rate is fine, push disable is all good. These are all, in turn, is good. Probably more laggy, but yeah, it’s, Definitely, definitely a good idea. Cool, yeah, so seems like everyone is, now pretty comfortable looking at what this… what, the whole guardrail thing looks like. Okay, so we’re gonna get into Cloud Code in a bit. Just wanted to set up the demo scenario here. So, in the AI Analyst Agentic system that we open sourced, and again, feel free to download it, take a look, poke around, and build on top of it, the… Kind of like the, the, the… the fictional dataset that we included in the, in the repo is, a fictional company called, Nova Mart, which is an e-commerce company, think of it as, like Amazon, like, yeah, think exactly like Amazon. It’s just a, a fake company. And, Let’s say… if a product team is looking to ship a feature called Save for Later, so some sort of button that’s like, oh, you know, instead of buying now, like, I’ll just save for, save for a future time, and then you can come back to it pretty, pretty, pretty easily. Hopefully this is straightforward for folks because of, the whole Amazon mechanics, so think of it as the same thing. And, you know, our goal is to understand you know, for this feature, and given this, primary metric of 30-day purchase rates, for example, what would be possible guardrails that, that we can, that we can have? And then we’re gonna get into Cloud Code right now to use it. Pointing at some of the, the skills that we have in the… in the system to help us answer this question, and then we’ll see how it works into, kind of, like, future analysis. Let me find the share button here… Okay. Alright, so I am switched to the other screen. Can you guys see my Claude Cope thing? A black screen here?

Shane Butler: Yeah, it has, like, the overall conversion. Yes. Top in there.

Hai Guan: Awesome. That chart has nothing to do with what we’re talking about, just a default thing that, that I was looking at, previously. So… Okay, so for folks who are, not familiar, this is VS Code. It is an IDE environment. Think of it as, like, a workspace that you set up, such that you can have a very intuitive, easy to use navigation panel to look at your folders and files. So, I’m in AI Analyst Plus, but think of it as, like, the AI Analyst repo that we shared, and within it, it’s got different agents, different skills, different slash commands, and stuff like that, so… If you look at it, it’s like, it’s got a forecast skill, it’s got a, metric skill, and what we’re gonna be looking at is guardrails. And if you’ve never… done, clock code or anything, it’s actually very, straightforward. No coding background required, and the workshop that we run tomorrow exactly shows how to, you know, get this whole thing set up step by step. What I’m doing here is, clock code and terminal, so right here at the bottom is basically my terminal. And, I just signed into… I just went into Cloud Code, and, and started working there. So… what I’m going to do, because there’s so many different, components here in the… in the system, I’m just gonna ask it to tell us what, what… what sort of a guardrail or metric, pipeline does it have? So, I’m gonna…

Shane Butler: Maybe while you’re typing that high, just to get a feel for the room. I’m curious who here has used Cloud Code for analytics before. Maybe drop a zero in the chat if you’ve never used Cloud Code before, a 1 if you’ve used Cloud Code but not for analytics, and a 2 if you’ve used it. for analytics. Kind of helps us, know the audience. Cool. So, a lot of people have used Cloud Code, got a couple people who haven’t used it, a couple people who have used it for analytics, but most people have used it, but not for analytics so far.

Hai Guan: Got it. Cool. Nice. Wow, that’s a really good poll. Alright, so I just… I’m firing this prompt. Can you show me an ASCII diagram of the metrics and guardrail pipeline? So it’s gonna go look and explain to me visually what it is that we have in the Agentic system as pertains to metrics and guardrail metrics. So I’m just gonna fire that off… Let’s see… Cool. Alright, it’s… should be relatively quick, I am hoping. Yeah, so what it’s doing is, burning tokens, but also looking at, doing its own thing to… Show this at the end. Okay, so… Go here. So, this is the illustration of the workflow in the, in the AI analyst system. So, metric and guardrail pipeline. It’s visually trying to show us what this looks like, what the sequence of steps is, and what it does when it encounters metric and, and, you know, like, how it how it kind of, works with success versus guardrail metrics. So… step one in this whole pipeline, where, you know, if we use it to do analysis and stuff, it will have to help it define certain metrics. So it’s got different, components to it, and it gets very specific around, hey, what is a reusable piece of metric when we say. you know, whatever, like, monthly active user, how is it actually defined in the most granular way possible? So it’s got different skills around metric specs, and then it forces the system then to pair a guardrail against it, at least one guardrail. So, at least one guardrail per success metrics, and then measures a different dimension, it gives some… examples around, if it’s quality, then look for something that’s quantity, speed versus accuracy, that kind of stuff, and then figure out the acceptable thresholds and what things should move by. It could figure out, based on the data that you have. but the general recipe is that it will figure out… it’s… it will take… it will make sure to bake that in as part of the, the job, if you will. And then there’s the, kind of like the… what do you call that? Like, the examples of, common pairings that we went through. It will start to… it will… it will… when it actually does the analysis, for example, in an experiment setting, then it will, you know, obviously compute the two, and then it will check against the guardrails, and then it will have verdicts and the threshold that we just talked about. Right now, it’s just an example, but you can imagine, like, it could be specified based on the… based on the business and the domain, and And your use case. And then it’ll give a report, that kind of stuff. So, what I am going to do is, We’re gonna… we’re gonna do the exact setup that we just looked at, from an earlier slide. I am planning… Actually, let me just paste something here. I already have it written down. Alright, cool. So, this is the prompt, the next prompt that I’m gonna give it. I’m planning to ship a save for later feature on Nova Mart, so I have the dataset already loaded. It comes pre-installed, if you will, with the AI Analyst repo. I’m telling it the primary metric is a 30-day purchase rate. What we’re looking for is a plus 15%, increase in that, for the save for later feature, and, just help me design the guardrails for it. And then we can actually use it to, to brainstorm slash confirm what would make sense to do. And so… yep, and so it’s gonna keep thinking a little bit. It will go into the… the, the guardrail skill. It will go into some of the metric skills to, to figure out, you know, hey, in this scenario, where is my question fitting in the, in the sequence here? Cool. Claude Opus 4.7 is not the fastest model, but… oh, okay, that’s, fast enough here. Alright, look at the fat of the flattery language. Good feature, the guardrail.

Sravya Madipalli: That’s one thing it’ll never forget, to give some flattery.

Hai Guan: Yes, exactly. Okay, cool. So, good feature to guardrail, big risk is cannibalization of immediate purchase. So, notice we do… we did the, like, it understands that the save for later feature is the… the thing that, you know, like, that you can, that you can come back to, hopefully, in the future when you’re not as ready to make the purchase right now. And then it’s flagging that the big risk is that maybe people would be over-reliant on the fact that, hey, I can always just punt it. down to a later time, so people would actually not do immediate purchase. So that’s the big idea here, and then it’s gotten… it’s giving some recommended guardrails based on that. The… So the guardrail could be same-session purchase rates, average order value, it would have the… the, what do you call that, like, the acceptable range as well, based off of the data that, that it has access to on the Novo Mart, company, and it also has Its own judgment around why these things should, you know, should exist, or what the rationale is for picking some of these, picking some of these things. So it’s pretty detailed, and you know, like, right now it’s just the interactive, kind of like me talking to Claude, on the screen, but in the actual skill, for example, it will actually be able to write out a detailed report once we make some decisions around, you know, these are the guardrails that we care about, that kind of stuff. So… Even says, hey, three things to nail down before the launch. What are these, attributions, measurement, design, blah blah blah. So, when I need you to… Let me write the full spec. So it has access to a metric spec skill, so that it can log when we make these decisions that, for example, hey, if I think return rate is really important here, then it will log that in the system, such that in the future, when I actually use it to analyze an actual experiment or an analysis, it understands that that’s what I care about. So, let’s do something here. Let’s say, we don’t actually track Return rate… What else would you suggest? So, I’m trying to show the interactive nature of using it as a brainstorm buddy, so almost like a, you know, like the thing that… the phrase that we like to to say a lot in the AI analyst lab is that you’re getting an AI companion, effectively. You can delegate your execution to an AI assistant that thinks, like, the best, you know, as much as possible, that we bake in the best product data scientists out there. Cool. Thought partner, yes, Attendee. Great point. Thought partner, not a replacement. Okay, so two reasonable proxies, it says, hey, we don’t track return rates, then it would be like, hey, do the ticket rates, pre-shipment cancellation rates, and then it will tell me why this would actually proxy the return that we don’t have, or that we don’t track. And then what are the caveats, and what, what are some… acceptable ranges that it then, kind of pivots to, with a new metric like this. So, also has recommendations and, and stuff like that. So, So, you can go pretty deep with it, in terms of thought partnership, brainstorming, and stuff like that. So, what I’m gonna ask next, and then we’ll open a… go back to the slides and stuff, hopefully this gives you an idea. What if return rate increases, but 90-day purchase rate? remains flat. So, I’m trying to see… so, right here, at the top, the… some of the recommended guardrails says, hey, return rates and 90-day repeat purchase rates, I want to see, or not, yeah, 90-day purchase rates. I want to see, hey, if these things move in different… in a given scenario, what does that mean? So that, you know, you can… you can, you can figure out pre-aligning up top why… what you can expect and what you would do if you see certain things. So this is really helpful to just, you know, like, get your, get your rationale in the right place. remains flat, what does that tell us? And so, it will continue to think and give us a scenario, and then you can imagine this could be done for a billion different scenarios. So, the… the… The really helpful thing to do is just to be very comfortable with some of the choices and scenarios for things if they happen this way or that way, if it’s unclear. 36 seconds… yep, Opus 4.7 is not the fastest model, so… If folks have a way to make it run faster, let us know.

Shane Butler: Use 4.6.

Hai Guan: Yes.

Shane Butler: We talked about this in the chat, but I also like to toggle the effort on 4.7 to something that’s not as high.

Hai Guan: Yeah. Okay, so translation feature is changing what customers buy, not whether they keep coming back. So… Yeah, so, most likely story, it’s generating incremental purchases, but meaningful shares are speculative, so it gives some… possible explanations for, if this were to happen, then that means that. So, assists you to not have to think, you know, like, I mean, you know, it would be a lot easier to just chat like this and get a decision than to, than to, I guess, figure out all the scenarios and then all the different, different combinations of what would have happened. Okay, cool. I am going to switch back to my slides here.

Shane Butler: You had a question in chat here, hi, from… or Attendee, you wouldn’t ask your question?

Hai Guan: Let me see… Yeah, which question, sorry?

Shane Butler: From Attendee here, are the acceptable ranges derived from your historical data, or are Claude using best judgment?

Hai Guan: Yeah, so, great question. So, I’ve done some analysis in the past, with this exact dataset, quite a bit, actually, because we recorded some lessons in the past, and things like that. So… As you use the repo more, it will know more of the shape of the datasets and understanding of it, and what you care about and what you don’t. You can also name it explicitly to it, like, you know, hey, I care about the threshold being less or more than 5%, for example, and then it will… it will codify it. So, pretty flexible. For this one, I think I just figured it out based on my previous runs.

Attendee: Got it. Thank you. Appreciate it.

Hai Guan: Alright, no problem, great question. So, okay. So, let me go back to here. Am I showing the right screen? That’ll be safe.

Shane Butler: slides.

Hai Guan: Cool, okay, same, black screen here, so I can’t really tell. Okay, we did the demo, Cool. I wanted to leave the last at least 10-15 minutes for questions, so, couple ways to take the next step, if you’re ever interested, just like what we talked about earlier in the lesson. Tomorrow, again, we have the intro to Cloudco Analytics. We’re gonna help you set up the whole thing, install. clone… do actual real analysis, and we’ll be there to help you guide every step of the way. In a week and a half-ish, during the weekend, we have the Analytics Bootcamp, the Claude Code Analytics Bootcamp. For attending this Lightning lesson, you’ll get 20% off with the code CLAUD20, and this expires tomorrow, end of day, Pacific time. So, something to… We would love to have you join us if you’re interested. This is where you’re gonna build your own AI analyst system. Such that, you know, like, you have… you can use it on your own data, you can, you can have it tailored to, your own use case, your own habits and, and your own context. So, you know, learn to… you can come and learn to build with us. We’ve run this once, and, pretty good, really, really good reviews from, from our… from our students.

Shane Butler: maybe one more plug for the workshop tomorrow. So, like, the reason we’re running this workshop tomorrow is, based on feedback from our initial bootcamp, and just kind of what we’ve seen kind of bringing, Agentic Analytics and just Cloud Code more broadly to all the companies that we work at, is that… the hardest part of… and I think we saw in the poll earlier that a lot of people were… were, say, one, so they’ve used Cloud Code before. But a lot of the… the hardest part of this, I think, is just getting started, and I think that, like, you know, it can be daunting for a lot of people to get started with cloud code, or working in terminal, especially if, like, you’re not a super technical person, but I mean, like, once you get started, I… and it’s, like, a couple hours of friction, we find that you’re kind of just, like, really off to the races and dreaming up stuff that, like, no one else would think of. Like, for instance, like, my… wife’s in marketing, she’s also, like, a ceramics artist, and she’s not technical at all, doesn’t know any sort of coding, but, I got her set up with Cloud Code to help her, like, build out her ceramics business, and then she’s just asking it, creating and developing things that I would have never, thought of before, so… we want to do this live, even though there’s, like, a lot of, I’m sure, tutorials around there, around getting set up. We think that, like, just spending a few hours with people live is a great way for us to unblock people, and have other people, like, unblock each other as well, and kind of support each other. So that’s why we’re doing tomorrow, you know, it’s only, like, 25 bucks. So yeah, give it a… give it a check out, or feel free to DM us if you have more questions about that. It should be pretty fun.

Hai Guan: Nope. Looking forward to seeing quite a few people from, from this crowd tomorrow. Okay, cool. So this is the last slide. We’ve got roughly 10 minutes left, so open it up for Q&A if there’s been any questions, maybe in the chat, or if anyone wants to come off mute to ask anything top of mind.

Shane Butler: Or feel free to drop, questions in the chat.

Sravya Madipalli: I think Attendee had a question to the follow-up on the thought partner. Attendee, do you want to go ahead and ask a question?

Attendee: Yeah, yeah, sure, sure, thanks, thanks for bringing it up. So, I had a question about, how… what do you recommend? I think, Sean, you mentioned about, I’m just, just looking at my question, sorry, you had a, you know. that you can have, no, so my question was, when you want to, have a thought partner, do you go about changing your original Cloud Code, MD file, to… to give a brutal, honest opinion, about your metrics, or do you go about using another agent that you have… you can create, I think you can create autonomous agents through Florida right now, and use that to stress test, you know, what you have what results you get. So I’m trying to see what would be the best approach.

Shane Butler: Yeah, I definitely do the, sub-agent route. A couple reasons. Like, one, you have all this kind of stuff baked in to your system already, that’s kind of biasing it, but you also have, like, if you’re working with, like, your primary kind of, like, orchestrator agent in Cloud Code, they’re gonna be biased by all their context, they’re gonna be… like, do… they’re creating the work, so, like, when they judge their own work, they’re gonna be biased to say it’s correct. Even though I do find it, it does… it will catch itself. So I’ll either… if I… if I… if I’m kind of done with where I’m at with it, I’ll sometimes just clear context. And then if I clear context, I’ll have it recheck. But the sub-agent is usually what I do. And I even have, like, a skill that’s called, like. it’s called, like, it’s like a, it’s called Architect or something, but it’s like a plan mode where I’ll have, multiple personas created specific to whatever I’m analyzing or developing, and each of those will… Basically provide… will read whatever’s the analysis is, and then they’ll separately critique it, and then they’ll come together as a group and argue with each other, and then they’ll go back and revise their critiques, and they’ll come back together into line. on, like, a, like, a line on their feedback, but I do find, like, forcing them into their own context that isn’t polluted by everything they built before is a good way to get more, kind of, reliable Answers… We’re also gonna talk about… we’re not… we don’t go over it… In this coming bootcamp, we have an advanced bootcamp in June, where we’re going to talk about leveraging both, like, codecs and Claude code together, so having totally other, you know, frontier models critique the work. I find is pretty useful as well.

Attendee: Yeah.

Sravya Madipalli: What happens with me too is when you ask it to review the… whatever plot code generated, it tends to, like, be biased and say, yes, oh, this is the right way, but if you kick off a sub-agent with minimum context, it does a way better job at validating the plan, whatever it is that you’re working with. So I do the same as Sean, what Sean shared as well. And, we plan to share all of this as part of our advanced bootcamp. We just realized there’s so many questions around this, and we don’t have time to do that as part of our bootcamp that’s coming on May 23rd, so probably two weeks after that, we’ll have an advanced bootcamp, where we’ll share the best practices, like, how do you go around, like, multi-project, like, you know, multi-project flows, how… what… goes into cloud.md, and, you know, all of that stuff. Yep.

Hai Guan: Let’s see… Looks like Attendee, you have a question, if you’re still on? Or, you have two questions. Do you want to ask them life?

Attendee: Sure, can you hear me right?

Hai Guan: Yep.

Attendee: Yeah, so first one is when feeding data to an LLM, what format is best? Markdown, CSV, JSON, whatever? Does that matter on the purpose of the analysis, or what should I think about there?

Hai Guan: Yeah, I’ve found it doesn’t really matter. It knows… it knows how to… how to figure it out, basically. And even if, for example, like, within… your… let’s call it, like, CSV, for example. If there’s, like, bad formattings or, like, very irregular stuff, it figures it out pretty quickly, pretty easily.

Attendee: Right. And what about the numbers? So, how detailed should they be, or should they be rounded? Does that matter if, let’s say, I have a dataset with thousands of, entries?

Shane Butler: Yeah, they… they can just be, like, raw numbers. So if you think about this, it’s gonna be, Everything that’s ran here, the numbers themselves aren’t fed into the LLM. Basically, the LLM reasons through Python functions to run that will calculate over the numbers, so… from that point of a Python script, it doesn’t matter if you are gonna give it a, like. Rounded number, or… it’s more… most raw form, so… I’d probably go with, like, whatever its most raw form is, because it’s… The most accurate, and then you can… work with Claude to decide, like, from the… at the output stage, like, what’s the right kind of, Rounding you want for your audience. I don’t know, hire, Shravi, if you have any… Thoughts on that?

Hai Guan: Yeah, I… I think… I think that… that sounds… that sounds right. And then there’s…

Attendee: But.

Hai Guan: questions… sorry, go ahead, Attendee.

Attendee: Just a follow-up there. So, obviously, if it’s going to be handled by a Python script, then, yeah, for sure, raw would be the best, but what if we’re actually going to use the LLM itself to do pattern recognition? Would that then be a benefit with rounded numbers?

Shane Butler: That’s a good question. I mean… I think you could. I think I would actually… I think I would want the LLM. if it was gonna do pattern recognition, I would think I would actually want it to create Python helper functions to do the pattern recognition itself. Basically, my kind of, like, philosophy with anything with numbers and analytics with LLMs is I want it to do as much of that in code as possible, because whenever you start putting the numbers directly, like, through the LLM, you start opening yourself up to, the probability of hallucinations, since it’s now, like, in that mode of, like, predicting the next BEX token. Whereas if you have it, like. create some… Functions to, to do the pattern recognition itself, then you know it’s just reading output. But, that being said, yeah, I think, like, probably rounding, if you’re gonna do some sort of pattern recognition, I imagine would be a lot easier for it to, like, not get tripped up on if you had something with, like. Point and, like, 8 decimals or something like that. I don’t know, hire Shwabi, if you have any follow-ups on that. It’s a pretty good question.

Hai Guan: Yup. No, I agree. And, I think on the point about the put things in code, or reason over code, it also expands beyond analytics as well. So, you know, like. I think quite a few folks have used cloud code, but not for analytics, and so if you’ve ever built workflows or automations or whatever it is, something that becomes repeatable when you start putting it in code is actually going to guarantee reliability much more than if you have a reason over its… in its raw form all the time, so I think that carries over to how we think about the analytics side as well.

Attendee: Thanks a lot, that makes a lot of sense. Thanks.

Shane Butler: Nice, thanks for the… thanks for the question. Any other questions?

Hai Guan: Got a minute left.

Shane Butler: Or will the recording be available? We’ll… we will send out the recording in an email, probably later today or tomorrow. Anything else? Yeah, let’s see. If no other question, yeah, tomorrow… Join us for the workshop. If you’d like. Should be an interesting one getting set up. And then, the bootcamp’s really fun. The May 23-24 bootcamp’s really fun, and we’re gonna run it monthly, if you can’t make it in May. We’ve got another one in June, I think it’s, like, June 13th and 14th or something, I can’t remember exactly, but we’ll run that monthly. It’s got, like, a… I don’t know, I think it’s got, like, a 4.95 out of 5 star on Maven right now. It’s honestly just a good way to, like, meet other people in the space as well, too. Yeah. If you have any questions about that, feel free to… to ping us.

Sravya Madipalli: I also want to let everyone know we have a Slack workspace where you could join and, you know, meet other people with the same… in the same area, so join the Slack works where I just gave the link. We have around 600 people, and we’ll keep sharing all the workshops there, and, you know, everything about the courses and other things, so… and you could.

Shane Butler: Oh, yeah.

Sravya Madipalli: Yeah.

Shane Butler: And we have two workshops next week. We have two other free workshops next week.

Hai Guan: Oh, Create Lightning Lessons, yes.

Shane Butler: Yeah, I’ll probably send out an email to everyone. That has those as a reminder, but on Wednesday, we’re gonna go through root cause analysis in Cloud Code. And on Friday, we’re going into Turning Insights into Action and Cloud Code. So, those are… those are free kind of walkthroughs, just like this, where we’ll talk about concepts and do a demo, and… Yeah, and do some Q&A.

Hai Guan: Cool, awesome. Well, thank you all for… joining us, Hope to see you tomorrow. Otherwise, have a good weekend.

Shane Butler: See you, Attendee.

Hai Guan: Do you all.

Free, every week

The next one is this Wednesday.

10 AM Pacific, live on Maven. One topic a week. Bring a question from your own work.

WED SEP 30
Ace Analytics Interviews with AI
Register
WED OCT 7
Metrics 101: Define a North Star with AI
Register
WED OCT 14
Build a Semantic Layer So AI Defines Your Metrics
Register
WED OCT 21
Experimentation 101: Run an A/B Test with AI
Register
WED OCT 28
Trust Your AI Analytics: Know When the Number Is Right
Register

Next cohorts start Oct 19 and Nov 2.

AI Analytics for Everyone
$1,800 · Oct 19 · ★ 4.9/5
Enroll on Maven
Agentic Analytics: Build an AI Analyst
$2,500 · Nov 2 · ★ 4.9/5
Enroll on Maven
Or come to a free workshop this Wednesday. Register free