Shane Butler: Alright, welcome, everyone. Welcome back.
Hai Guan: Hello? Hello, hello.
Shane Butler: I’m gonna turn off the waiting room, hi, if you don’t mind, just letting people in in the meantime while I get that off. Okay, waiting room should be off now.
Hai Guan: Okay.
Shane Butler: And let me share my screen for everyone. I think I’m gonna share… the full window… So it’s easier to hop between the Sync Cloud code.
Hai Guan: Where’s, where’s everyone coming from? Where are you all based? Drop it in chat.
Shane Butler: Yeah, if you can see my screen, drop in the chat what you’re… located from city or country. It’s always… I’m always curious to see how far These reach, because sometimes it’s like, hey, everyone is, everyone’s in… the US, or it’s like a… Europe spread. I feel like this one’s too late in the day to get the Australia-New Zealand crowd.
Hai Guan: Yeah. Dallas, Austria. Wow, welcome, Attendee. Brazil. Philippines, San Francisco, Sweden. Wow. It’s always good to see how international these sessions are.
Shane Butler: Nice. I’m located out of Tahoe, highs in the Bay Area. Shravia? I think Shravia’s calling in. Yeah, there she is. Let me get you as a… Co-host, Ravia.
Hai Guan: Fremont, Roseville…
Shane Butler: Okay, well, we can… we can kind of… I have, like, some… while people kind of filter in, I have a… we can do some intros and do a few little polls and stuff of the room. So… for the next about half hour… okay, hold on, confirm. Alright, so today we’re talking about, developing Northstar metrics. We’re gonna talk about, like, about next half hour, or maybe 40 minutes. I’ll try to leave 20 minutes at the end for questions, though. It seems we can never do that. It seems like we always just go to the end, and they’re like, we’ll stay over. So we’ll see if we can do it today.
But yeah, we’re gonna talk through how you actually pick the one metric that kind of tells whether your product is really working, and then I’m gonna put that in front of, we’re gonna use Clock Code. and show you a couple of things, live in the demo. We’re partnering with Amplitude on this one, so we’re gonna be using Cloud Code for, like, the system, but actually, everything we’re teaching here comes from… well, it comes from, like, our experience, comes from a bunch of other research that we’re pulling in, but actually, like, the whole kind of, like.
Amplitude, North Star Playbook, a lot of the, frameworks that The agents and skills and functions that we’re gonna run code today, are based on that. It’s actually, like, baked in. And so I’ll show you the skills and agents today are a little different than what I usually run. Usually, we kind of have a skill that has a bunch of information around it, what it should do, and it might fire off other skills or fire off agents that then fire off helper functions. In this case, a lot of the skills and agents go more in depth to, read through, research, and that’s curated in Knowledgebase.
Read through different rules that apply to different industries when it comes to metrics, there’s actually an agent that’s called, like, the North Star Librarian, so… who will go and make sure that any sort of decision that’s made around your metric has some sort of source, that will pull up to you as well. I don’t know… Maybe, hi, if you want to look up the North Star Playbook… from Amplitude. It might be one of those things where you have to, like, put your email in, and then they’ll send it to you, but it’s pretty good. So I wouldn’t mind… I wouldn’t, I would… I would recommend reading through that as well, although we’re going to talk about it here.
To kick things off, though, first, let’s do… oh yeah, let’s do some intros. I think I recognize a lot of the folks here today. You’ve probably come to some other ones, but I’m Shane, I’m a co-founder of the AI Analyst Lab, We do a bunch of free, workshops around, Agentic Analytics. We put out a bunch of, free, kind of like, email courses, async stuff around learning how to use Agentic Analytics. We’ve done a lot of work primarily with Cloud Code, although this past month. We’re getting a lot more into codecs, and open source models as well. We also offer some, paid courses. We actually have a boot camp coming up this weekend where you build your own Agentic system.
It’s like a no prereqs required, no code required, teach you how to build, your own agentic system. And then, a lot of that’s built… a lot of that’s primed on, like, our open source Agentic analytics system, AI analysts that we’ve… Released, I’m sure many of you have access to that repo. And then next week, we have a pretty cool course called AI Analytics for Builders, which is, like, a 5-week course around analytical thinking, but the execution layers in… in Cloud Code. Maybe we’ll get into codecs on that one a little bit, too. So think of it as, like, your kind of, like.
you know, I’m, product data scientist or data analyst, and I want to have analyses that go faster, that go deeper, that go further. This is, like, that… or I’m a PM, or engineer, designer, or banker. We had, like, a chef in here the other day, actually, and I want to, like, learn how to, to do analysis, but I need to, like, understand the frameworks first, and then now you don’t have to, like, do the coding in R, SQL, or Python, you can do it… do it first called code. So we’ll talk more about that later, though. That’s us, hi, you wanna do a quick intro about yourself, Savia?
Hai Guan: Yeah, hey everyone, my name is Hai, co-founder of AI Analyst Lab. As Shane said, we put out a lot of content, both on LinkedIn, our emails, and Maven, just like this one. I’ve got 20 years or so experience in the data science and analytics space, and currently utilizing AI for analytics and agentic systems for analytics quite a bit. Chavi, you want to go next?
Sravya Madipalli: Yeah, yeah. Hello, everyone, I’m Stravia. Sorry I had to be off video. But yeah, I have 15 years’ experience in data science. I started off career at Microsoft, and later, Nextdoor, and most recently, Grammarly, which is, branded as Superhuman. At Nextdoor is where Shane Hai and I met, and that’s when we started a bunch of things together, and most recently, AI Analyst Lab. And, very excited about the journey that we’re taking from here. Yeah. Shop.
Shane Butler: Sweet. I also just saw… Attendee. That you said you’re located in Miami and Bali. It sounds tough. It sounds like a tough… life. Yeah, sounds rough. I lived in Indonesia for a couple years, actually, so I, I spent… I’m in East Java, but I would go over to Bali every few months. Alright, cool, let me get into it. So, a couple questions before we kick off, just to get a read of the room. We can kind of change the… the format of this, knowing where people are at. So, first thing here, I just want to see, like, people’s, kind of, where they’re at with Cloud Code, where they’re at with Agentic Analytics. So, if you’ve used Cloud Code. If you’ve never used Cloud Code before, put a zero in the chat.
If you’ve used it, but you haven’t used it for analytics, put a 1. I’ve used it… I mean, it’s for analysts, but it too. Cool, we got…
Hai Guan: Wednesday?
Shane Butler: So the people who haven’t used it before, lots of ones. nice. Attendee, welcome back. Sweet.
Hai Guan: tooth.
Shane Butler: Well, yeah, a lot of our… Oh, nice. Hey, I think you get a… I think you get to be a 1 or 2 if you use codex. We’re just… we’re saying Claude Code. I feel like they had, like, with the coding stuff, a little more of, like, the first mover thing, but Yeah, it’d be like codexes. Doing the same thing these days. Yeah, and then… I don’t know. This is, like, a total aside, but I’ve been talking to a lot of people as we’ve, like, been working with some of the other models, too. I just feel like the open source model stuff’s gonna catch up so fast to the… to being able to do what these other things are doing.
It’s like the… I think the stuff I’ll do today will do, like, Opus 4.7, but to be honest, it’s, like, 4.6, 4.7, 4.8 Fable. In terms of analytics tasks, if you have a gigantic harness around it, there’s not much of a difference. Definitely some step up if there’s, like, vanilla It’s asking Claude stuff, but… I feel like we’re gonna hit some… some threshold where the models are so good. But, you know, that incremental lift You don’t necessarily need it to… to accomplish what you’re doing with your day job, but cool, good spread today. We’ll definitely be showing the capability of Level 2 today.
And then second one… Alright, more on North Star metrics, because we’re going to do, kind of like, some framework discussion here before we even get into analytics. Maybe a 0 if you haven’t really worked with a North Star metric before, and a 1 if you have. Let’s see… Yeah, go… go connect to us, Attendee. I like the community building. Cool, nice, this is good. Good, even split, 50-50.
Hai Guan: Yeah.
Shane Butler: Nice. So some of this is going to be new, some of this will be review for people in the beginning. You’re one… if you’re a one, some of this will be a bit of a review around what North Star Metric is, but we want to set the foundation for… because there is, you know, I’d say at least half the people in here haven’t worked with, Just north of where she’s 0 and 1, Attendee, alright, cool. I mean, I know a lot of people who have heard the word Northstar metrics that I’ve worked with who, I would say they’re basically at the zero level when I see them try to make the metric, but no longer, because AI can help them, so they don’t have an excuse. Yeah, so… so let’s get into it.
So where I want to start is, like, think about, your team. Your company’s, kind of, like, primary dashboards for a second. And, you know, this… maybe this is a little dramatic, but if it’s anything like most, like, it’s not one metric. It’s really easy to add a ton of metrics on there, especially when you’ve got multiple tabs of dashboards going. sometimes I’ve seen them, like, especially… especially some of, like, the high-level ones, where it’s just, like… or even at the very low team level, it’s just, like, 30 or 40 numbers on it. At least, most of them have at least, like, 10.
And the question I’d ask is, like, you know, for… when you look at a dashboard like that, that’s, like, just inundated with all these different metrics on it, like, 40 is dramatic, but say even 10, like, how easy… how easy is it for your team and for you To all agree, and point out one of them, and say in a sentence, like, this is the one that tells us that we’re winning. a lot of people can’t do that, that’s why they have so many metrics on there. We track a ton of activity, and we kind of call that measurement. And honestly, AI makes that worse, not better, because now you can measure basically anything.
So, it’s easier than ever to pour all your effort into the wrong number, or collection of wrong numbers. So, what we’ll do, I’m gonna give you the framework here that companies like Airbnb and Netflix use to pick the one metric that really matters, and then I’ll show you how we can use AI to really pressure test that metric live in the demo. And we’ll turn, like, a full year of kind of, like, raw data into a decision in about… I think that even takes, like, 5 minutes or something to run through the process. But the problem, basically, is that a lot of teams What a lot of teams measure is activity. instead of value. So… It’s really easy to make this mistake.
Things like logins, or signups, or your monthly revenue, those go up. and to the right, and it feels really good when that happens. But if you’re honest, none of them really tell you whether your customer or your users got more value out of the thing you built, which is… Like, the reason you built it. It’s, like, you know, revenue is the reward for us creating the user value, but we build the thing to create the user value. revenue and whatever is just, like, a kind of steps removed that says, like, oh, that maybe happened. you know, I pay for plenty of subscriptions that I have just, like, forgotten to unsubscribe from.
I’m not getting user value out of them, but some companies getting revenue out of it. So it shows up in two ways, either a vanity metric, so one number that looks impressive, but means nothing, like daily active users is a very, very, very common one, or the opposite, a dashboard with, like. you know, 30, 40, 50 numbers, and no single one everyone’s rallying around. So, every team picks their own, and you end up arguing in a lot of these meetings instead of being aligned. And the reason it matters so much is that this metric sits at top the very top of everything.
Like, if the one you picked is wrong, then every decision underneath it inherits that mistake, and you end up shipping things that moves the number without actually moving the customer. So, So I’ll come back to it, like, we’re figuring out the right metric… actually, figuring out the right metric for your own product is week one of what we teach in our AI Analytics for Builders course, starting on Monday. We’ll come back and talk to more about that, but, like, it’s… this isn’t an easy thing to do. We’re gonna talk about it, right, in, like. We got 44 or more minutes here, but we actually spent an entire week, on this topic, because it’s so important, to make that one number correct.
Would love to get a poll from the room, like. For those of you who said 1, or even those for you who said 0, who could think about this, like, what’s your team’s kind of current North Star. Do you have, like, a clear… North Star for your… for your team chats. Conversion rate, conversion. Case time. Mal, ROI. You’re not gonna like what I say about Mao, for the Mao people. And you know what, I’m not gonna say it, AI’s gonna say it, so it’s not even gonna blame me. loss ratio. Nice transaction rate. Cool. Alright, we’ll keep going. So what actually makes a metric a real North Star?
So there’s kind of, like, 7 questions you can go through, and I’m actually gonna send you guys… So, this repo and stuff I’m gonna work through is our, like, plus AI Analyst Press repo. We have an open source one called AI Analyst. everyone has access to that. This whole Northstar thing, which I’ll show you later, is in our Plus repo. That’s currently only open to our core students. I think we might just open source that, too, later on. We haven’t, like, decided yet. this Northstar stuff, all of us got pushed to main on that one even yet, actually. We’ll probably open source that, but in the meantime, Come on, I cannot click on this.
I do have this… I’ll send this to you guys at the end that kind of has, like, just, like, a little worksheet with, like, some checks around, so you can check your own North Star, so you don’t have to… remember all this. But we’ll also send a recording of this, too. So what actually makes the metric a real North Star? So, 7… questions, does it capture customer value, reflect your strategy. Is it leading rather than lagging? Can your team move it? Can a normal person understand it? Can you measure it? And is it free of vanity? And then 4 of those 7 are really, real deal breakers. Like, you get all 7 of those, and it’s like, oh yeah, this thing’s strong.
But 4 of them are pretty, are pretty big deal-breakers. We’ll send a deck out after, Attendee. This is just… I don’t have it, post it anywhere right now, but I’ll send the PDF out. So, customer value? leading, and can you move it, and not vanity, those are… those four are really important. If you miss any one of those four. basically your North Star metric, I would say, is out, no matter how it does on the other three. The easiest way to see is an example of this, right? So, An example everyone knows, like, Airbnb’s North Star is nights booked. And so if you walk it through that check… checklist, a booked night is real value. Someone’s actually staying somewhere.
It leads the money instead of trailing behind it. The team can genuinely move that metric, and it’s not something you can fake. So it checks all those boxes, and that’s why it holds up. I’ll tell you, there was a company I once worked for that counted vanity metrics that counted their weekly active users, not just as people who were on the platform. But of people who would just, like, open any email from them. So it’s a very gameable. Vanity metric, especially when certain, certain, like, mail things, like iOS, just automatically say your email’s open, so if you just send more emails, suddenly your weekly active users go up. Don’t do stuff like that.
definitely don’t create that metric and then go public and not be able to change it for years. Makes things very, very hard. Northstar metrics are hard enough when you’re, like, internal to your team, but some of these inevitably will… get out, trickled out to the board, now you’re reporting something to the board. If it’s like a… if it’s not hitting those, seven things in a checklist, you’re given bad metrics to the board. If you’re a public company, and then the board wants to have that, like. reported on in your, whatever, quarterly earnings call, now you’re, like, you’re just stuck with this thing forever. You have to basically go private in order to get rid of that metric.
So, it’s really important. And all these decisions happen off of it. So… Let’s get a gut check from folks here. I kind of already teased this, but daily active users… 1 or 0? Good North Star or not? Type 1 for yes, 0 for no. Attendee’s like, it’s freaking good, man. Shut up, Shane.
I guess it depends how you actually define active, which is a very hard thing in itself to define, because… oh man, yeah, I mean, I’ve been in… I’ve been in… months back and forth with people who try and do, like, a daily active user, and then people are like, oh, but active is, like, when you do this thing 3 times, and this other thing 2 times, and this thing maybe, if you don’t do this thing, you have to do these other 7 things on that day, then we’ll cancel as active. It’s like, oh god, you have just turned… 20 metrics into 1 metric somehow, and that doesn’t make any sense. Alright, we’ll keep going. We’re gonna come back to this. We’re gonna have Claude tell us the answer to this.
So… One metric at the top, Northstar metric alone isn’t enough on its own, because you can’t reach up and move it directly all the time. What you actually move are things underneath it. These are, like, input metrics. And there’s a simple way To make sure you’ve kind of covered exhaustively all the, input metrics that drive our North Star metric. So there’s four kinds, you can think of them as levers. So there’s four kinds of levers, there’s, like, how many customers. How much each one of those customers does? How reliably it goes through. and how often they come back. And so those link back to breadth, Depth? Efficiency and frequency. BDEF.
Those are your levers, and the North Star is the outcome they add up to. So as those levers go up or down, there’s probably some metric at your company, or multiple metrics that relates to each of these, you’ll be able to drive your North Star metric. We don’t just work on these… the BDF ones, because, those are, a lot of them are gameable, but, like, the North Star metric itself is not. Let’s see, yeah, one thing that trips a lot of people up, your input can’t just be, like, your North Star metric wearing a different hat. This is really easy to do, actually. So say you pick, like, weekly active buyers. From your site as your North Star. That’s a headcount. Right?
And then… for how many customers you write, like, active buyers, that’s basically the same number. So you didn’t break it down, you just kind of renamed it. The way else, like, to pick a metric, you can break into pieces, a count of something happening, like orders are nice booked. Then the levers underneath are real and separate. And in a second, you’ll watch a tool catch this mistake on its own. Okay, so let me bring in the… AI… And… I can kind of show you a little… About… show you around a little about the, The repo here, and then we’re gonna run some stuff. Yeah, so… true, Attendee. It all comes back to prioritization of what we’re gonna work on, honestly.
Alright, so… this is our AI Analyst Plus repo. It’s like our AI… open source AI Analyst repo, but it has It’s, like, the currently… active, developed one, so we’ve been… the AI Analyst repo we released, I think, in February. We’ve been constantly developing on it, but in this Plus repo. And we’ll probably push some of this stuff out to open source soon, but we do… anyone who’s in any of our courses, they get access to… to this private repo as well. What I’m going to show you today, I don’t even think, is pushed into main on that one, but there’s a skill here North Star, And it has, you know. We could take a look at it.
It has, yeah, I’ll just read this out to you for now, but at the high-level description of it, North Star metric. It’s a North Star metric lifecycle Coach. It helps PMs design, audit, defend, it doesn’t have to be PMs, it could be anyone, diagnose and evolve their team’s strategic anchor metric across, their North Star Metric Journey, it’s cited. So it remembers the product across sessions, it grounds every claim in the Amplitude playbook, which is why we’re partnering with them today. It’s curated to that casebook. And, then it just has some things of, like, when it’ll… what it’ll trigger here. But it has a few different, For seizures it runs. audit.
So this, in this case, a user has a… a candidate metric already. And they want to go through, like, evaluating to see if it’s good or bad, see if it’s weak or strong. We’re gonna run the audit today, it’s the primary thing we’ll run through. We’ll also run through drivers, I think, if we have time. A triage, should a user once, like, is this even worth the audit? So this is, like, just, like, a quicker run of it. Although, honestly, the audit takes not that long, you’ll see. explain, user wants a cited explanation of a North Star metric concept, so this is a little bit different where we went with this one.
Rather than having just, like, a single skill that has kind of all the information in it, we have a whole, like, Wikipedia of, of concepts, In beliefs of cases. That… to point back to strong North Star metrics of anti-patterns, right? Like, gap thinking. insisting on multiple North Star metrics. So it’s like, it’s not just how to, like, create your metric, it’s, like, all these things that happen in real life. debates… it has different information for different verticals, like consumer subscription versus dev tools versus B2B SaaS.
So this thing goes into pretty far depth, and you can basically have it explain any concept to you as a teacher, as well as when it makes any decision, it’ll cite where it’s pulling that from. draft, this is something I’m still kind of working on, but this is, like, when you want to design one from scratch, and you don’t have something to audit. You know, usually where you want to start with is working with your team and trying to come up with a few yourself, and then it’ll make recommendations off you for you, but having a place to start from is really important. And then… Yeah, inputs and drivers are pretty similar.
modes of this skill, where it kind of decomposes the BDF thing, and goes through your, history from whatever window of time you select, and tells you how much is each of those, metrics driving your North Star metric. Okay, so that’s the real high level. Yeah, if you join our course, you can go through this more. Eventually we’ll get this open source, but it might be a few weeks, because we’ve got some other stuff going on right now. But let’s go ahead and go into the demo. So, I’m going to… earlier, right, I asked you, like, is Dow a good one or not? So, let’s see. We’re gonna use a North Star skill. Come on. If you can’t… see this that well? You can zoom in.
on Zoom yourself, so I don’t have to zoom in here and zoom out all the time. But when we do Northstar, I’m gonna go into audit mode. And I’m gonna give it a metric. Dow. These first few ones, Should run pretty fast. Later on, it’s going to… I think this next one I’m going to run, so I take a little longer, so we can take questions while it’s running. So, in this case with Dow. So, it’s opening, it’s triggering the skill, right? It knows that it needs to run the… Audit mode, so it opens this verb. Under the skill, audit. In camp and preview, it’s a little easier to run.
So, when it does this, you know, it, it, Pre-filters, anything already cleared, so, If, if, if, if the, if there’s any anti-patterns, it’s not even gonna run everything downstream. It’s pointless to go and run this rubric on it if we already know, it’s like a… anti-pattern, metric. So you can see here, pattern refusal confirmed, anti-pattern page has a real fix recipe with a concrete reframing example. So in this case, Dow is a really common one that people think are North Star metrics, and so, and we can… we can have this explained to us a little bit more, but you can see it refused, Dow matched a canonical bad pattern. We can ask… we’ll ask it about this, actually. In a moment.
It counts logins, not value received. So it climbs with notifications. Required logins, or addictive patterns without improving customer outcomes. So it conflates value delivery sessions with vanity sessions. You can definitely define something where you have a… this is a… a DAO of a value delivery session, then you have to define what a value delivery session is. That’s fine, but DAO as itself, it’s gonna reject that as a vanity metric. So, it has a test here. If a metric only goes up, counts raw activity, and you can’t explain in one sentence how its movement reflects customer value, it’s probably a vanity metric. And then it gives you… some other options.
So, Try replacing daily active users with the value-delivering action a user takes when the product worked for them. This way, happy deliveries dropped people opening the app in favor of deliveries with no issue. which research shows correlated with retention and CLTB? And so… and then it sources everything, right? So we could go into… the wiki… cases… Happy deliveries. And, you can read all through this kind of, case study. around, why this is more… a better metric to use. So we try and really, for this one, because you’re gonna get people who ask you why, why this, why that?
We really try and, like, ground everything in, like, research that’s already done, studies that have already been done. So you’re not making… this is just, like, this is… just isn’t new stuff, so we don’t need to start from scratch. Now, if we want. We could even have it explain some of this stuff. Oh, it even recommends it here, right? So it says, next steps, you can either, we’re gonna do this, actually. We can first have it explain, like, oh, we want to dig into why activity versus value. Maybe we were, like, pushing back on this. So we’ll say North… Star… explain. And you could look, if you wanted to look at the explain verb over here.
So basically what this does, it kicks off an agent, the librarian, agent… And, that agent is going to go, Read through our glossary and articles. And resolve the, kind of, The concept that we’re asking about. If you want to look at agents. Let me close some of this stuff. There’s an agents folder down here, and we have a special group of agents that are all around, North Star. And you can read more about, like, what that agent does here, so… The dispatcher gives you a resolved article, Adapts the format to the expertise level, so we can… Have more in-depth or high-level explanations. Applies all the citations to different, cases, or to the amplitude playbook.
Yeah, it surfaces if there’s debate around this topic, and there’s a whole thing in our knowledge base around debate, because everything’s not cut and dry, everything’s not, you know, people have different opinions, so you don’t have to… You can make your own, opinions on your own by reading through the debates. but what it gave us here is it says, you know, Vanity Metrics Northstar. A lot of this is going to reiterate what it already just told us, but in a little bit more depth. Yeah, so we already read, basically, the decision rule, and… The high-level thing before… DAO, ad impressions, downloads, page views, registered users, story points delivered.
Yeah, these all have, like… oh yeah, StoryPoint’s, like, feature factory stuff, right? Time on page. each of these kind of share that same shape of a raw activity count, standing in for value. And I just guarantee… I’ve been on teams where these are the North Star metrics. Every time. Like, but they are actually… they’re so gameable. And then, so, the fix, they say, is like, like I said before, like, yeah, you could tweak it, you could try to be like, hey, a daily active user that actually gets this value, but actually, the fix is more like, just find what that action, that value action was, and that’s probably your North Star metric.
Okay, so… That’s a little bit around vanity metrics, a little bit about how the audit works and triggers when you get one, So it’s… it’s… it is… a lot of it’s brought in from the Amplitude playbook, Attendee, but there’s also a bunch of other resources around metrics, So that’s not the only source, but it is… a lot of it is derived from that, and it’s derived from the cases that it’s based on, because that playbook itself cites a lot of, material also. It goes into that source documentation as well. And it’s kind of something that’s ever-expanding as well. But it’s a pretty good playbook. Okay, so let’s reframe… to another… Things, so instead of… Dah, we’re gonna do… Where we at at time?
We’re pretty good, I think. We’re gonna do North Star… This is gonna take a little while to run, so we can go into some questions while this happens. Audit… And we’re gonna do weekly… Completed. orders. Okay, so now I’ll try and do a real one. This time… it should run the full, like, 7… a full, like, seven-question check. Those four that really matter, like we talked about… and then those deal… the deal breakers, customer value, leading indicator. You can actually move it. It’s not a vanity. If it flags anything. Like, if it’s soft, like, like, completed orders is a little generic on strategy. That’s the tool, you know, being honest with us, not rubber stamping whatever we type.
We actually, like, you know, we don’t want this thing to be sycophantic when we’re creating metrics, so we… this is kind of a make it hard on us. Okay, so… It did not pick it up as a vanity metric, so it is not refused. It, is gonna keep going with it. Nova Mart is this dataset we use. It’s a synthetic data set we built. It’s like an e-commerce site. It’s just flagging here that e-commerce isn’t one of the industries We have put into the, The industry list here, so it needs to kind of, like.
reason through that on its own, which is why we are leveraging agents, and it’s not just, like, a pure DAG, because sometimes stuff’s gonna miss, but it’s smart enough to figure out, what’s good for e-commerce. Okay, while this is running. Because this is going to be a little bit. Do we have any questions? Any questions coming through the chat? Yeah, Attendee, that’s the right… that’s the right one.
Hai Guan: I think Attendee had a question earlier. I don’t quite recall what… what the… what the, what the context was, Attendee, do you want to ask?
Attendee: Oh, no, no, that was just a comment to someone else in the, in the chat. You can, you can skip that one.
Shane Butler: Cool. Thanks. If anyone has other questions… Deal free, it doesn’t have to be Northstar metric or related to, could be pure Agentic Analytics related. Alright, so you can see it’s going through the questions here, right? So, customer value…
Hai Guan: You have your hand up.
Shane Butler: Yeah, yeah, what’s up?
Attendee: So with a lot of these metrics. how much help, or how do you suggest providing context, I guess, around the business that you’re at? Like, do you actually suggest doing that, like, throwing together some sort of a markdown or something, saying, this is my business, this is kind of where it’s at, to, like, also guide that conversation? Or… like, right now, I’m assuming it’s… I don’t see that context for, like, what your business is specifically.
Shane Butler: Yeah, I would definitely recommend that. So we ha- so it does have, like, some… some, let’s see… some kind of businesses in here, but it’s… it’s nowhere near exhaustive, it’s just, like, some examples, and they’re pretty high level, like, it has, like. Because, like, each type of business is gonna kind of play a different type of game, in a way. Like, e-commerce is kind of like a transaction game, where it’s like, we’re just trying to get people to have more transactions, buy more stuff. A… say, like, a social media site is more of, like, an attention or engagement game. We want people on there… they’re… they’re getting entertained, right?
They’re not… we’re not… they’re not there to become more productive. And so, for your business, the more you can add. and you can make, like, a full wiki on your business and clear out all this other stuff. That’s gonna help it, have the context of, like. What metric is most relevant to you? Like, I worked, previously in legal tech, and I don’t think there’s anything on here for… Legal tech. So we need to build that out if we were to use it there.
Attendee: Good, thank you.
Shane Butler: Yeah, no problem. Okay, it finished the checklist, so what it did is it went through those 7 questions. That I showed you in the slides earlier, and it actually said, this is a… it passes as a North Star metric, but it’s weak. So, it had 5 of 7 checklist criteria pass. There’s no fatal failures, so, like, the 4 fatal questions, customer value. Not vanity. Actionable… camera, wherever the other one is. Measurable. Like, it passed all of those. But it didn’t pass. All… all seven, so… Okay, and it actually says, like, this is one of the reasons it says weak, for vanity. We can complete orders as a classic vanity. It counts a paid transaction, not a login, but it’s issue blind.
It counts orders regardless of returns, defects, or late delivery. And then it gives you a fix. So this is pretty good, right? So the fix would be add a value received qualifier. Weekly completed orders with no issues, meaning they’re on time, not returned, and no escalation. that’s a pretty strong North Star metric over just, like, pure orders. So you can see it gets pretty deep. Like, if I was on a team, it’s like, first of all, like, would we have gone to orders over just, like. active users, I don’t know, but if we did get there, I don’t know that we would have necessarily, filtered out, like. returns and such. And then vision strategy, it also said this is a bit weak.
It says generic, could be any stories, North Star metric, non-fatal. optional polish. So I think, this probably is where it comes in to be, like. kind of like you just said, Attendee, like, providing more context, it could probably give us, like, a better fix on something that’s more specific to This specific e-commerce site. I don’t know, like, if there’s some… some sort of… some sort of order or some sort of product they really, like, specialize in or something. Okay, and then it saves, a more in-depth readout of everything here in a Markdown file. And so, you know, you can go through and read in more detail if you wanted to. around, like, Did it pass, or fail, or just, like.
partially pass each of the questions, and then why that occurred. And then for… the ones that failed, what is the way to fix it? So, it gave us, like, a very, like, one-liner fix here. We can see It goes into more detail around how to fix it. And then it also ties you back to always any sort of, Kind of, like, relevant stuff in the industry, or cases, so it gives you similar cases, gives you some recommended next steps. Okay, and then this frozen context block thing, once you run this thing for a certain metric.
I cleared this out before we came here, but it then saves it into a JSON file, so the next time that someone runs it, it can just be like, oh, like, you ran this last week, we don’t need to wait 5 minutes, we already know the answer to this, and it’ll just draw it up for us. Okay, 15 minutes. I think we’re good. I wanna… I wanna show you guys a little more… a little deeper demo. So, let’s say we have actually, Decided we’re gonna move forward. With this, and now… What we want to do is, we want to understand the drivers. of this North Star. So we’re gonna say North Star… This one might take a little longer, so we can do questions again. Remember, we will run through some slides and come back.
So Northstar drivers. What was it called? Weekly… Completed. Orders… And it might have some follow-up questions based on the metric… But it’s gonna kick off this driver’s workflow. Which… If you wanted to… Again, you could, Read about this skill… up here… Yeah. So the purpose is to decompose what drove the North Star. Window over that, you know, breath. Frequency, efficiency, and, depth. So, it’s giving me follow-up questions. It’s basically asking if I want, like, a full like, what date window I want to look at, right? Because the, when you’re… we’re now getting into something, like, of like, hey, what’s driving it in North Star?
It’s going to be, like, what’s driving it over a certain time period, because that can obviously change over time. We’re going to say we want the full past year. This data is fictional, but… There’s, like, a few years of data, I think 2023-2024, so I should pick, like, 2024, Jan to December. So it’s going to, basically, decompose this into those four… Okay. That was way too fast. what happened is I ran this earlier, so it just picked up my last run, actually. But what I would… and that’s probably fine, because we only have 12 minutes left anyways.
But, Basically, what it would do… when you do drivers, I’m just, like, I’m telling it, take weekly completed orders as a North Star, break down what drove it this year, It’s pulling the orders, breaking the metric into the levers we just talked about, and looking what actually moved over the year. And then it hands me back an actual report, not just, like, the numbers. in the terminal, So you can see it has this report here. what it found, or the orders grew 6 times over the year, which on the face of it looks like a fantastic year, right? But when you look at the breakdown. almost all of that growth, like 99% of it, is coming from one single lever, which is just more buyers showing up.
So, how often people come back to buy didn’t really move. The size of the average order, AOV, which is our depth guardrail here, actually, went down. About 9%, and the membership program that’s supposed to bring people back in is, like, sitting at 3% or 2%, so it’s, like, basically… Not moving, or… Or, adding anything to the, to the North Star metric itself. So that’s… that’s interesting, right? Like, the headline number is telling you it was a great year, but when you actually read the breakdown, what it’s actually saying is you’re growing by kind of, like, renting new customers. You’re not keeping any of them.
And the one lever you’d use to fix that, of trying to get people to return seems to be broken. So it’s the same 6X either way. The number can’t tell you which one of those is true by itself, which is why you have to look at the inputs, which kind of goes back to, what was Attendee was talking about, where it’s like, okay, if we’re going to prioritize work right now, we know breadth is good, we can get a one-time person to come buy stuff really easily, we’re actually seeing a ton of growth in there, but we’re not having, returning customers, and then whatever these customers are buying over time is actually decreasing in their value.
So, you know, you probably want to look at, frequency and, obviously, depth as, things you want to increase here. And efficiency, too. But breadth, you’re good. Okay, I think… We had some slides on that, but we could just skip through them. Yeah, I was just gonna talk through… what it said. Yeah, a step back. everything I showed you, grading the metric. breaking it down throughout the year by the different input metrics, writing the report. The AI did that, like, really quick, right? Like, it does it in minutes. What it can’t do is then decide… it can’t really decide what you measure in the first place. You know, your team, your company has to, like, come up with that context, to start with.
I’ve been working a little bit on having it just, like, come up with metrics for you from the start, based on, like. kind of, like, the industry practices haven’t kind of rolled that out yet, but it’s still going to be unique to whatever your team within your company’s doing, especially if you’re trying to do something new that people haven’t done before. I can’t look at that The North Star metric, just, like, honest. own and know, like, whether to celebrate or to worry has to drive it down to those input metrics. And then the judgment around, like, hey, looking at these input metrics, which one of these should we prioritize the team works on based on, like, how much lift that’s going to be?
That’s, again, like, that’s the human’s decision. So everything else kind of stays the same, but it’s pretty good at, like, helping you create a strong metric and understanding what’s driving it. So, if you want to learn more about the kind of, like, building these systems, and also how to figure out, like, that kind of human decision part I just talked about, so turning decisions… turning data into a decision, and not just a chart, there’s a couple ways to do that. We have a boot camp, we have it coming up. Tomorrow. Actually. And then another one, we’re trying it a little different. We have a boot camp tomorrow.
We have a boot camp for a week, weekdays in July, for mornings for a couple hours. That’s where you build an agility analytics system yourself. And then we have a follow-on that’s our advanced, AI analytics course in the end of June, and that’s, like, taking it more into Production, so it’s getting into, validation, context engineering, multiple models. And then we have our 5-week builders course. That kicks off on Monday. The next time it kicks off is… in August. That’s how to frame questions. Basically, it’s an entire week. On this kind of stuff we talked about today.
But then it gets into, You know, obviously, how to frame questions, pick the metric, and dig into the data, but then in the subsequent weeks, it goes into, like, root cause analysis, correlative analysis, trend analysis, segmentation, and week four, it gets into causal inference and experimentation. In the fifth week, it gets into presentation and sharing out with stakeholders. The overlap between these courses, Attendee, I’d say, is just The main overlap is basically just, like. setting up, installing Cloud Code.
Sravya Madipalli: Three of us is the overlap.
Shane Butler: Yeah, the three of us is the overlap. You guys are the overlap, because you’re going to take them all. No, I’m just kidding. Although we do have a two-for-one deal with the five-week and the first, like, the build-it course, not the scaling one. But the overlap, truly, is just that, in the beginning of the bootcamp, we just make sure everyone’s running, got cloud code installed, can clone the repo and run through the analyst the first time, and we do that same thing in, like, week two. of the 5-week course, but everything else is different. So the build it and scale it are all about building agentic systems. So it’s more around development, and then the AI Analytics for Builders.
This is more about learning the analytical frameworks. So it’s pretty similar to what we did today, actually. We spend, I would say, about 70% of the time, 60% of the time, teaching the kind of, like, best-in-class analytics thinking and framework end-to-end, and then we teach you the remaining, whatever, 30% of, like, okay, how do you execute all that? Leveraging AI or agentic systems, rather than writing SQL or Python or R directly. Yeah, so that’s… I don’t know if that… that answers your question. And then we do run… a… two-for-one deal. So, let me save… Builder, bundle, bootcamp Advanced, yeah, okay, and then, so… the, if you do, the… Builder’s class?
The 5-week cast, we just… we toss in… The Build It one for free? And so… yeah, it’s basically a two-for-one. deal. I think they pair really nicely. They’re complementary, they don’t really overlap. You can do one before the other. You could do the bootcamp this weekend, and then decide you want to take the 5-week course. 6 months from now, and I just take whatever you paid for the bootcamp this weekend and deduct it from that. You could take the 5-week course on Monday, and then decide you want to take the boot camp. 6 months from now, and I just add you to the bootcamp for free, or you can do… add them both at the same time. Yes.
And then we also do a discount with, The advanced course, if you take, the bootcamp, and then when I move on to advanced. There’s some codes down here, any of these if you want to take them solo. You get 20% off. But if you take the 5-week, you should just take the bootcamp as well, because it’s free. So today… oh, yeah, I’ll send the recording out. in… I don’t know. Not too long. The price for Boot Camp Plus Builder, so it’s $1,800 is the, Builder’s course. The 5-week course is $1,800. And then we put it in the bootcamp for free, so… it’s $1,800 for both, so you save yourself $900. The bootcamp on its own is $900.
And then you can do 20% off, too, so if you just want to take the bootcamp, you know, it’s actually, like, $720 or something. And if… and if you guys wanna… have more questions about that, you can… DM me in Slack, or on LinkedIn, or… email me. I’ll drop my email on the… Chat. Yeah, bootcamps this weekend, and then again in July. We’re doing some new stuff with the bootcamp, so we really take the feedback from people who’ve gone to the cohort seriously, and we iterate a lot each time. We’ll continue to iterate. We do it for a couple reasons. One is because we’re learning what people want. So we want to, like, constantly iterate on the feedback with that.
Also, this… the industry’s moving so fast, so we just are adding new stuff to the bootcamp. All the time, like, we’re gonna add, some stuff on open source, this round, for instance. Because seems like those models are really catching up, and also… I don’t know if you guys have used, like, the Fable model. It’s really freaking expensive. Like, it just, like, eats your tokens, so I feel like there’s also, like, we want to be on your side and make sure that At your company, you’re just making, like, the best decisions, and not just, like, on the newest, best model because it’s the newest, best model, but actually beyond the thing that works get your job done, which I would say is, like.
a few model releases ago right now, like, probably Opus 4.6. And then… It says, can this be run in VS Code with Cloud as plugin, or does it need Cloud Code IDE installation? Yeah, you run VS Code. I run in VS Code. But I… I do it in terminal on there, rather than the plugin. And then the terminal, the other reason to do that is because, like, you don’t necessarily want to be tied to anthropic models. You might want to use… Google Models, or, OpenAI, or open source models, too. We’re not going to get that fully into that in the bootcamp, we’ll get to it a little bit. The advanced bootcamp, we’ll spend a bunch of time on that.
Something else we’re doing for… we’re testing for this bootcamp, too, is we’re gonna have… Two days, we usually just do two days, Saturday and Sunday, four hours a day. We are gonna add a bunch of bonus content this time, based on the kind of questions we get people from people. We always get questions around open source. Context management, invalidation, and data warehouse connections. So, throughout the week, after Sunday, the bootcamp doesn’t really stop. Monday, Tuesday, Wednesday, Thursday, we’re going to be releasing async content on those four topics. Validation, context management. Open source model and data warehouse connections. We’ll have practice exercises, so those are optional.
Bonus, we wanted you to be able to, like, have it sink in as you use it at work, and then Friday, at the end of the week, we’re gonna have optional office hours. So you can do the boot camp on the weekend. apply it at work throughout the week, go through the async content, and then on Friday, if you want, we can answer a bunch of questions live. Of course, the Slack’s always open, too.
And then in July, we’re… experimenting even more, where, we are not gonna host it on the weekend, we’re gonna host it on weekdays, so it’s not a, like, intense four hours, but we’ll do the two hours each morning, and then you can, like, go apply stuff at work all day, or on your own personal projects, and then come back the next day, ask questions, and it’s a little more piecemeal that way. Maybe we’ll… and we’ll experiment with other stuff in the future. If there’s other ideas. Okay, I know Route 1? I can stay over, because I quit my job last week, so this is all I do. But if other people have jobs to get to. That’s cool, no worries. But I can stay over if people have questions.
I don’t think I had anything else here. Yeah, my last site’s questions. Oh yeah, I’ll send out the recording, of course, and I’ll send out the PDF I’ll send out the PDF today. with everything. The recording, either tomorrow, I’ll probably send it out. And then, I had that little, like, worksheet thing that just has the questions and stuff to ask. I’ll send that out today, too. And then sometime in the next few weeks, if… for those of you who aren’t joining our courses and don’t have access to AI Analyst Plus. I think in the next few weeks or a month or something, we’ll probably open source all that.
When that happens, I’ll kind of blast out our whole email list, which all of you are on, so you can get access to that, too. But if you want it sooner, yeah, you would get it basically, tomorrow if you took the course. questions. No questions is fine, too.
Sravya Madipalli: I think Attendee had a question about, so according to this framework, does the North Star metric change over time?
Shane Butler: Oh yeah, like, so, your North Star metric is definitely gonna change over time, and that’s why we have the… why we have this, like, audit thing. So, like, like, it just depends what your team’s working on and your team’s mission. I guess, like, if your team’s mission is consistently, like, the same thing, then, like, yeah, it might… Stay the same for a long time. I would say you want to audit your North Star metric quarterly. I’d say the first time you make your North Star metric, you want to audit it weekly for, like. the first month, or month and a half, to see if you actually picked a good one, if you’ve never created one before.
Because it’s really easy to be like, this is the best, everyone agrees, and then you start building stuff, and then… you see the actual data and how it’s moving, and it’s like, do some analysis, and it’s actually not a great metric, and you kind of flag all these other things. And then I would say probably revisit it. Quarterly, if you’re kind of, like, what your team’s working on and your mission stays, like, the same, it’s probably not going to change much. Could last throughout the entire year. At a very minimum, you definitely want to visit it annually. But I would say, like, any quarterly planning cycle, think about your North Star metric again.
And then you can, like, rally around it for that quarter. Depends on, like, the cycle of, like, the product and company you’re working on, too. Okay. Other questions? Cool. Alright, this is great. Hey, thanks for the 20 of you who stayed over, and… Oh, how to define Northstar metrics in the B2B space? Yeah, I think it’s just, like, it’s just gonna be, like, when you… when you have B2B, like, B2B is when it actually becomes really important to, create a North Star metric, because it’s so easy to get caught up in, like, I’m just developing for, like, the loudest customer in the room. But, like, it’s the same thing.
You want to do it at the user workflow kind of level, and hopefully your prod dev teams are kind of organized in such a way that each of them are focused around a user workflow. So, like, an example of, like, a good North Star metric for, one of my… one of the team I used to work on, where we were automating out part of someone’s workflow, trying to make it faster, and so it was, like, time from the start to the finish of a successfully completed task without any, like, errors in it. But it’s like, it’s, I’d say, like, yeah, think of, like, what is the user trying to do in the real world? And then, how can we reflect the success of what they’re trying to do in the real world?
Within our product, and then what are metrics that can basically reflect that action in our product? Yeah, balance.
Attendee: So I got another question on… this is more on the AI analyst, lab, and it’s… I’m trying to think about how to apply it, like, pick a pet project, and so my area I’m trying to pick is… employee engagement surveys, like around an HCM. go through the analysis process. So, is the AI analyst the free… the free repo you guys have, a good starting place to sort of… kind of build that out and learn more about, like, how you set up that repo and adapting it? Or… I mean, are there any gotchas or places? And I can post this in the Slack, too. Just thought I wanted… I wanted to ask that.
Shane Butler: I think it would be pretty good for employee engagement survey, like, because that’s not going to be… I think that actually would be a really good place to start, because that’s going to be, like. either one table or, like, a CSV or something, you’re not gonna have to worry about Adding a bunch of, like, semantic… models around, like, hey, you have to join to these different tables. Although you could if you wanted to, like. learn specifics about employees later, so I think that’d be a pretty good… start. I would basically… download that CSV, or can… Within… into the repo, and then start, kind of, like. asking, like, what question… or asking it to, like, profile the data.
There’s, like, a data profiler function, and then start asking it, like. You know, can you help me frame some questions around, and then give it some topics, and start from there and see where it goes?
Attendee: Okay, and… and my next step was, is thinking about it, like, in a wider platform of, like, an HCM system, where now you have, like, payroll and, like, job, promotions and stuff like that. That’s when you start looking at, like, okay. defining the table schemas for what those other tables are, and some of those other… I think there were schemas, and you mentioned something else too, right?
Shane Butler: Yeah, there’s gonna be, like, yeah, you’ll want, like, the schemas, you’ll want to have, like. Kind of like these, semantic models and kind of, like, knowledge bases around kind of what we were talking about earlier, the context of your company, like, what are the metrics that make sense? What are the definitions to all these things? Like, there’s just gonna be so much stuff that it can kind of take guesses at it, but it will need you to, Kind of, like, formalize in a document somewhere in the repo that it can, like, read every time, so it’s not guessing every time.
Attendee: And then just the last one, what do you recommend for, like, do you just recommend raw CSVs if you’re just trying to, like, synthesize data to, like, play with and, like, check… test how well it works and, like, kind of verify? Is that…
Shane Butler: I think it’s a good place to start, but you can connect to, like, a data warehouse or something pretty easily, and it’ll run SQL for you. Obviously, there’s just, like, more… a couple more hoops you have to… Jump through to, like, do that connection. But, as long as it’s not, I mean, if it’s sensitive. like, information, you probably don’t want to, like, download the CSV onto your computer or something, but, like, if it’s not something.
Attendee: It would be synthesized and fake. It would be synthesized and fake.
Shane Butler: Yeah, yeah, I would just play with it in CSVs to start. That’s what… yeah, I mean, in our… in our, like, bootcamp stuff, we make, like, a local database, that it queries and stuff, since that’s, like, a little more what a lot of people are going to use it for, but a lot of the time when I’m analyzing, data. I’ll just, like. download some stuff from the internet as a CSV. If there’s multiple tables, I’ll have, like, I’ll make it create a DocDB database. It’s really good at that. Okay. If it’s all in one place, yeah.
Attendee: Thank you.
Shane Butler: No problem. Yeah, I think, Attendee, there’s, There’s an interesting thing there around, like, yeah, customer… requirements. I mean, I’m sure you could create there’s, like, customer wants, and there’s customer needs, and one of them provides real value, right? And so it’s like, this is, like, just, like, a hard thing with B2B of, like. Am I building something that provides value to multiple customers, or is just this something that this one person really wants, and then they’re gonna leave? That’s just, like, a balancing act. for any… B2B company that they have to make a decision on. But I think, like.
You know, something like, finance compliance or something, where there’s, like, true compliance things, I think you could make a, like, did this… you could go through and ask the North Star metric thing to run this, but, like, I’m sure it has some stuff in there, because there’s a whole fintech industry thing it has, where it’s, like, I bet there’s, like, a thing, like, how fast And how often do you run through, like, this without hitting any, like. I don’t know, compliance issues in the workflow. I don’t know, because I don’t work in that space, but I can imagine there’s probably something like that. We’re gonna talk about that a bunch in week one of the Builder’s course, though.
And then the other thing with the Builder’s Course, just to, like, give you guys a little more detail, it’s a little different from the bootcamp. The bootcamp’s all live, except for this bonus material we’re making. The reason it’s all live is because the industry is moving so fast that when we’re teaching people how to build agentic systems, like. I don’t really want to spend a bunch of time recording myself teaching that for it to change in 2 months and have to record again. That’s why we do that all live.
The, the 5-week course, a lot of that is around kind of, like, best of practice industry standards for, like, analytical thinking and applying these stuff that hasn’t changed for a long time. It’s a little more, like. evergreen, and so what we do is all the lesson material is async. They’re anywhere from a 5 to a 20-minute videos, about 3 to 4 hours per week. One hour is, like, you should watch all this, and then there’s, like, 3 hours of kind of optional async content. Totally self-paced through the 5 weeks.
And then we have live office hours for 3 hours a week, where we do more Q&A and discussion about whether we talked about that week, or what we talked about in prior weeks, just depends on who comes. And we stagger those office hours based on the time zones of the students who are in the course. So we’ll kick it off in the morning Pacific time for the first week, but then we’ll do a poll, and we’ll pick office hours based on who’s in there. So, like. Yeah, like, it’s probably gonna be, like, evening, morning, so we can hit every time zones. But we’ll talk a lot, we could talk in extreme depth about, kind of. your specifics of, like, B2B SaaS metrics there. in those live office hours.
Or, you know, I’m gonna talk about another free thing here at some point. Alright, my wife has called me 3 times in the past 5 minutes, so I think I have to leave. I just keep rejecting her phone call on my… on my watch here, so I don’t know what’s going on, but I think I have to go. But thanks everyone for joining. And thanks for all the engaging questions. And, I think next week, next Friday, we have one on validating… validate… validating the output of the Gentec analytics platforms. That’ll be pretty good, it’d be like, Kind of, like, similar thing to, like, today, we’ll talk about framework, and then do a little bit of stuff in cloud code. So, check that out.
Hopefully, I’ll see you there. Thanks, everyone.