Shane Butler: Hello! We’re gonna get kicked off here in a couple minutes. But we’ll, give folks a couple minutes to enter. Hey, Attendee and Attendee, while we’re waiting, do you want to drop, kind of like, your roles in, in the chat? Love to… See what kind of audience we have today. Hey, everyone. Hey, everyone. We’re gonna get started at probably 11, or… oh, I’m on a different time zone. 103 PST. While folks are rolling in, Feel free to drop your roll. In the chat. Data scientist product. Welcome, Attendee. Saw Attendee’s in the house. Hey, Attendee. Long time. Attendee, product management. I’m Attendee. Hey, application developer. Welcome, Attendee. Hey, Attendee.
We’ll give it another couple minutes, folks, and… Feel free to drop your roll in the chat while we wait. AI product manager. Transitioning from data science. I feel like we’re probably gonna see a lot more of that. Give it another minute here. Applying scientist. Attendee’s dropping her LinkedIn there, that’s how you know she’s serious about the networking in the chat.
Attendee: Yep.
Shane Butler: Okay, let’s go ahead and get started. Can y’all see my screen alright? Sweet. Okay, so, hey everyone, welcome to, redesigning product metrics for AI evals. It’s more like supplementing product metrics for AI evals, but I feel like, ChatGPT gave me redesigning as, like, the better hook, so you just have to do some of this stuff for marketing. But I’m Shane, a little bit about me. So, Yeah, I’m a principal data scientist at ENTRE. We’re a legal tech AI company. I’ve been there for a bit over a year, solely focused on AI evaluations for our data science team. Been working in data science for a little over a decade.
Across both B2B SaaS and consumer products, so you can see some of the companies I’ve worked at up here. And previously, actually, before grad school, I was in education a bit, so I was a teacher for a couple years, which… is why I like doing lessons like these, and also doing them at my day job, and stuff like that. One other kind of fun plug, I’m a co-host of a podcast called Data Neighbor. We started about a year ago. We talk all about, data as it pertains to AI especially, and we have leaders from, you know, top, Topkanite. data leaders, top AI tech company leaders, some exciting stuff coming in 2026, I think that’s relevant for this.
group is we have a new series we’re going to be launching. We’ll talk… I’ll talk more about that, but it’s called Measuring the Machine, and it’s solely focused on, AI evaluation, so we’ll do episodes on that for, For a couple months, actually. Okay. What else? I also am offering… so today, what we’re going to talk about is gonna be a little more of a high-level kind of framework, kind of how I think about approaching Metrics as they pertain to AI products when I’m kind of, like, working with a prod dev team. We’re not gonna be doing, like, any coding, or really, like, be looking at any data or getting into, like, the nitty-gritty stuff today.
I do have a 6-week… kind of applied hands-on course that I’m kicking off in April, where we’ll get way more deep into this kind of thing, so we’ll talk about the foundations, which this is kind of like a mini preview to. We’ll talk a little bit around… instrumentation and observability. We’ll talk about, more importantly, you know, how do you interpret success and failure of AI outputs? And actually, more and more importantly, how do you connect them back to the actual user value and the business? As well as the whole, kind of, like, flow, once you’ve created this metric, how do you kind of operationalize it?
How do you set up different Forms of operation, like, if we’re operationalizing it where you have safe layers, within, like, you know, your development environment as well as in production, and kind of, like, in the background in production. So, who should join? This isn’t, like, AI Foundations course, like, I’m not gonna teach folks about, like, what a, you know, representative versus generative model is.
The assumption coming in here is you’re someone who’s working on an AI product or are trying to kind of, like, move over to working on AI products in the near future, so product managers who are shipping AI-powered features, you know, data scientists, analysts who are measuring it, and engineers and MLEs who are building and maintaining those systems, either currently or sometime close to the future. And then, last plug on here, and we’ll get to the good stuff, and out of my, ad I have squeezed into this presentation, is that we have a little discount code, here, AIMetrics30. For those of you joining the session. Okay. the ad’s done. Let’s get into, product metrics.
So… Why do we have product metrics is kind of, like, a good place to start. I was actually, doing a podcast interview this morning, and… we… he mentioned this quote I really liked, that a metric stops becoming useful once you… start trying to optimize it. So, product metrics are not about the metric, the metric’s not the goal. Product metrics exist to support product decisions. So, they’re effectively just tools that allow us to answer, what are we gonna do next?
And they… should be… directly tied to real decisions, like the cost of… A wrong metric should be the cost of a wrong decision, and the benefit of a really strong metric should be a benefit of making good decisions in your product over and over and over again. You know, when to ship a product, when to roll back, when to invest further, when to stop. Why do we make product decisions? So, product decisions, these all exist just to help our users and our customers accomplish the goal. they came to the product for. So I told this joke yesterday to Attendee, who’s in here today, because he works at ENTRE with me, but, like, like, humans don’t wake up and be like, oh.
I… I can’t wait to engage with SaaS platforms today. Like, I’m so… I’m gonna go spend 90 minutes on Amazon right now. That’s not their goal. Like, the ways in which a lot of us, like, measure them around, like, Dow and retention and these kind of metrics don’t actually reflect the behavior they’re trying to do. You know, someone’s like, I’m going on a trip. to Hawaii, I want to book my tickets so I can get to Hawaii, or the holidays are coming up, or my friend’s birthday’s coming up, I want to get them a gift that, they really like, so that’s why… that’s why they’re on Amazon.
That’s why they’re on, like, sites booking their Their… their flights, they’re not on there to, have I don’t know, long dwell times. So, this creates this sort of, like, gap between what a user is actually trying to do and what we see them doing in product, right? So, you can have a user… you can have two users who have very… similar observations within our system. Like, they both viewed a page, they’re on there for 5 minutes, they have the same amount of scroll events, but one of those users in the real world might come out with, like.
say we’re talking about, like, Amazon or some sort of e-commerce site again, that they’ve identified the right size and great fit of a product that’s gotten to them, and they’re gonna keep it and use it forever, and someone else could find, like, some basic info, or not even buy it, or they get it and they return it. And so we’re… a lot of what product metrics are is, like, just trying to bridge that gap between these system outputs, these product outputs, and, the actual users. job to be done, or value. And so… if you worked in product analytics, or product data science, or product development, you’re probably familiar with the concept of North Star metrics.
And the reason I go through all of this is because, you know, I want to lay the groundwork of, like. how strong metrics work today before we get into how AI kind of upsets it. But Northstar metrics are really that, that bridge between these kind of, like. Really frequent and measurable, but gameable system outputs, and the actual business via the user value that someone’s acquiring. And so if you think about it kind of in terms of a hierarchy, you have these leading indicators, these input metrics, these system outputs at the very bottom. These are things that, like. You can get down to the millisecond. They’re really easy to measure and quantify most of the time.
At the very top, we have, like, super lagging indicators, things like business outcomes, like, I don’t know, net revenue or churn, things that could take us months. So if we make a change right now in the product, we’d see those input metrics change immediately. But if we make a change in the product that’s bad. We’re not gonna see that user churn, potentially, for months later. And we can’t wait months to understand what’s going on in the business.
The North Star metric helps us connect these two together so that we can leverage input metrics that we know are grounded in, like, user value, this metric that we basically prioritize, like, what the team’s working on this quarter, or something like that, to, like, these bigger business outcomes. So it’s kind of like our, like. It’s a way to focus the team And prioritize their work. But also to tie the really granular day-to-day movements back to the business. And the user. So, what we’re going to talk about mostly today is around that lowest level, those driver metrics, or input metrics, or system outputs, or whatever you want to call it.
So these are really attractive, because they’re frequent and measurable, as I mentioned before. But they’re also pretty high risk, because they only work if they directly connect to that North Star, so they’re, like, highly gameable. For instance, You know, if you have an input metric, like, Users clicking on my email. And we know, like, okay, if someone clicks on an email, like, they’re probably going to come to my site, and they’ll get value out of the site. Well, one way to increase those email clicks would just be to, like, spam them with a bunch more emails.
And that’s not actually going to end up providing them with whatever user value we’re trying to get them out of, when they come to our platform later on. So that’s kind of what I mean about Gameable. I mean, you could… There’s all sorts of, you know, dark patterns that arise in products, especially consumer products, that, like, will, like, keep people on certain sites longer and engage them longer, that got optimized around metrics, but don’t actually end in user value. I think it’s really common in, like, engagement metrics. We’re not gonna go through how to kind of ladder up that input to Northstar to business outcome causality today.
We’ll actually do this, like, hands-on in that course I mentioned, but I’ll also host a lightning lesson, In January, that goes into some of the methodology, because there’s different approaches to this. There’s stuff that’s, like, pretty lightweight correlation, there’s stuff that’s, like, you know, the gold… golden, standard of, like, an experimentation, actual, like, A-B test, but that’s often not always feasible, and then there’s this middle area of causal inference modeling. So we’ll go through some of those approaches in this upcoming Lightning lesson in January, if you want to check it out.
So, that all being said, traditional metrics and traditional measurement assumes this, like, stable product logic, this deterministic code, and that… so we know that a user input will result in a A, predefined, or at least some, like. predefined set of system outputs. There can still be some variance there, but it’s, like, a more understandable distribution. That we can then more confidently connect to our North Star and business outcomes. Ai breaks that assumption around that, like, really deterministic output, because we’re no longer dealing with fixed code, we’re dealing with a stochastic model.
And so you can have, you know, many possible outputs, many outputs, like, we’ll never see, and then, like, an actual output. This is a challenge, because, you know, when you’re… In the early stages of a product, you might see AI products with, like, amazing demos, That are working on, like, a very small amount of inputs that they’ve honestly, like. may have nailed to some degree. Even those demos from some of the AI products, like, don’t even go well on their kind of, like, predefined, inputs, since they’re still using a probabilistic, system, unless they’re faking it.
But, As your user base grows, as the use cases change, even as it doesn’t, the same input can have wildly different outputs, and so that’s why I think you’re seeing a lot of, like. AI evaluation become really important now, because we’ve had a few years of companies, since ChatGPT came out, since GPT 3.5 came out, start working with this stuff in their product, and some of them have nailed some good, kind of, like, stuff And development, but now as they scale, they’re seeing… all this output that they weren’t able to predict, and they’re not hearing about it in their metrics, they’re hearing about it from their customers.
And by the time you’ve heard about it from your customer, unless you’re working with some, like. customer who’s in a beta and is in, like, an experimental mindset, a lot of the time it’s too late, because people are coming in with this. Yeah, exactly, especially with all the Agentic workflows, which are, like. even harder to diagnose. And so… people are already coming into AI products very skeptical, and so it’s really, really easy. Like anything, it’s really easy to lose trust, and then it’s really hard to earn it back, which is, I think, is probably one of the reasons, at least, like, in my work, where I’ve seen a lot of… teams kind of decide, like, hey, we need to invest in AI evaluations.
I kind of just talked about this already, but yes, this is kind of like… this is where, at least where I see, like, the kind of AI… metric living in the, existing hierarchy as one of these input metrics. You know, some of these input metrics are really deterministic. Previously, some of them are more around, like, model scores that are coming out of things, other existing models But now you’re… you’re ending up with these, like, non-deterministic answers. that kind of screw up and comp- or complicate, at least, the North Star interpretation. So… kind of high level, how I think about this, if that output is… the unit of observation.
Like, if previously the unit of observation for that input metric was, like. The output of a model, Where we could, have some sort of, like, precision or recall or accuracy score. Now I think that output of observation isn’t just the output itself, but it’s the output as experienced by the user. And this makes it really complicated when we were trying to answer the questions, like, what do we evaluate? Because it’s… it really depends on the product and use case and the user. For some people, evaluation may look like Just, like, correctness or faithfulness. For others, it’s gonna be completeness. Some could be, like, safety or tone or style.
But at the end of the day, I think… A really good place to start, for me at least, is just to start with user experience. That’s what I try to ground myself in every time, and then after I can understand the user experience, I can begin to form hypotheses and translate those to metrics. So… process I kind of go through is start by defining the user value in the workflow in the real world you’re trying to improve, outside of your products. Like, name the workflow steps, the thing that the person’s trying to achieve. define what bad would be like when they’re trying to do that work, what good might it look like. So these two answers, these two questions are ones I try to answer at this step.
What’s the job the user’s working to accomplish? And then, hey, how is AI supposed to help? If you base yourself in that. Like, you’re gonna be so far ahead than folks who are trying to take kind of, like, cookie-cutter metric suites from other products that don’t pertain to their users and plug them into their own workflow. Like, if you start from where the user is, This’ll be, like… Just levels up from what most people are trying to do in terms of, like, getting good output. Once you have an understanding of that, you can begin translating it into hypotheses. So, if we increase… this should say AI quality… we would expect to see X resulting in Y user behavior.
So you’re effectively converting those thoughts around what good and bad look like to claims about how it might show up in the product, and then… how the user may behave in the product once they realize that. So what would AI output that’s helping accomplish that job look like? Can we develop a measurement of the AI output that reflects that. And what would we expect the user to do when the AI feature works well? There are many different ways to kind of measure that and judge that. A lot of it comes down to working with some form of ground truth. There’s many different approaches to getting ground truth, and some are gonna be more robust than others.
For me, I typically opt for Ground truth that can be as close to something that’s coming from the user themselves as possible, whether that’s historic user data or some sort of Output within the product itself that you can align to. that’s… if you don’t have users, you’re not gonna have that, so that’s not always the case. This is why we see a lot of subject matter, experts and domain experts becoming really important now, with their, leveraging them for review. Other ground truths could be authoritative sources, so other documentation you’re matching to, or simple rule sets, and how we judge against these can vary quite a bit, too. It could be purely human review, some sort of, like.
rubric across many humans. I’m sure you’ve heard of model judges for scaling some of that human-based review. And even programmatic checks. One more plug for another Lightning lesson coming up in February. A lot of this kind of… Development of ground truth and validating ground truth and review requires some form of annotation. at least what I’ve found is that the best kind of way to do these annotations and just remove as much cognitive load and effort from that annotation as possible is to reflect it as close to the user experience and the product itself as you can. Like, if you’re annotating chat, for instance, like, having some sort of UI that has a chat interface.
It’s a lot easier than trying to flatten chat into roles on, like, a Google Sheet or a database. So I’ll… we’ll cover this in depth in the course as well, but I’m also gonna do a little lightning lesson just on, like, how to kind of, like, vibe code your own custom annotation UIs for your product. So check that out. Okay, back to… the regular programming. So these kind of, let’s see if I go back here… These, like, judgments or ground truth options, the ways you judge these can vary in different ways. You can create metrics that are structural. simple checks.
You can create ones that are similarity, like rouge or blue or some, bespoke version of that, but you’re looking at similarity between texts, or you can do some sort of judgment checks, like human review or element of the judge, to understand if they’re good. I think the main takeaway… though, is that you probably want to use more than one of these things. Even if you find out that you can use similarity… similarity scores that work really well for your product, you’re still going to probably want some sort of semantic check as well in order to diagnose them.
So you’ll have… you’ll end up with many, probably different metrics, and you’ll have to figure out a way to basically prioritize and filter them. a great way, I find, for doing that is kind of now returning those metrics back to our, kind of metric hierarchy, and… in identifying at least, if not a correlative or causal relationship, a logical relationship between them and a North Star, validated by the business outcome, so we can identify a smaller suite of, kind of, key quality indicators. Alright, we’re getting close on time here, so I’m gonna kind of… Get through these last two.
Slides here, so… Due to its unpredictable nature, even if you’ve gone through that whole, you know, define user value, formal hypotheses, I’ve linked them back to my North Stars, and I’ve filtered and prioritized some suite of metrics I’m feeling really good about, it doesn’t really end because, as we mentioned in the beginning, as… as anything changes with those inputs, The outputs are going to change, too, and new cases could potentially arise, or that… or you may even… You’re just not going to have the foresight to basically account for every kind of definition of what the user’s needs are up front in the beginning, and so it is a bit of a continuous process here, not a one-time setup.
Which isn’t necessarily the case with regular kind of product metrics anyways. And so where that lives and the outcome here… you know, I think you kind of insert this new loop. Your system output isn’t going directly to inputs as well into your North Star metric, you’re creating this whole another loop of defining how do we interpret that system output, rather than directly correlating them to North Stars. And if you do that, I’ve found, like, you’re able to iterate a lot faster. Because you create a metric that people can look at.
in the leading day-to-day, you get better alignment across your team, which just makes conversations so much easier when you’re trying to figure out what you’re all working on, and you also are able to set up a lot more safety nets for these launches, as people are a bit more risk-averse to AI products. Okay, nice. This is good on timing so far. So, few upcoming things, Yeah, check out the course, feel free to, yeah, connect with me on LinkedIn, I think my email is on that course page as well, if you have… any questions, or you see if you even see anything in the syllabus that you don’t think is covered that should be. This isn’t until April, so not set in stone.
You’ve got the promo code there. I also write on Substack with my colleagues, Hai and Sravya. We both run that podcast, Data Neighbor, so we talk a lot about data in general, we post our podcast episode kind of readouts on there, and then we’ve been talking about AI evaluation a bit lately. We’ll probably be talking about kind of, like, AI-powered analytics as well, a lot more on there, too. We’ve got that new series I mentioned coming up in February, so we’re talking to the VPs and leaders from, like, Brain Trust, Arise, PromptLayer, Traceloop. A few others, Compose, though. Check that out.
And then we’ve got those lightning lessons, I mentioned around validating the impact on the business and creating custom UIs. Sorry I didn’t leave enough time for questions. I did want to pass it over to, Attendee, because she has a really great course.
Attendee: Oh, thank you, Shane. Thank you for leaving time for my shameless plug here. I also teach AI Evals and Analytics course on Maven. It’s called AI Evals and Analytics Playbook. Where we share our battle-tested playbook that we use in practice and industry on evaluating different types of AI products. Yes, so you’re all welcome to join us. Our next cohort starts in January, January 16th. It’s going to be a two-week course. We have one session on Saturday, one session on Sundays for two weeks. So, four sessions in total. Thanks, Ty.
Shane Butler: And, Attendee, you have a great.
Attendee: Substack as well, so… I’d mention…
Shane Butler: Bone them on that.
Attendee: Yes. Oh, thank you for mentioning that. I’ll also drop it in the chat. I also share a lot in Substack of my experience working on AI evals. So… You’re all welcome to join the conversation. Feel free to, connect me with… connect with me on LinkedIn, and also comment in my Substack posts.
Shane Butler: Sweet! Thanks, Attendee!
Attendee: Thanks, dog.
Shane Butler: I’m gonna close it there, we’re at time.
Attendee: We’re at time.
Shane Butler: Probably that was a bit of a speedrun, but thanks for joining, and if you have any questions, Connect me on LinkedIn, or email me, or whatever.
Attendee: Or email me, or whatever.
Shane Butler: Thanks!
Attendee: Thanks.
Shane Butler: Everyone’s.
Attendee: Thank you for everyone.