← All free workshops
Free workshop · Friday, January 30, 2026

Validate Business Impact of AI Features

Live on Maven, Wednesdays at 10 AM Pacific. About 90 minutes.

Transcript

Auto-transcribed from the live session and lightly cleaned. Attendee names are removed; their questions are kept.

Shane Butler: Alright… Welcome, everyone. to… Validating business impact of AI features. We’re gonna… we’re gonna get started pretty soon here, As people roll in. But, let’s… I do have some questions. First question number one, in the chat, let’s wake up the chat a little bit. Can you just drop in the chat if you can see this screen? If you can hear me… And bonus points if you say something other than yes. What’s… where are you located at? I’m located in, South Lake Tahoe, you can say where you’re located. What’s the weirdest thing on your desk right now? Oh, Vienna, Austria? In Dubai? Paris? Okay, this is way cooler a group than I thought. I was like, everyone’s gonna be like, I’m in San Francisco.

Oh, this is awesome. Okay. Boston, Texas, nice dang, you guys had it rough lately. You guys have more snow in Austin than I have in Tahoe. Belgium.

Hai Guan: This is cool. So international.

Shane Butler: Yeah. Oh, you guys are all, like, up late for us, too. Is it late or early? I think you guys are later, yeah. Like, 8 hours or something. More than that. 9, okay. Dang. That’s past my bedtime, Attendee, so I appreciate you coming on. Okay, cool. 12 a… what, 12 AM? Alright. You really want to validate business impact of AI features, I like it. Sweet. Okay, question number two. Okay, so I’m Shane, I’m here with my, colleagues, Hi and Sravya, And, okay, question number two, who has no idea who we are? Who has no idea who Einstrappy are? Or you can say, you do know who we are. Any… Anyone know? Or you don’t have to say anything. That’s fine, too. So I’m Shane, I’m a principal data scientist at ENTRE.

We’re an AI legal tech company. I lead AI evaluations there for our data science team, also an AI educator and instructor with Maven. Which I assume you can all, kind of figured out. So I have two courses here on Maven, AI Evals for Product Development, which is kind of what we’ll talk about today, and then, another course on AI analytics for builders that I run with Hai and Saravia. It’s about kind of, teaching folks how to build agentic workflows and use AI tools to effectively, like, offload a lot of data science analytics tasks. Maybe… hi, Saravi, do you want to say… Quick what’s up to the room.

Hai Guan: Yeah, hey everybody, my name is Hai. I lead data at the same company as Shane at ENTRE, and I’ve had, probably almost two decades of experience leading data science teams across Silicon Valley big tech companies, names like LinkedIn, Pinterest, Meta, things like that.

Sravya Madipalli: Hi, I’m Sravya, I lead a data science team in Superhuman, previously called Grammarly, and pretty excited to be here, and also, you know, having another course, with this group here called AI Analytics for Builders. Yes.

Shane Butler: Sweet! Thanks, guys! So… Hi and Sabi are gonna be monitoring the chat. If you guys have questions as we go through, or comments, I mean, we’ll try to make it as interactive. as possible, as much as you can for something like this. But, got questions, drop them anytime. You can kind of, like, plus one, or smiley face, or whatever, or downvote, whatever you want to do. I don’t know if you want to thumbs down someone else’s question, but you can. And hi and Saravia will monitor that, and they’ll, stop me, otherwise I’ll keep rambling. This is gonna run for… we’ve got it scheduled for 60 minutes. My next meeting’s not till 1.30 Pacific, so I can stay about 30 minutes over.

This might honestly go over anyways, depending on how many questions we have during it, but… I want to try and get as much of your questions answered as we possibly can, so… Let’s jump into it, though. So, yes, validating business impact of AI features. So, what we’re going to talk about today, AI evaluations, AI quality, I think a lot of folks, I assume, in this room have heard about this over the past, you know. 6 months to a year, as a lot of these companies who are building AI products and features, who have, you know, really, like, pretty cool demos, things work pretty well, start to scale their products, and, start to see, oh, shoot, like, I don’t really have a measurement around, like.

quality here. So there’s been a lot of, kind of, content, literature, and courses, coming up. Actually, I think there’s more… I was talking to, my friend Stella the other day, who also teaches the AI Evals course on, on Maven that I recommend checking out. But we were talking about how there’s so much, actually, like, content right now around, like, how to do AI evals, but then there’s not that many people actually doing it. So, a lot of the content right now is kind of focused really on that AI quality piece. The focus of this is then to connect that Up the chain, too.

business impact, stuff like growth and revenue, and… If you think about it, the reason for that is because you’re… Come on, slide. There we go. Your CFO, or your exec team, your leadership team, your investors, your customers and users, honestly, like, they don’t care about, like. your F1 score, or precision, or recall, or, like, your success rate of, like, some obscure… A slice of, like, an error analysis of, like, oh, my, my, we increased the precision for this very, bespoke part of the product to answer this question by 3%. Like, no one cares about that, besides, like, the team that’s directly working on that. And what they really care about, like, all those people, I mean, like.

leadership, execs, investors, like, they care about, realistically, you know, did this increase profit? How much did it increase risk? Because they’re trying to make a lot of decisions around investment and prioritization. Users, what do they care about? They really just care about, you made their life easier, you made their life better, you provided them some value. And so. the critical answers, the critical questions we really want to answer as we think about AI evaluation and, you know, evaluation of products in general is… is not just, like, this quality aspect, but, you know, what changed in the product? What changed for users, and what changed for business.

So, I wanna get, I want to get a little idea of, like, what’s… where everyone’s at today, so what I’d like to do is if you could kind of drop… drop in the chat a 1, 2, or 3 here. And I’ll tell you what these kind of relate to. So drop in a 1 if you’re… and this doesn’t have to be, like. you’re with your team at some big company building shipping AI features. It could be a personal project, a passion project, whatever, but drop in a 1 if you’re kind of building, or playing with or developing some sort of AI feature, but, you don’t really have, like, a metric or metric suites to understand, like, is the quality good or not.

And then… you know, if you got two, two is like, okay, you got the quality metric, but you don’t really know, like, okay, I tweaked this change, the quality went up, but I don’t actually know how that translated to dollar actual value. And then number 3 is, like, you got features… You’re building, you understand quality. Your comments. Okay, ones and twos… I expect a lot of ones and twos. If you have… if you have some threes, then you can help me lead this session. That would be great. Cool, yeah.

Hai Guan: Boom 1 and 2.

Shane Butler: 1’s 2 is… 1-2 is great. There’s also zero, there’s… 0’s okay, too. 0’s like, I just… I’m here, I’m not building something yet, and I want to learn. 1 and 2s are really good. So let’s talk about, kind of, those different levels, though. Our goal today. is to provide you with a framework, a simple framework, to go from level 1 to level 2 to level 3. So if you’re level 1, you’re gonna climb all the way up. If you’re level 2, you’re gonna get up to 3. We’re trying to get you from… Level 1, which is where you’re… you don’t hesitate metrics, so you’re effectively kind of shipping blind, in a way, right?

You’re… it’s hard to make kind of confident decisions around investments in a company around AI features when you don’t have any sort of Measurement around the quality, because you’re not sure, like, what to work on, The rough thing, honestly, with shipping blind is, like, you will… get AI-quality feedback, but it’s gonna come in the form of qualitative feedback from customers and users. Usually, the people who give you that qualitative feedback are, like, the pissed-off ones, so it’s, like, not even… it’s not gonna be fun. And then, in the worst case scenario, it’s gonna come in the form of, like.

User, or churning user, you know, So, shipping blind, it’s, it’s like… Building’s, like, the first stage, obviously, but we really want to get to level 2. Level 2, I think a lot of the existing content and curriculum around AI evaluation out there right now is getting folks from Level 1 to Level 2, and it’s a really, really important step. You need that before you go to Level 3. Level 2… you’re, like, AI quality confident, you’re… you’re iterating and developing, offline, and seeing quality go up, But then, when you ship it out, like, you’re not actually sure, like. in, like, a month, or three months, or 3 quarters, like, what’d this actually do to our bottom line? Did it expand users?

Did we acquire more users because of it? Did we get better retention? Did we increase our profits? So that just makes it hard, because with every company, there’s, you know, tons of things we could be investing in. We could be investing in AI features, we could be investing in UX, we could be building new products, we could be iterating on existing products and features. It’s really hard to make those decisions and calls when we don’t know when what we built is good enough yet. A lot of that can be achieved by validating up to, like, business impact. And so, level 3, not a lot of folks are at level 3 today, not just in this room, but, like, in the entire industry.

and even companies where, like, they are at level 3, they’re at level 3 at, like, a few features, of, like, many features, like, the huge companies, like metas and stuff like that, like, they’re at level 3 for some… for some stuff, but not everything. And this is where you can, like, really confidently connect here, like. AI improvements to business outcomes. So that’s where we want to get you today. And the framework? we’re gonna use to get there is something that we call the AI Impact Chains. This framework It’s… there’s 4 links in this chain that bring you from, like, an AI feature change to… Eventually, revenue. Maybe revenue via some other business outcomes.

But there’s kind of 4 links in this chain. There’s your feature change. the first link, then there’s AI Quality, And there’s customer value. And then there’s business impact. It’s not super mind-blowing, crazy chain. Has anyone kind of, like, seen anything like this before? You can drop in the chat. Yes or no? Any sort of, like. This sort of chain from feature change. All the way up to business impact, with some of these links in between. Okay. That’s good. then we’re here to learn. So… You need every single link in this chain to get from feature change to business impact. Anything that’s missing here, and you’re just kind of guessing.

what your business impact is, or what your customer value is, what your quality is. Yes. form of the pyramid, maybe, like, North Star metric pyramid, where you have kind of, like, your business impact, your North Star metric, which is kind of customer value, input metrics at the bottom, which is, like, what AI quality in this, yeah. Attendee, good call out. So… If you get all these links to hold together. Or whatever, if your whole pyramid is in place.

I like chains to pyramid because you can break a chain and… No, pyramids, like… I don’t know how… you could have missing things, but if it all holds together, then you’re proving with evidence and some level of confidence, that whatever feature change you’re making to your AI product or feature there is actually leading to some business impact in the long term. Okay, why is this important? So… Why can’t we just have… feature chain AI quality. So… AI quality is not necessarily customer value. It’s kind of like… you think… I would think it is, like, quality, customer value goes hand in hand.

But if we kind of really think about what types of metrics these are, it becomes very clear how they cannot be connected. So, when I think about quality metrics, and there’s probably more than this, but if I were to roll it up into, like, 3 categories. And today, we’re not going to go into how to create quality metrics, but we did a free workshop about a month ago, and you can go to DataNeighbor.com, and you can watch the recording of that if you want. It’s called Redesigning Product Metrics for AI Features. A quality metrics, they kind of take these three forms, I think. There’s, like, a correction rate, so this is, like.

your AI output, how close did you get to what the user actually wanted? And so it’s kind of like, Keep, like, an edit distance. It’s usually, like, a range. from 0 to something, maybe 0 to 1, 0 to 100, whatever. But it’s, like, some kind of, like, closeness metric of how close you’re getting to what they want. Acceptance rate, this is, More of, like, a 0-1, like, Is the output you did. Is the output your product created? Accepted by the user, or they say, like, oh, this is 100% good. Move on to the next step in my life. That’s like a proportion, usually. It’s like, okay, 70% of the output was just, like, one-shot good for the user. And then there’s error rate.

you’ve… if you’ve kind of read some of the existing AI eval, literature out there, or some of the other courses out there, they spend a lot of time focusing on error rates through, like, error analysis. It usually takes the form of working with a subject matter expert, creating some ground truth. going through and annotating a bunch of AI output, identifying, through those annotations categories of things that went wrong, and then… forming, like, precision or precision recall F1 accuracy metrics around, like, are we reducing these very, specific types of errors? It’s kind of more like debugging, I would say, similar to, like, MLOps debugging.

Really important, it’s extremely diagnostic, it can help you get better customer value, it can help you increase your acceptance rate and your correction rate. That’s just one very small piece of the puzzle. And… These three things do not necessarily mean customer value metrics. So when I talk about customer value metrics. Depends on the product. Say, like, we’re gonna talk mostly about, kind of like, SaaS product today, or more like a, kind of like a workflow product, where it’s like, hey, we’re just trying to make someone’s life easier, trying to offload some of their work for them, automate it, or make it faster.

Obviously, there’s other types of value you can provide, like, I don’t know, Netflix or social media or something like that, where it’s, like, engagement, entertainment-based. We’re not really gonna get into that. Mostly because I don’t like social media that much. But, customer value metrics, so today we’ll talk about, like. Time to finish a task. That’s, like, a very clear workflow, customer value. Are we speeding up something that would take you longer before. That’s success. That’s customer value. task success rate. It’s like, just like you go into your job, you’re trying to get something done, you succeed or you fail, are we increasing… the number of times you succeed.

Is our product increasing the amount of times you’re successful? These are more around, like, kind of customer value metrics. And so you can… kind of see where you could end up with AI quality metrics, especially around, say, like. error rates, but also correction rates. Like, error rates for sure, where it’s, like, our output is, I don’t know, constantly, saying that the sky is turquoise when it should say the sky is blue. That’s an error. Let’s, like, fix that. But the customer, for whatever product is, this is a terrible example, but maybe they don’t really care about that distinction at all. That doesn’t actually… Give them a negative or positive experience.

But in terms of your output and error rate, that this could show up as a metric, and there’s many different metrics that can show up like that. You can also have, like, not just, like. wrong metrics around AI quality, but, like. over-optimized, like, trying to make them perfect. Everything doesn’t have to be perfect, like, a correction rate, for instance, you don’t have to get to 100% one shot all the time. Sometimes, like, humans, like, sometimes they wanna, like. hold the reins and, like, know they have some control over what they’re doing, too. Maybe you just have to get to 70%, and that’s what gets you the incremental customer value.

Effectively, quality metrics, they only matter as so long as they are providing incremental customer value. And so you have to have That connection in the chain. Customer value, then… Only matters if it… Okay. We should always bring our customers and users value, but there are ways we can create customer value or user value product metrics that don’t actually reflect true value in the real world, and don’t actually… consequentially lead to business impact.

Like, an example of this would be something like, a lot of companies over… might over-optimize for, say, like, engagement, so… like, Amazon probably definitely doesn’t do this, because they have their shit together, but I was on Amazon a couple weeks ago. I bought… I recently bought this, like, Ooni pizza oven. I don’t know if you all have heard of these things, they’re kind of like, wood-fire or gas pizza ovens to, like, try and make, like, those Italian-style pizzas is, like, one of my things I’m trying to learn this year. And so, for these pizza ovens, you have to have this thing called a peel, which is, like, basically a big paddle.

metal or wood paddle with a stick on it, and you, like, shove your pizza in, and then you take it out, and you’re like, oh, it’s, like, kind of burnt, and, like, turn it around, put it back in, out and in, because this thing’s super hot, it’s, like, 900 degrees Fahrenheit. And, It’s like, okay, I gotta get this friggin’ pizza. peel thing, and I go on Amazon, and, like, it’s just crazy, like, they have… there’s, like, ones that are $15, and then there’s ones that are, like, $125, and they have… tons of reviews, there’s, like, hundreds of these things, I don’t know which one to buy, so… to them, on the back end, it’s like, oh, this dude’s, like. really highly engaged with our site.

He’s going all over, clicking on things, but really, I’m just kind of lost, and I don’t know what to buy. Anyone ever kind of experienced that in terms of metrics, where they have created some sort of… user value metric that they think kind of, like, is reflective of… engagement’s a big one, of something… good behavior the user’s having, but it’s actually, like, a bad behavior. Anyone run into… I know high in Saravia. We have both argued that because we worked together at a company where that happened many a time. So this is kind of, like, where business impact comes in, because… It’s really, really easy to make up messy metrics. We’re actually gonna do a free workshop on this.

It’s called Designing Metrics That Matter. with high, I think it’s, like, February 18th. Again. Go to DataNeighbor.com, you can kind of see us. I would just say go to dabernier.com and in another window right now, or hi if you want to drop in the chat, because I’m gonna, like, throw out These names of other free workshops that are upcoming. You can check it out. See if there’s something that resonates with you. Usually treat moving one step ahead in the journey is positive. Yeah. exactly time spent. Time spent is, like, one of those things, just like… it’s so tricky. It’s like, no, we want to decrease time, but then there’s all these platforms that also want to increase it.

But yeah, that’s a good point, Attendee. I think moving through task success and getting through the workflow is definitely, like, a very positive, thing. So, I’m… I’m very, very sure that Amazon’s not optimizing on how much time someone’s on the page. It’s probably, like, add to cart. put in your billing checkout cart. It’s probably what they’re more optimizing on, but just an example. So, business impact metrics are really important because they can validate if your customer value metrics are good or not. You can correlate those, you can identify causal relationships, and you can see, oh, if this customer value metric’s going up or down.

what’s it doing to my business impact metric, my business outcome in the long run? If those are also going up, that’s probably a good sign. Like, if you’re… moving customer value in such a way that you’re expanding, like, the seats of a SaaS product, or retaining through, like, increased renewal rates, or getting more revenue per account, or if you’re lowering costs, or reducing the need for, like, your CS time on cases. that’s, like, a really good signal that you are truly providing customer value. It’s just very tricky because it’s a lagging indicator.

Which is why you need this, Kind of… Four links in your chain to get from, like, the very leading indicators around feature changes all the way to that lagging indicator of business impact. Common failure modes, so this is just kind of stuff that comes up that, like… It seems like they’re positive sentiments. But you’ll… can end up being… getting a lot of pushback on it if you don’t have numbers to back it up. So things like, hey, the model got better. Ai quality seemed to increase because of these metrics, but, I’ve definitely been questioned by leadership, where it’s like, what does that actually mean for the users, though? Or, like. So how… how good is it?

Like, how… where do we have to get it to, to where the users are going to be happy with it? Users seemed happier, like, if you get, like, a bunch of qualitative feedback, that could be, like, highly biased. You know, it depends who you’re talking to. For instance, if you have, like, hand raisers who are, like, in a beta or something that you’re talking to. Maybe they don’t want to be as negative with you, again, like, makes it still really hard to make, like, investment resourcing and prioritization decisions just based off, like, users seem happier.

And then, kind of like the other side of the spectrum, if you don’t connect back the chain, if you say, like, oh, like, our… revenue increased, or our expansion growth increased. Basically, like, you run into these scenarios where a lot of people like to take credit for those situations, right? It’s like, okay, was it… was it your future change? that increased revenue? Was it another team’s feature change that increased revenue? It’s really hard, because it’s super lagging. This would be, like, 3 quarters later. Marketing sales, like, we did these marketing campaigns. That, that increases. Sales, like, oh, no, we, like, really, like, grinded, like, we were the ones who increased revenue.

Well, sometimes it has nothing to do with a company at all. I worked at a company, we were… Shravi and I worked at this company next door, it’s a neighborhood social media app, and I remember there’s multiple times where Suddenly, super engaged. Everyone’s, like, daily active users. posting everywhere. It’s ad-supported, we’re getting a bunch of revenue from this. Everyone’s claiming, like, oh, this thing I did, like, a few weeks ago and rolled out, that must have done it. There’s, like, a campaign going on. And then, like, every time, it was like an ice storm in Texas. Yeah, we had a dude from Austin here, like, it’s like, I bet you neighbor… I bet you Nextdoor blew up this past week.

So it’s just like seasonality, external events, a lot of stuff can move that, so you really have to have a strong understanding of these… of this change. Future change, what do we ship? AI quality, did they applicate better? Customer value, did users complete Work faster or better. Business Impact. Did profit improve without increasing risk? So… pretty straightforward framework. That’s, like, the thing with frameworks, right? It’s like, it’s like… Once you see them, it’s like, oh yeah, it’s like… Of course, that’s, like, a no-brainer. how we should think about that. If they seem… if a framework is, like, convoluted and hard to understand, it’s not a good framework.

But also the thing with frameworks is, like, when we don’t have them, it’s kind of really hard to just, like, piece it together from nothing, because we’re all focused on a lot of other stuff. We’re all focused… Like, a lot of us are focusing on feature change AI quality. It’s really hard to find time to think about Customer value, and then if you get these Three parts of change, it’s really hard to find the time to carve out for business impact. So… Fairly straightforward concept. As most frameworks are. Alright. About 30 minutes in. We’ll keep chugging. I want to walk through this in terms of a kind of example here. So, in this example.

Let’s imagine that we are all on a team building a product. It is a… Agentic AI data. scientist, analyst, converts, like, text to SQL. And the reason I picked this is because we have a whole course on this, AI Analytics for Builders. Hi and Ravi and I actually, on the side, we also do this podcast called Data Neighbor. Maybe, hi, you can drop the YouTube link in there or something. But, the past few months, we’ve interviewed bunch of, CEOs and leaders and, VPs of, like, BI and Agentic AI… Agentic Analytics, companies, much like what we’re gonna talk around here. So, if you’re kind of interested in this space. A little different than AI evaluations, but around agentic analytics.

check out our podcast. Over the next couple months, every week we’ll be releasing an interview with, you know, people from, like, the VP of, VP of Tableau, to, like, CEOs of, like, startups in the space, like LiveDocs and Count, to a bunch of stuff in between. Alright, but back to our use case. So we have this AI Data Analyst text-to-SQL product. If we think about it from the user workflow side, the user goes in. They ask a question in the chat box. AI. writes some SQL. Presents it to the user, they check if it’s good or not. If it is… It runs the query.

Then it develops a chart, And then it exports that chart, and the user has this nice… answer to his question, or her question, in chart form. They can share it with someone else as evidence, or whatever it’s needed for. Maybe it’s like… Yeah, there’s a very important customer talking to. an account executive, and they need to, like, get chart relief asked for them. That’s definitely happened to me in the past. So there’s many steps for, like, failure in this, so I could… A user could write the question, and then it could return SQL, but it could be totally wrong, right? Wrong, like, group buys, wrong tables it pulls from, wrong information it just pulls.

It could get it right, but it could have syntax errors, so it doesn’t… actually run, it could get it right and run, but the chart could be total trash and just, like, not be helpful at all, so now I have to go take this data and, like, go into, like, Excel or Tableau or whatever, make my own chart. Maybe I have to go back and forth and, like, edit the query myself, ask the question a different way, or I can go perfectly, and I can ask a question and get my chart and be on my way. It’s like a seat-based, we’re gonna say, like, it’s a seat-based model, so, You know, for every additional seat we get.

for people within an account who are able to use this product at, like, for our customers, you know, $200 a month. And, the current state, what’s happening right now, is that there’s a new version of this. There’s a V2… of this, text-to-SQL AI Analyst, and what we all need to do right now is figure out if we’re gonna roll out V2. And so, I’ll tell you a little about V1 versus V2. V1, it’s like direct SQL generation, there’s no kind of planning step involved. There’s no really grounding in, like, the existing data warehouse schemas or anything, or context around that. No validation, just kind of, like, does a single loop thing.

And, you know, because there’s not, like, this planning step or, like, validation step, it’s, like, a lot lower compute. V2, this new thing we’re trying to roll out, there’s some reasoning steps when it generates the SQL. It actually creates a plan before it generates SQL. So, like, maybe it figures out, like, hey, here’s all the different CTEs I’ll have to get to, here’s things I’ll have to filter out. There’s more context around the existing schemas in the data warehouse, and there’s another agent that, like. does a double check, validates, repairs to SQL. Obviously, a lot higher compute. So, let’s imagine we’re all… Remember those level 1, 2, 3? Level 1, we’re building, shipping.

Level 2, we know AI quality. Level 3, we know about business impact. Let’s all imagine we’re level 1 right now. What do you guys think? Right now, this feature, V2, it’s out to 5% of people. You kinda know some stuff about it. Or do we want to roll it out to more people? Do we wanna roll it back? I’m gonna stay at 5%. V2V1, I guess, job in the chat. Which would you… Which would you prefer right now? V2… B2 sounds pretty nice. Very smart. Okay. Let’s get a little more information about V2. And V1. So, we won’t get into that quality yet, but we do have some latency figures. So, latency is how long It goes from this first step, ask a question, to AI writes SQL. How long does that take?

V1, it took about 30 seconds, and then V2, now it takes 2 minutes. And latency is really important, because you can imagine yourself in the… in a position, kind of like what I said, like, maybe… like, there’s some customer that’s gonna churn, it’s your biggest customer, and customer support’s talking to them, or the AE’s talking to them, and they’re chatting with the CEO, and the CEO is like. what do they want?

They’re like, I don’t know, they want… they have all these data to ask, they have all these things they’re asking us for right now, like, we just gotta give them what they want and make them happy, and the CEO is now going to the product teams and be like, you need to pull this data for me, and, like, you’re in the middle of building something else right now. You know, from the customer point of view. And so… the customer wants us to use this AI analyst to, like, pull that data. for the person who’s complaining. Yeah, the longer it takes, like, the more, kinda… pissed off. They’re gonna be… what do you guys think now?

So the… We know that the response time… we added a bunch of, kind of, planning stages and validation, but now it takes 4 times longer to Answer the question. Yes, and we got V1. speed… it’s kind of like, what do they say? It’s like, you can compete on speed, quality, price. Speed’s important. Maybe V1 was contextual questions for V2? Yeah? Okay, let’s keep going. What’s the… oh, man. Attendee, you’re… you’re skipping, like, 5 slides, man. You’re skipping to the next chain. I gotta build, I gotta build these examples first. But yeah, user experience… I think we know where it’s going. Alright. But we’ll play along for a little while. Okay, so it took longer, but then check this out.

So, it took longer to get that initial query running. But then… the user edits per query. So let’s say we go back to our… Little flow here, ask a question, AI writes SQL. The amount of times someone has to, like, go back and re-ask the question, or kind of edit the sequel and try to run it again… that decreases significantly. So you used to… you used to have to go back and re-ask your question, like, on average, like, two and a half times. And then… Now, it’s taking, like. way less times to go back. So, like, it does take longer to get that initial… That initial response, but… It’s a lot closer to a one-shot. What do we think now? Yeah, good comments in here.

I mean, the duration is about the same thing, right? Yeah, it’s like… yep. Reasoning adds a check, yeah. Attendee. Add a context, tables, columns, plus E1, yep. Alright, I’m gonna skip through a couple of these ones as we build up. So let’s say now we get into, like, you know. it takes… it takes less edits to go back and forth, even though it takes a bit longer. We find that maybe, like, the syntax errors are a little more, so you have to go back and tweak those. And then… what you kind of have here is you have your… your kind of, like, completion, that completion rate, when I talked about how close you can get to it.

You have kind of your acceptance rate, how often does it just, like, actually run through? And then, now we’re getting into some of those, like, error rate metrics. Something like… the precision on a customer support questions. These are, like, the kind of things where, like, you have a golden ground truth dataset, and you’ve gone through and annotated and figured out, like, we categorize these types of questions as customer support questions, how does it perform on them? And we actually found that in V1, It performed… it got it right 50% of the time, like, exactly. But on V2, It gets it right 25%, so 50% increase in the amount of times it’s getting these customer support questions right.

What do we think now, V1 versus V2? Yep. Quality question with weights to just get a number. Yeah. We definitely need to figure out one number. we have too many numbers going. This is, like, kind of, like, the challenging thing with, quality metrics, is that, like. Especially the error analysis metrics, they are never-ending. You’ll constantly be in the iterative loop of expanding your, kind of, failure modes, and you can just end up with an unlimited list, and you really have to assign some sort of weights and prioritize them So, love what I’m seeing in the chat. And then, you know.

we’ll stop playing this game, but, like, you could imagine, like, maybe the customer support questions it gets a lot better at, but then the finance questions, it gets worse at for some reason. Like, it used to, like, finance questions are hard, like, there’s all these industry, you know. Financial calculations, but then within a company, there’s definitely internal interpretations of all those of those metrics, you just filter them and edge case, like, the hell out of, like, financial metrics. And, and so maybe it actually gets worse at those. And now in this situation, actually, it’s pretty interesting, because you’ll end up with, like.

a team of people, like customer support, that could look at a readout that you have on AI quality metrics, and they’re like, hell yeah! Ship V2, like, this is gonna save me a bunch of time. And then finance is like, no way, are you insane? Like, we have to report those numbers out publicly, like, we’re not gonna use this at all if you just decreased it to 30%, 40% already kind of sucked. So, Yeah, I think a couple of you got it in, like, the very beginning, but, like, you need to get to that customer value, so hopefully that kind of drove the point home a bit. That’s why customer value is important. It’s, like, it breaks that tie.

And it helps you prioritize which of these metrics is really important. Best way to do that’s gonna be an A-B test. There’s other ways to do it, too. So, like, some… I hear questions before… we’ll talk about this more, but there’s other ways we can… we can kind of find these relationships, too, without an A-B test. But… Let’s see, like… V1, V2, alright, V1, the time to correct try to export. So, from the question to exporting, used to take 12 minutes, now it takes 3 minutes. V2, accelerated the chart export workflow by 4 times. Really reducing time. Yeah, exactly, Attendee. Yeah, cost is a huge one we ship in here.

And I don’t… I mean, I don’t know what’s gonna happen with, like, the AI industry and all these, like, proprietary models, but, like, I know a lot of them are operating at losses right now, so I don’t know, like, how expensive this stuff gets down the line, so keeping an eye on… on transaction costs is really important. Yeah, so what do we think, V2? B2… ROI for the win. Yes. Okay. So, really cool thing about A-B tests. It’s that tiebreaker there, for sure. What’s actually even cooler about it, I think. and comes into this whole chain thing, is if you do enough A-B tests.

you can actually start to form the relationship, understand the relationship with confidence, statistical confidence, between those AI quality metrics and your, and your customer value metrics. So we’re actually gonna… I don’t have time to go over all that today. As you see, we’re already 40 minutes in. I will stay over for questions, though. So we still have some stuff to get through in the next 20 minutes. We’re gonna have a session just focused on this, designing experiments for AI features, February 19th, totally free workshop, pretty similar. We’ll do today, if you’re… so if you want to join that, if you’re liking this so far.

take out your phone, yeah, scan the QR code, or go to bit.ly Experiment AI, and that will get you… To our sign-up page, you can kind of read about it. It’s like the same sign-up page for this one, so… you can tell me if you’re interested in it, but what we’re gonna try and get there, talk about there, teach you how to do, is basically, it’s like. We can now know with confidence that, say, like, when we decreased edits by these 7.5%, like this metric back here, user edits per query average. maybe that had, like, we run many A-B tests over time, we can draw the conclusion that, like. That actually is a thing that decreased the workflow. time by 75%, but forexed it.

maybe these other ones, like SQL syntax Rate, or Response latency had no effect on customer value. At all. And so, with enough A-B tests, you’re able to kind of, like. Form those relationships, which is pretty cool, because now. You should still run the A-B test. But you can effectively, run your offline AI quality evals, and then predict what you believe the customer value will be when you eventually run an A-B test. Question I get a lot here, is… Should we just run A-B tests, then? No. don’t skip the AI quality stuff. You still need that, because the… I’m sure you’ve heard this many a time. This output from this stuff is, like. non-deterministic, right?

It’s like, it can feel like you’re locked in with the output sometimes, and other times it feels like you’re just, like, pulling a slot machine, so you don’t just want to release this stuff out in the wild. A quality metrics are extremely important, guardrails when you’re developing early before you get something in the customer, and then once you get those to a place that you’re comfortable, it all depends on the cost and risk as well. You can release that and do an A-B test. You can find those relationships earlier, so you know, like, how high you need to get that up stuff up before you do a test. We’ll talk more about that. At then as well, and obviously we’ll have this whole deep dive.

Yeah, so Bitly Experiment AI. No, probabilistic, non-deterministic, yeah, it’s, like, the same. I mean… Yeah. Good question. Now you have me thinking. I mean, non-deterministic in a sense, like, you really don’t know. what it’s gonna say. I guess if you, like, tweaked the temperature and everything of the models, then you could get to probabilistic, but in some ways, I feel like it is, like. non-determinable. But I use them interchangeably. Right. So, we’ve gone to customer value, we’ve gone to feature change, AI quality, customer value. This tells us if we want to ship a feature or not. But we still have to find that last piece in the chain, business impact. Can we validate it?

How do we understand if that feature change we’re making is impacting the business? This is really hard to do, because it’s a lagging indicator. It could be months, it could be quarters, That’s why, like, we don’t ship, like, we don’t… we’re not gonna wait 3 quarters to ship something. We ship off customer value. because we can measure that in days or weeks. We can measure AI quality in however long it takes you to get the output, could be minutes or seconds. But, we still want to know, like. what business impact is, because it helps us with investments. So basically, what I think about AI quality and customer value and business impact.

I think about… AI quality… I use in development, it helps me know if I want to test this with customers, get in front of them. Customer value tells me that I’m testing it in customers and production. It lets me know if I want to roll it out or ship it to everyone else. And then business outcomes, business impact, this tells me what do I want to invest in and prioritize next with respect to the product as a whole. Lots of gray area and overlap between those things, but that’s kind of, like, the three kind of ways I think about it. And the way we do this, Business Impact, is something called… causal inference.

A-B tests are a form of a causal inference, they’re actually, like, the gold standard of causal inference, but there’s Many other methods of causal inference. stuff like… I don’t know if any of these sound familiar, but, like, regression discontinuity design, difference in difference regression panel regression, lots of regression. Regression in econometrics has been around for a long, long time. People have been doing this stuff. Double-bias machine learning, propensity score matching. course into exact matching. There’s lots of different models around causal inference, but what the output effectively gets you is learning the relationship between an input, and an output.

So in this case, we have our customer value. We can ignore AI quality and stuff for now. We have customer value that we have, Yeah, R… I mean, R is like a economet… economists, or econometricist’s dream. Yeah, I started in R, too. But… I went to the dark side of Python. Measuring customer value. We know this, we know that we reduced 12 to 3 minutes, now we want to learn the relationship between that and some sort of business outcome that we can translate into revenue.

Regression allows us to make that translation, so in this case, we talked in the beginning that Each seat that we get into an account gives us another $200 a month for the product, so we’re gonna try and learn the relationship between That time to correct chart export to… The number of seats will get gained by an account that month. This is, like, made up, but this is, like, how it works, basically, like… The coefficient, like, in this case, we could say something like, Every 1 minute faster, We get someone… to export a chart we will… we expect to gain 5 more Cs. This would kind of be the output of those causal inference models. And if you think through this, why does this, like, make sense?

Well… if I can get data faster, I can make decisions faster. If I can be more confident and make decisions faster, then I can acquire more of my own customers, or make my customers happy, or… divert that time that I spent on making decisions or finding data to building other products, and so now I’m willing to buy more seats of this AI analyst tool, because every additional person I have doing this, offloading that analytics and problem-solving work that they had to do themselves before, Is now able to effectively give me more time, resources, and money to put and invest elsewhere in my business.

So you can see, just like anything, as they make that product better, as we make the product better and faster, we could probably expect increases in, like. expansion of our existing customer base at this, like, seat per account level. And then, after you have that, it just becomes a math equation. Like, once you have that coefficient that’s output by the causal inference model. You can do something like, okay, well, if we get 5 more seats per 1 minute faster export. and we increased, exports by 9 minutes, then we get 45 seats more per account. Say we have 100 accounts, that’s 4,500 seats. We know that a seat is $200 per month. So, that translates to $900,000 monthly recurring revenue.

additional, and then if we bring that out to the whole year, now we’re at, like, $10.8 million in incremental annual recurring revenue. And you can put confidence intervals on this. This is obviously, like, made up. I don’t know if people are gonna spend I don’t know, $200 a seat for this, but chatGPT Pro and all these companies do charge something like that.

So you can put confidence in rules, like, it’s statistics, it’s, like, a distribution, so maybe it’s, like, in a range of, like, 8.6 to 13 million ARR, but what’s really cool here is now… Because you’ve connected that chain, we’re able to, like, discern… What the customer value impact is on The business, and then we can move even higher up. to, the AI feature change. So we’ll… I rattled off a bunch of names of causal inference models. Obviously can’t teach causal inference right now. We are gonna have another free workshop on that, on February 26th. And, yeah, that first number is, Attendee, is, like, what would come out of our causal inference model. One minute.

To 5 seats, 1 minute to 1 seat, 1 minute to .25 seats. But that is the crucial number. When you can get that, when you can control for all the other variables that are impacting the acquisition of more seats, when you can understand, like, the relationship in all of the pieces, steps between, between, like, we increased their speed that they got this answer, this person made decisions faster. making that decision faster made their manager happy. That manager said, hey, we need to get this tool for more people. Finance approved it. Like, when we can, like, figure out all that links through a model like this, the rest is just simple math.

So we’re gonna go into, actually, what all those types of models are at a high level. I’ll probably go, like, 3-4 of them, maybe, like, panel regression, diff and diff, and, like, propensity score matching or something. If there’s certain models you want to learn about, you can drop them in the chat, and I can think about changing it too. But take out your phone, scan the QR code, check that one out. Bitly Causal AI, that’ll be in late February. And yeah, why this matters, like… When you get this whole chain now.

when you can identify all those relationships, when you can do enough A-B tests or causal inference analyses to understand, with some level of confidence, the difference… the relationship between your AI quality metric and your customer value metric, and your customer value metric and your business impact, then you can literally change a feature. Get your AI quality evaluation output metrics. That takes you… like, that’s just how long you… it takes for your eval pipeline to run. That could literally be minutes.

And then you’re just forecasting with those coefficients, and you can begin to make predictions of, like, hey, this small tweak I made to this feature change may have this huge $10 million impact On the company in the long run. if we want to, like, get that, let’s put more headcount in this area, or let’s stop focusing on these other things and work on this feature. So it’s extremely helpful, for investment and prioritization, decisions to get this full length through. And… kind of went over this, but like I said from the beginning, the CFO doesn’t care about your F1 score, leadership doesn’t care about that. This is what they care about. This is, like, the slide they want to see.

They don’t want to see, just the… edits went from 2.4 to 0.3. They want to see how much do we predict within some interval, some confidence interval, our AR will go up. how’d you get… calculate that? Oh, this is how 4,500 seats went up. Why are the seats going up? Oh, because we made it 4 times faster? Awesome. this is what… this is what you want to be able to report, or predict, or forecast, and there’s definitely uncertainty around it. But it’s… it’s how you get from level 1 Where you’re shipping completely blind to level 2, where you have some idea around call metrics, all the way to level 3, where you can, like, begin to really have some forecasts around the business impact. Okay.

We have 5 minutes left. I’m gonna stay… I’ll stay up to 30 minutes over. And depending on how many questions we have. But, if you enjoyed this, obviously. mentioned a few of those things we have coming up. Designing Experiments for AI features, Dub19. Designing metrics that matter, FEB18. Was everyone causal inference analysis for AI features? Feb 26 or something like that? End of February? Check those out. we have, like, 10 or 11 other workshops. We’ve got some really cool ones, actually, planned. We have some ones around, doing the experimentation analysis itself. Leveraging… AI, we have some really cool ones about, like, running analysis or designing analysis, with Claude Code.

I have one around, like, building custom annotation tools to help, help you create those error analysis kind of metrics. A bunch of other stuff. Go to DataNeighbor.com, hi, just put it in the chat, check them out, sign up for whatever you want. Also, we have… In addition to all those free resources, we have a 6-week cohort. We’re gonna kick this off in April, where we are going to deep dive into AA evals. From… all the way from just, like, observability instrumentation to, making decisions with, like, your leadership team and everything in between, so think… Developing ground truth, looking at offline versus online evaluations.

Of course, getting into the running experience, going really deep into causal inference. all the, like, common kind of failure modes you run into, how do you get this into, like, a CI-CD pipeline, how do you, yeah, constantly, like, iterate and improve on what you’re working on here. So that’ll be a 6-week cohort, starts April 6th. You can check it out at AIEval.ai, AIEval.ai, or you can just take out your phone right now, scan the QR code. It’s usually 1500 bucks. Discount if you use this promo code, BizImpact. This is valid for a week, so through Feb 5. It’s 30% off. So, knocks off about $500.

I do find that people, like, they want to sign up, and then it kind of, like, they forget about it later, so if you’re… if you’re really interested, and you want to sign up immediately, I’ll give you an extra 10% off today if you sign up by midnight. PT, and just… just for reference, like, Maven… Maven has, like, a satisfaction guaranteed kind of thing, so you can actually drop the course two weeks in and get a full refund, or any time between now and then. If something comes up, and if you have to do something like get, budget approval or something from your work, like, I don’t know, we’re all, like, humans here, like, I can extend the promo for you, DM… just DM me on LinkedIn.

You can find me on LinkedIn, Shane Butler, or you can email me, sean at AIEval.ai. You’re probably saying, that says Shane at AIeval.ai. Yeah, it’s kind of weird. It’s pronounced Shane. Spell shame. That’s my little sales pitch, but I do want to go into a couple questions. I can stick around for 30 minutes. We can talk about anything we talked about here today. We can talk about the course, we can talk about those upcoming sessions. A couple common questions I do get, I have them here, another tab I actually wanted to pull up, because… I feel like these come up a lot. One I get a lot is around, like, hey, I can’t run an A-B test, I don’t have… Enough users.

Or my leaders are, like, really… Oh yeah, there will be a recording, of course, sorry, I should have said that. I’ll send out a recording, I’ll send out these, this promo code and stuff in the recording, too. And yeah, just… yeah, and if I forget for some reason, I won’t forget, I think it’s, like, automated, but you can just DM me or email me, and I’ll get you what you need. So… Some people say, like, we can’t do an A-B test, because either we don’t have enough users… well, first of all, you have to have some users, so if you’re in the stage where you’re building a product, but you don’t have customers or users yet, on one hand.

that’s kind of rough, because you can’t calculate business impact, but on the other hand, it’s like, that’s kind of good. You can just focus on AI quality at that stage, because you can’t understand user impact without impacting users, and if you don’t have users, you can’t impact them. So AI quality is going to be, like, your focus there, and then you’re just gonna get some early users and do more of, kind of, like, qualitative user research discussions with them early on. Don’t focus as much on the metrics at that stage. does usage per user matter? Because if it’s high, then maybe you don’t need too many. Yeah, exactly.

It’s all, like, when you’re… it depends, when you run your experiment, what you’re going to, what you’re going to, like, have your exposure on. If, say, like, you’re… wanna know, like, the… average time to export per user, then you kind of have to have a lot of users, because you have to hit you’d have to have a representative sample of, like, the population of people you want.

If you’re looking at something like time to expert per page visit, then you could… you could use messages, but the problem with, like, not having a lot of users is, like, and you have a power user, is you end up biasing just to how they act, so… I would say, like, you actually, in the case where if you have, like, not that many users, and you have some that just, like, use it a lot, you may… Even want to be, like, segmenting that out, and be like, this is how it affects high usage users, this is how it affects low usage users. That’s a good question, though.

Something else you can do, just like the causal modeling from customer value to business impact, you can do causal modeling from AI quality to customer value. So, an A-B test, an experiment, is just a form of causal inference. It’s like the gold standard, because it really controls for everything. But there is still some uncertainty in there, there is still some randomness associated with it. But… if you’re… if you can’t A-B test because you don’t have enough users, and it’s gonna take you, like. You know, 3 or 6 months to get enough, events to, like. reach power and confidence in your results, then what you can do instead is do a causal inference.

You can basically backfill all of your AI quality metrics, like 6 months. And then you can do your causal inference relationship between that historic data to come up with, what the relationship would be. It’s not… as accurate… as an A-B test, but it is a totally… Accept… acceptable, kind of, replacement for that. And some folks, like, it’s not just, like, if you don’t have enough users, it’s, like, if you’re risk-averse. So, I work in legal tech, it’s, like, very, you know, buttoned up. industry, banking, finance, health, a lot of that stuff, like, really, like, you kind of don’t want to just, like, throw stuff out there.

There’s a lot of risk associated with it, and so doing causal inference before the A-B test, you can still do an A-B test, but have this kind of, like. Higher confidence calls or rinsup before it for really big changes that you’re not sure about, is a nice, helpful way to do that. Another, another question I get is around, like. leadership, or only cares about, like, latency and cost. Which are both really important, right? We talked about price, quality, speed. So, yeah, price and speed, latency and cost, those are huge. The way I approach that is basically… go into the building of your features with some agreement and contract around, like, established guardrails.

Like, hey, we’re not gonna go above this latency, we’re not gonna go above this cost, because we know it’s just… It’s basically, like, the business can’t handle that. It’s, like, not an option. And then, yeah, no problem, Attendee. Yeah, no problem, Attendee. Looking forward to seeing you guys in, in future sessions. So I go in there, I have those guardrails, and then that way, later on, even if latency or cost increases. you can reframe the discussion around, like, hey, yeah, it increased, but we are below the guardrails, so rather than focusing on that increase, let’s focus on the ROI with respect to the incremental gain in customer value and the incremental gain in business impact.

Latency is just a… one proxy for time, right? Like, resolution time is, like, the real outcome, like we saw here, like, the time from… Question to export. Another question… is around instrumentation, this is a lot of metrics. AI quality is a lot of metrics, customer value is a lot of metrics, business impact is a lot of metrics. the other, like, stuff that’s not even there, like, more like those system metrics around, like, latency and cost and stuff like that, that’s a lot of metrics. All that inventing stuff takes a lot of upfront investment. We are gonna have… In our course, An entire week dedicated to instrumentation and opera… operational, not operations… observability.

That’s, like, the first or second week, I think it’s the second week of the cohort. And, So we’ll be talking about how we set that up from… the system to the AI quality, to the customer value, to the business impact, how we can plan and design plans and work with teams to make sure we get all of that kind of eventing and tracking. ready up front, that makes the rest of this analysis possible. Unrelated to our other course, actually, I just thought of this, AI Analytics for Builders. There’s actually gonna be lessons on how do you use Agentic workflows to create tracking specs.

So… If you’re doing AI evals, and you want to try and automate some of the AI evals with agentic workflows, take both courses. Actually, someone yesterday signed up for both, but that was pretty cool. We’re gonna be spending a lot of time together in April with that dude. Alright, what other questions? Is there kind of some of the first ones I usually get? Any other questions? In the chat, or from people… And the audience… Hi, do you got any questions? Anything I… I didn’t mention? What? I keep hanging around here. if folks want to go to AIEval.ai or scan the QR code and take a look at the syllabus and be like, what’s this week talking about?

Feel free to do that, or if you have any questions about the upcoming Workshops, or anything today. Do segments and scenarios come up frequently in AI evals? Segments, yes, for sure. Because, the output of LLM or Gentic workflow can be drastically different from one user to another. It’s not just the output, it’s how they respond to it, too. So, like, their inputs are gonna change, but also, like, something that I might be totally cool with as an output, I may be like, this sucks. Like, so we have, my day job, right? Like, we have… there’s, like, a spectrum of, like. Lawyers that we work with, like, some are gonna be like, yeah, this is, like.

resonates well with, like, my style, and other ones be like, oh, this isn’t my style at all. So, segmenting by, like. by user persona is definitely a big thing. We kind of talked a little bit about it when we went through the AI quality metrics, and we had, like, the customer support question and the finance question, but segmenting is, like, a really important diagnostic part of it. Scenarios… could you tell me more what you mean by scenarios? Attendee, thanks for joining. I think Attendee bounced, but if you watched the recording, thanks for coming. How’d you think of any questions?

Hai Guan: I think, generally, I’m curious about, kind of, like, the audience here. What’s your biggest struggle in the whole AI eval space?

Shane Butler: I’ll answer Attendee’s, I don’t know if I’m pronouncing your… your name right, so I apologize if it’s… if I’m not. Yeah, what’s the user trying to accomplish? Buy a ticket, try to get some info? Is there a… definitely, definitely. Okay, sweet. Yeah, it’s kind of like, you know how… Thanks, Attendee. Happy to help. Attendee, alright. Okay. Attendee, so… Yeah, scenarios… Right? People have different… expectations for different products. So, for instance. Some products, it’s gonna be, like, if you have, like… yeah, let’s… let’s take your example here, buying a ticket.

If you have a Gentic workflow that, like, you want some… want to go in there and say, hey, I’m trying to plan a trip to XYZ, find me the best ticket. You know my preferences, like, I like to get there early, or I like to roll up late, or I want to go when the airport’s quiet. Or I don’t like connections, or whatever. You kind of want that to really kind of one-shot it, because… or maybe, like, ask you a couple follow-ups, but, like. if it takes too long, or if it gives you a bunch of wrong tickets, you’re just like, I could have done this myself by now.

But then there’s other things, like, I don’t know, when I interact with, like, ChatGPT or, like, Claude Code, and it kind of gets me, like, 60% of the, like, 30% of the way there, and I talk to it some more, and now I’m, like, 40, 50, 70% of the way there, and I’m like. Hey, that’s pretty good. That’s, like… 70%‘s good. And so, like, it definitely matters, like, what the use case is in terms of, like, how you’re gonna be building those eval metrics. Some, you’re gonna wanna have, like. A widespread of different error, error rates, different error rate metrics. Some, it’s gonna be really just focused on, like, what was that, like, kind of completion. rate, like, Yeah, that’s a great question.

We’ll definitely talk about that in the cohort, in terms of, like, how do we… think of things for, like, specific, use cases. There’s gonna be kind of like a… We’ll have… we’re also gonna have, like, a lot of time where it’s, like, just talking about what people are building themselves, so some of the projects will be, like, reflections and, like, your own work as well, so we can get, like, into the specifics of your use case. Any, any other stuff? Yeah, I know Hai said, I don’t know if there’s anything that people, like, are really struggling with. What tools are you using for annotations or analysis? Okay, this is a very good question.

So… First, our podcast thing we do, Data Neighbor, that I think I dropped on YouTube earlier. We’re doing this other series, pretty cool. We’re interviewing all these, CEOs of these, AI evals tools, like, CEO of Braintrust, we’re talking to, like, Senior Director of Product at Rise, we talked to their head of one of their heads of product, back in August, talking to CEO of TraceLoop, Composo, prompt layer. So there’s a lot, and I think I’ve got a few more on the docket, too. But I think a lot of these tools, something that I have found, like, doing annotations, like, these tools are really good at, like.

getting you the trace and, like, helping you instrument the data of, like, all the different steps within, like, your LLM calls and, like, the agentic workflows. They’re really good. Kind of, once you have the annotations in place, and you’ve created, like, your, your different, like, error analysis dimensions. They’re really good at helping you, kind of. iterate on prompts to, like, turn it into, like, LM as a judge, or, scale it a bit. annotations thing, I think a lot of things… these things are missing. The annotations, from the way I see it. In order for a subject matter expert to annotate the output of AI, especially when the AI output is effectively the product itself.

the user interface they have to go through should really reflect the same user interface that the product has for the user. So I’ve tried to build annotations through those products. A lot of it’s, like, kind of like palms and rows in a spreadsheet, and I’ve also tried to do annotations in, like, a Google Sheet before, and it just… It’s really a lot of cognitive load for the subject matter expert, so I can tell you, like, a couple of examples. One is, like, if there’s two blocks of text. like, the… the output, and then of our… of our AI output, and then, like, our kind of, like, golden standard output, like, what the user wants.

in these tools, like, or in, like, a Google Sheets or something, it’s, like, just huge blocks of text, and, like, 95% of the time, the subject matter expert has to just decipher, like, what the difference is, and only 5% of the time, probably, are they saying, like. why it’s good or bad. And so a very simple way to create a custom, like. annotation around this, like, which is what’s reflected, like, more in, like, products, is, like, if you, like, have, like, red lines for, like, the differences, just like when you’re, like, in a Google Doc, and you have, like, the suggested edits, and you have, like, the red lines, and then the green additions. That’s not gonna be in these products.

That’s, like, a pretty big one for annotation comparison, but what… I’ve done, and colleagues I have done… worked with, is basically 5-coded custom annotation UIs. It’s really, really pretty easy. We’re gonna do a workshop on it in March, I think. And we’ll go into it in the course, too, and in the course, I’m going to provide a repo that has, like, a few key examples of custom annotation tools. But, you basically… can make the backend pretty reliably. Like, all it has to do is, like. read data, annotate, write data. Like, those are pretty 3 simple things to code in the backend. And then, like, I don’t know shit about front-end.

development, but, like, you can vibe code all that kind of thing, pretty easily. It’s pretty… it’s pretty amazing how you can make your own custom UIs. Another example is, like, chat or something. Like, you just… you really can’t annotate chat well in these, in these annotation tools or in a Google Sheet or something, because there’s so many… turns in the conversation, you just can’t really accurately flatten it out into a bunch of, like, rows and stuff. So that’s another really good one where you can, like.

vibe code, a custom annotation UI, where your subject matter expert can come in and actually be chatting With the product as… with… just like a user would with the product, but then on the side have, like, an annotation, Kind of box where they can, like, say, like, oh, this response wasn’t quite right, and… If I connect back up this response. Earlier, you can see it’s, like, diverges from the conversation. So that’s my take. Yeah, and hi, drop the link in the chat there to our upcoming session on building custom annotation tools. Did that answer your question, Attendee? Cool, cool. I can stay for, like, 14 more minutes. And then I gotta hop, but any other questions?

Yeah, I think Hai had a good one around enough questions for the audience, like… It’s like, what’s the biggest challenge right now? Trying to think of what the biggest challenge for me right now is. Finding the time, I think. It’s like, it’s moving so fast, like, building product is, like, feels like a lot. Faster and cheaper now, with all, like, the productivity tools, but then, like. it’s like, you gotta find… find that time to set up the entire evaluation workflow, but it’s really worth the investment. In my day job, like, I probably spent 4… 3 or 4 months iterating on our AI evals, like, pipeline.

testing out different quality metrics, relating them back to customer value and business outcomes before I could get all those weights and then really figure out the suite, and it honestly wasn’t until that was all set up and did all, like, that kind of, like. early investment that I was able to, be like, okay, now I can actually analyze the data, and it’s like, oh, it’s not working for this segment, or that segment, or, like, we need to get to this goal. But once you have that all set up, it’s pretty cool. Cool. Alright, I’m gonna give it 2 more minutes. Oh, here we go, here we go. Stay alone. Figure out what are the bottlenecks in NIVLs. Yeah, yeah. Sweet Dom Dom. Awesome.

Hope to see you in some of the other… the other ones. Attendee, yeah, Common bottlenecks… I mean, instrumentation’s a big bottleneck, for sure. Instrumentation and tracking is a big bottleneck. I think agreeing on, like, what customer value means is Bigbonic. There’s a lot of… The custom annotation tool is a big bottleneck, too, actually. Yeah, annotation tooling and instrumentation, I think, are, like, the two biggest ones. Just getting started, also, is, like, a big, hard thing to do. It’s, like, the first step’s kind of hard. But that’s a good question. I’ll think about that a little more thoughtfully. If you want to DM me on LinkedIn or something.

Or email me, just to remind me, and I can, like, think about it a bit more. Nice, and then there’s… oh, nice. Yeah, that… that link that, Attendee put in there around redesigning product mixtures for AI evals. It’s kind of like, I’d say it’s kind of like the higher level version of this. We went into a little more details here. Where do AI PMs hang out? Where do they hang out? Oh yeah, this is good. Can’t find any communities. This workshop’s awesome, by the way. Okay. Where did they… how did any IPMs hang out? Man… Yeah, somewhere at Lenny’s. What? I’m trying to think of, like, communities to talk to people.

I mean, hi, I feel like you’ve met a lot of people going to kind of different, conferences and dinners and stuff in San Francisco that we brought onto the podcast. I live in Tahoe, I don’t hang out with anyone, man, I hang out with the bears and the coyotes.

Hai Guan: They’re everywhere here in San Francisco. If, if anyone want to drop by, I can intro you guys. But just, you know, certainly it’s geo-specific, I would say. There’s a lot of in-person events if you’re in the proximity in the Silicon Valley. I would say outside of that, it’s probably a hit or miss.

Shane Butler: Slack groups, forums, you know, Substack, there’s a lot of AIPMs writing… on Substack. If you basically go to Lenny’s, and go to… go watch Lenny’s FataCast, see their guests, or go to Maven and look up the people on Maven who are doing, like, AI product management work, and then you just, like, look up those people’s names on Substack, they all have, like, pretty awesome newsletters. And I think there’s, like, some good community in the comments there. And a lot of them are, like, super awesome response, like… like, we’ve had a bunch on our podcast before, like, super responsive, like, yeah, we’ll talk to you. I guess that’d be my advice. I don’t have, like, a link to a Slack group.

But we should… I mean, we should create one for this, maybe. We’ll… maybe we’ll create, like, a Slack group, or, like, Discord or something, and I can send that out. Alright, I’ll do that. next week. Yeah, I’ll do that. I’ll make a Slack group, I’ll email everyone who came to this with it. That’ll be cool.

Hai Guan: I think it would be also really cool to, perhaps in one of the future, either lessons or incorporate in the course, to bring in some AIPMs.

Shane Butler: Yeah. That’s a good idea. Okay. We did… we’ve had some pretty good ones on our podcast that we could probably bring in. And some other folks we’ve just talked to that have been kind of helping us with, With this, yeah. I’m trying to think of some people, like… Iman Khan’s really good to follow, Colin Matthews… Those are two we had in the pod that are, like, super responsive, and they have really good Maven courses. Yeah, we’ll try to bring someone in for either a lightning lesson, or indefinitely in the full cohort, but… Aman Khan is, like, head of product at Arise. It’s, like, that AI Evals platform, yeah. That dude’s pretty cool. He did a Lenny’s episode, too. It’s really good.

Definitely worth a listen. Alright, hang out a bit more. Got 7 minutes. Nice, already following perfectly. Perfect. What else? Other questions? What are you guys up to this weekend? I’m gonna try that pizza thing I told you about. Been, like… Gonna try that peel thing I eventually got. Try and make some pizzas in the pizza oven for some friends. What else should we talk about? Ski class? Oh, hell yeah, where’d you say you’re from again, Attendee? Austria. Man, I’ve never skied in Europe. We have… our snow’s so rough right now, it hasn’t snowed for 3 weeks in time, been getting out, but… Oh, you’re just watching, okay. Okay, Attendee has some AI evals thing.

Yeah, like I said, I could talk about skiing and snowboarding for all day, but for Attendee’s question around structured output in your work or chat, or combined. Yeah, so in my work, I work more… It’s like, my day job, it’s more like… It’s more structured output, I would say. We do have, yeah, thanks, hi. I’ll see you later. We have chat product, too. So, I would say the structured output really lends itself well to, like, that completion rate thing, where you have, like, a scale. of 0 to 1, how close did I get to what the user actually wanted, or to acceptance rates. And then, I would say the chat stuff is a lot more of that error analysis.

You really do have to do a lot of that with the chat kind of stuff. Anything multi-turn. Same thing with, like, agentic workflows, too, because you have to really look at, like, all the different steps that are going on, and kind of, like, check out errors between all those. The structured output stuff is easier, because you can kind of go, like. Alright, what was the input? What’s the output? What’s the ground truth? And… you’ll want to have those other metrics also, but they’re more, like, diagnostic to when you see, like, oh, my, like, my completion rate is quite low, can I figure out why it is?

Yeah, something like chat or, like, anything multi-turn or multi-step, you’re gonna want more of those kind of error analysis ones. We will go over, like, yeah, each of those, though, in the cohort, for sure. Great question. Oh, I haven’t tried this. BoundaryML.com. I’m just pulling it up on my other screen. First language for building agents. Okay, let me check this out. Thanks for sharing this. I haven’t tried it. Okay. I think I’m gonna call it, guys, because… My dog is barking like crazy out there, I don’t know if you can hear her, and I have another meeting in 4 minutes, so I gotta figure out why she’s barking. But this was pretty fun, and you guys had some great questions.

Yeah, I’ll send out the recording. Check out AIEL.ai for the course, check out DataNeighbor.com for all the other free workshops. Hit me up with an email if you’ve got any questions, or hit me up on LinkedIn. And, yeah. This was fun. Hope y’all got something out of it. Talk to you later. Thanks, everyone. Oh. Email. Let me drop it in the chat right now. It’s Shane… Shane, spelled Shane, at AI eval.ai. Here you go. Cool. I’ll stay on for 10 seconds, in case anyone needs to copy the email. 10, 9, 8, 7, 6, 5, 4. 3, 2, 1. Cool, alright. Thanks, everyone. Appreciate it. Have a great weekend. Hope to talk to you soon. Bye.

Free, every week

The next one is this Wednesday.

10 AM Pacific, live on Maven. One topic a week. Bring a question from your own work.

WED SEP 30
Ace Analytics Interviews with AI
Register
WED OCT 7
Metrics 101: Define a North Star with AI
Register
WED OCT 14
Build a Semantic Layer So AI Defines Your Metrics
Register
WED OCT 21
Experimentation 101: Run an A/B Test with AI
Register
WED OCT 28
Trust Your AI Analytics: Know When the Number Is Right
Register

Next cohorts start Oct 19 and Nov 2.

AI Analytics for Everyone
$1,800 · Oct 19 · ★ 4.9/5
Enroll on Maven
Agentic Analytics: Build an AI Analyst
$2,500 · Nov 2 · ★ 4.9/5
Enroll on Maven
Or come to a free workshop this Wednesday. Register free