← All free workshops
Free workshop · Wednesday, August 12, 2026

Open Models 101: What Are They and How to Use Them

Live on Maven, Wednesdays at 10 AM Pacific. About 55 minutes.

Transcript

Auto-transcribed from the live session and lightly cleaned. Attendee names are removed; their questions are kept.

Hai Guan: Let me… My screen… Whoa, everything kind of changes now. Okay, cool. Can you guys see my desktop here with Open Models 101?

Shane Butler: Yes.

Hai Guan: rates. Okay. Awesome. Alright, so why don’t we get started? Hang on one second, let me move my styles and stuff around. Cool, okay, great. Good morning again. Welcome, everyone. Today, the topic is around open models. So, AI, open models, you might have heard of them quite a bit, especially lately, on the news, on Discourse, on LinkedIn, things like that. We are actually going to talk about what they are. How to use them, the general landscape, and then we’ll get into a demo on using it in in a real analysis, for example. And, the… This lesson is… in collaboration with, actually, with Moonshot AI, which makes Kimi.

You might have heard of them in the last few weeks when they came out with Kimi K3, the model that has the most parameters ever released for an open source open weight, frontier model. So we’ll be, we’ll be taking a look at that in our demo, and, yeah, excited to, excited to sort of, like, bring you guys the, along on the, on the concept of, of open models. We’re gonna keep it pretty high level, so that we can keep it moving, since we only have an hour today. So, what we’re gonna be doing is we’re gonna have roughly 20 minutes of content, so slides.

We’ll get into a demo, probably around 15-20 minutes or so, and then at the end, leave The rest of the time for any questions you might have around, around, open models and, and, related, related topics. All right, cool. Oh, and, you know, like, we have other content available. We have upcoming lessons and things like that. You can check us out on AIanalystlab.ai, or on Maven directly. We’ll share We’ll recap all of these in an email at the end, and we’ll share recording, certainly, with you all. And so, yeah. Cool, okay, great. So, before we get into it, maybe a quick question for the folks here. So, another poll. Have you ever run an open model?

So, if you want to drop a 1 on… in the chat, if you’ve never touched one. Drop a tube if you’ve played with one, let’s say through something like an llama, and it’s totally cool if you’ve not even heard of… if you have not heard of this name at all. Or number three, you use it for real work before. Okay, cool.

Shane Butler: Lots of ones, we’ve got some twos, a few threes. We’ve got what’s an open model.

Hai Guan: Nice.

Shane Butler: You’re in the right place if you’re asking that question.

Hai Guan: Amazing.

Shane Butler: Attendee is the TA. He’s already answering it for us.

Hai Guan: Oh, cool. Attendee, do you wanna… Control on… over… over this deck.

Attendee: I’m not answering any questions either over here.

Hai Guan: We’re gonna put Attendee on the spot to be the Q&A person. Cool, yeah, lots of 1s, some twos, and actually some threes as well, so that’s, that’s awesome. So yeah, in the right place, we’ll get into some of the… some of the advantages, benefits, and what the whole discourse is all about. And hopefully by the end of this, you have a good sense of where this is headed and where they are now. Lots have changed over the last, I would say, 3 months or so, and, you know, certainly wanted to share as much as possible with you all. Okay, cool. So, let’s see. So, what is an open model, right? Like, I think, that’s a… that’s a very valid question.

Saw that on the chat, and thank you, Attendee for answering that. So imagine, just, you know, think of it as, everyone here has used ChatGPT, you’ve used Claude, probably, on the, on the chatbot, like, on the website, through their chatbots. So the idea for, you know, like, what a closed model is, is it’s similar to those, where you… type in a question, and then you get an answer back. But then all the computation for getting you the answer happens over at Anthropic or OpenAI servers. Like, it’s happening on their computers. You have no clue what they do. You have no idea what their model does for getting you the answers that you’re getting. So that’s kind of, like, the way to think about it.

And, you know, like, you’ll see a token count of some sort where, you know, like, hey, this is a complicated question, it takes some amount of computation, and you’re getting charged for that. So that’s roughly the idea, where it’s, you know, it’s closed by definition, so you don’t see it. Open, on the other hand, is, the weights are transparent. So, think of the weights of a model as just the numbers that really make the model the model. So, you know, like, you can You can download the entire model, you can run it on your own machines, you can even run some of the smaller models on your laptop, for example.

You can change things, you can ask it, it will compute and give you back an answer, and you know exactly why it gave you the answer, for example. For the, For… based on your questions, for example. So, the analogy here is almost like renting versus owning. So, renting is like, you know, you’re renting models from the, Frontier Closed Labs. They never make it, open to, for, you know, like, it’s their secret sauce. That’s how they’re able to kind of, like, monetize on it, have the char… have the pricing power and all that kind of stuff, open is, is the opposite, which is, you know, you have the ability to own it. They just make it available to everybody.

So hopefully this is, this is intuitive, this is clear. And so… if you kind of, like, get into one step into this, like, what does owning actually get you, right? There are a few things that are really important for why this is, you know, like Jensen Bong, all the tech leaders have signed some sort of, like, open letters to to the government to be like, you know, hey, we need open models. And there are a few things. Thing number one is privacy. So, your data, technically, runs on your own machines. Like, if you actually, let’s say, like, download the model and then run it on your own hardware. So, your data actually never gets sent anywhere, right?

Because all the computation is happening on your own machine. And so, you know, like, I think there’s, you know, if you’re renting it from a vendor, let’s say Anthropic or OpenAI, you are at their mercy and their promise to be like. oh, like, you know, we’ll never train on your data, or whatever it is. You know, that’s what they say. Whether that’s true or not is, like, you know, you take the risk. But if you’re running it on your own machine, then even if you cut off your internet, it will still run, as long as you have electricity. So that’s kind of, like, the idea. The second piece here is, the ability to fine-tune the model.

So you can actually change the model with your own data, because now you know the weights of it. And so, you know, like, maybe you have proprietary data around let’s say, like, healthcare or, like, some… some sort of, verticals that are, that are extremely, you know, beneficial for the model to behave, or to have that, to have that, knowledge, intimately, then you can actually do that, with open weights. You can do that in closed models as well, but, like, again, like, you don’t own any part of that. And then the third is, just simply nobody can take it away, because now you own it, by definition. If you, remember, maybe, like, a couple months ago, or maybe, like.

yeah, like, 2-3 months ago, forgotten by now. Fable 5, from Anthropic released for a couple days, and then, they had to be pulled back because the U.S. government says, hey, this is too risky, like, we can’t… we can’t, We can’t allow this until… until we have further review on this. So, anything that you sort of, like, depend on others, is basically subject to Availability, terms of service, and things like that, where it can be pulled away, like, you know, with, in a moment’s notice. So, you know, that’s another sort of, like, thing that, that open models, give you, where you have a lot more control over it, and it can’t just be yanked away, because now it’s yours.

So really, just the main idea is it’s… it’s yours. You know, like, I think that’s a little bit nuanced, but high level, that’s kind of, like, what, what, how you can think about it. Cool. Okay, so, So, who makes them now, right? Like, we’ve probably, or many people have probably heard of, oh, open models, you know, like, challenging the closed frontier models of them, of the U.S, and, like, you know, like, a bunch of, like, stock market movements because of those, Because of those chatters and confidence and capabilities and stuff like that. So, as of August 2026, I’m just listing some model providers out there.

There’s a way… there’s way more than this, out there that, That have really capable models, but, you know, obviously we just don’t have the, don’t have the space to… to put them all out. So the, you know, like, a few, a few models or model providers that, that, that might be interesting to… To… to learn the names of, if you’re not familiar. Kimik3, obviously, it’s from Moonshot AI, coming from a Chinese lab. You’ve heard… probably heard of DeepSeek V4, or DeepSeek, and then they recently, in July, came out with DeepSeek V4 Pro. There is Alibaba, which comes… which has a model family called Quen. There is a company called Z.ai, and they make something like GL… something called GLM5.2.

These are all frontier models, by the way. Like, they’re extremely… Like, they’re really, really big models, and very capable, and we’ll see it in the next slide. You know, like, from… or I guess earlier this week, Meta released Muse Glimmer, like, literally 2 days ago. I had to revise this… this slide last minute because of that. So, there’s a bunch of different models out there. Many people are aware of the, the Cloud family from Anthropic, ChatGPT, or the GPT family from OpenAI, or maybe Gemini from Google. But there’s, like, hundreds of different, different model variants out there. Each has their own specialties, each has their own strengths and weaknesses, and stuff like that.

And then it only keeps improving. And a lot of these, open models come from, come from China, and especially the frontier ones that we’re seeing kind of, like, rapid releases. They are, certainly kind of, like, predominantly coming from, from Chinese companies. So, the whole idea is to not kind of, like, you know, like, you don’t need to memorize any of the lists or anything like that, but just know that, you know, these might ring a bell as, as the competition heats up over, over time. Cool, okay, so, couple things to note, why do we, why do we even care? Like, you know, like, we have submodels, we have, you know, OpenAI Anthropic, why… why… why do we even need to care about, any of these?

A couple of things. One is, capability. So, if you’ve… Again, if you’ve heard of, you know, the Kimi K3 launch, for example, a lot of chatters out there that says, hey, this is a really, really good model, rivals the top of the of the U.S. closed source models, and stuff like that. So this comes straight from, from their publication around their model capability.

And not gonna go into, you know, too much in detail around different… kind of like these, different benchmarks, but across just different, different things that they measure, how capable their models are, you can see, like, other than Fable 5, or maybe, like, GPT 5.6, SOL, which are the two, most advanced closed source models, they’re… they’re right there with, with everyone. And, and then, kind of like, if you look into it, it’s really not that far off, based on some of these benchmarks, from, models from kind of, like, open source families. And so, you know, this… this is actually not true, 3 months ago, back in May, I think, if I remember correctly.

Where, a lot of the, kind of, like, the, the, the open source models are still pretty lagging behind the capabilities of, of the, of the closed, of the closed ones. Now, benchmarks are one thing, like, you know, I… my advice is to take it with a grain of salt. you have to apply it to your own workflows to see, kind of, like, what the behavior feels like, what does it fail, where does it break, and where does it perform well, kind of thing. So, you know, like, I… Would never hold this to… be authoritative or anything. It all depends on what you’re trying to do with it.

And so, we’ll get into, kind of, like, what that looks like from an analytical perspective, that we’ll get a demo for in just a second. Cool. And then the other piece that is important to note is the price. Again, lots of chatters out in the industry around, hey, closed source frontier models are really expensive, and open source models are just way, way, way cheaper. And so, you know, like, for a given unit of intelligence. you know, like, how do we even think about this now? Now that it’s, you know, like, the disparity is so, so, so huge.

And so, just wanted to put these models and their pricing on one single chart, so that we can see, kind of like, you know, what that whole thing, what that whole conversation is about, right? You know, if you… kind of, like, the most capable model here is obviously Claude Fable 5, and it’s also the most expensive by a mile. $50 outputs per million tokens, that is substantially, substantially higher than any of the bolded ones here, which is all open. So these are the frontier open models out there, Kimi K3, GLM 5.2, DeepSeav4 Pro, and they are I mean, like, you know, orders of magnitude, cheaper than some of these, some of these, closed ones.

Now, you know, like, does it mean, like, oh, like, we will just go, you know, like, use open models everywhere? No. Again, it all depends on the use cases and where… where it makes sense. Like, if… Perhaps, like, you really do need the frontier intelligence to be able to solve some problems that you’re… that you are, That you are looking for solutions on, or you may need kind of like, you know, like, very good models at certain aspects to coordinate and orchestrate your workflows. So it all really depends. But generally, kind of like the… the biggest thing that people have been talking about, comparing closed versus open, is on price.

And, you know, obviously on capability, but capability is, you know, like, very dependent on workflows and use cases. But generally, orders of magnitude difference in terms of, just pricing for per unit of intelligence. Cool. This stat is really interesting, so that’s why I put it here, which is there’s a… there’s a publication from Open Router, back in 2025, where, even back then, when open models aren’t, you know, like, nowhere near today’s capability, is… sort of like, you know, Open Router is a, kind of like a service where, like, you’re, you’re able to kind of, like, swap models around, and so a lot of workflows goes through their service, because it’s very flexible.

Like, you’re not locked into any one vendor, for example. And so what they found was, Roughly 20% of their tokens, are routed through open models. This is back in 2025. And, you know, only 4% of the revenue, is, is, is accounted for, through there, which is just highlighting, kind of, like, the, the dynamic of demand and, and just kind of, like, price, between the, between the closed and open, models. And by mid-2026, I think, just a couple months ago, they’ve, they’ve said that, they crossed half of routed tokens going to open models.

So as models get more and more capable, it’s eating more and more of the share of, sort of, like, the, the workflows out in the Out in the, enterprise, and companies. And, again, like, you know, like, everything is very dependent on your use case, so we sort of ran our own, pretty, you know, simple test as well, where back in June 2026, when a flood of, kind of, like, new frontier open models came out. We actually ran some, we actually ran some open models through a very… through the exact same analytics question.

And then sort of, like, hit the answer from the, from the agents, and then we, we graded them using, I think Fable 5 at the time, on, you know, did they catch all the, all the gotchas? Did they explore all the different angles? Did they come up with the answers that were exhaustive, for example, and that they sidestepped the landmines that we planted in those datasets? And, you know, like, this is kind of like a proxy, again, like, nothing is, deterministic, nothing is completely exhaustive, so take it with a grain of salt. Directionally, what we were looking for is, like, are the models capable enough to do pretty open-ended analysis that has a lot of, just different gotchas and stuff?

with… with, you know, with the right answers, and safeguards around, you know, hey, no peaking. Just kind of, like, explore the data on your own based on the same exact question, and see how they solve it. And generally, the takeaway is that they were pretty capable. Especially at the time, GL15.2 was the latest, model, and the most capable, and certainly. It worked pretty well. And so, it’s been kind of, like, 2 months, right, since, since the test back in June. And so, you know, like, all the models have only gotten smarter and more intelligent. Okay, cool. So I think we’re pretty close to demo time. So generally, how do you actually run one, then? Like.

It’s great, all these, all these things I’m sewed, how do I actually run one? A few paths. One is run… you can run small ones locally. So, tools like Olama can get you a small model running in probably a couple minutes, on your own desktop or your own, you know, or, sorry, your own laptop or your own desktop. Like, literally, your own computer is doing the computation and turning the model through the questions that you ask it. So that’s one… that’s one way. You know, not super practical, I would say, because the small… the smaller the model.

the less capable they are, and generally, let’s say for a Mac Mini, 16GB, the standard Mac Minis that you can buy from Apple, like $600 or something, you can run something like an 8 billion to… 16 billion parameter, model. And for contrast, the Kimi K3 open model is 2.8 trillion parameters, and so that’s kind of, like, how you can, think about… think about that. So you really need kind of, like, state-of-the-art hardware for that to work, but maybe for some things that are a lot less, you know, ambiguous, or a lot more deterministic. For example, like, oh, maybe, like, you know, just reading some stuff and summarizing it, or, you know. Copyed something that might work already.

Second one is hosted API, so you can… you can rent it, just the same as you do with, for example, Claude. If we take Kimi as an example, they have their own website, they have their own chatbots, they have their own API keys once you sign up for it, and you can use it as you go, kind of thing. You know, it’s still open, it’s just that they host the model for you, kind of thing. And then there are other providers where they aggregate different, different models and stuff, and you only pay for, let’s say, one subscription or one bill.

And then the third path here is self-host in your servers, so you can spin up your virtual private cloud, or, you know, like, some, some private servers somewhere where those machines are still yours, and you can download the model, put it on there, and then run it and have it run. Over there, and sort of, like, you talk to… you talk to it via, through… through that, through that channel. So all of these are… all of these are just ways that you can actually run one. And, in just a little bit, we’ll actually do path number 2, where we’ll, we’ll hit their APIs.

Cool, yeah, so just one misconception here to, to clear is that, not all, or I guess, the frontier closed models are definitely not runnable in your laptops. So, you know, like, while they’re very capable It is extremely unlikely, unless you have… unless your computer at home is, like, you know. so powerful, that, you can actually run it yourself. So, most of the time, these are… these are gonna be some hosted servers somewhere that you pay for the token at that pricing schedule that you saw earlier. Cool. Okay, so we are gonna get into a bit of a demo here, where we’re gonna, look at, Kind of like, first, swapping. the brain of, Claude Code, actually.

We’re gonna use Claude Code but with a open model, which is KimmyK3, and then we’ll use the same kind of, like, setup in Cloud Code, but just a different brain to help drive the whole operations. Cool, let’s see… Okay, For folks who have been following us for a while, this should be pretty familiar with you. If not, that’s okay, too. Hopefully, you’ve… you are… you may or may… you may have seen, kind of like, clock code, or you may have used it if you have not followed us. One way to use plot code is, in Terminal, and the… the workflow that we like to use, and that we teach others, in our courses, is to use it via, an IDE, like.

VS Code, which is what I have here, or kind of, like, anti-gravity, or cursor, or some other editors. So this one is, so this is VS Code, and, I have a terminal down here at the bottom. I am in a folder, that is our AI Analyst Plus repo. So maybe, Shane, you can share that link. There is an open source AI analyst system that Shane has built, that is open source, anybody can use it. anybody can play around with it, and what it is, is a, sort of like an Agentic system that mimics the behavior of a, senior product data scientist. So… Everything, all the best practices for how to frame questions. How do you analyze data? How do you do your visualizations? How do you tell a story?

How do you construct all that into a deck or an artifact, such that you can convince stakeholders on the findings and recommendations? They’re all encoded in, kind of, like, this, this, AI Analyst Plus repo. So we’re not going to go into too much detail. You’ve probably seen us, shown this quite a bit, before. So what we’re going to do is we’re going to launch Cloud Code using this setup, because, you know, it’s already built, but with Kimmy as the brain. So, usually, and if you can’t see it, hopefully you can see it, if you can’t, you can, Command-Plus, if you’re on a Mac, in your Zoom, and you can actually zoom into… zoom into your, Zoom in in your, in your, in your Zoom window.

So usually, if I go to… if I launch Claude, right, so I just type Claude on my terminal. And I’m in this, AI Analyst Plus folder. I mean, it’s renamed to Open Model 101 Demo, but, you know, it is… It is, it is AI Analyst Plus. on… generally, when I type Clod, what I get is, you know, I would launch Claude code, but then the brain is Opus 5, and the model families would be, Fable, Sonnet, Opus 5, so on and so forth, so everything from Anthropic. But this is actually very easily swappable. So Cloud Code is just a application, and the brain, think of it as an engine that could be decoupled.

So I, in this, in this demo, I actually want to swap this out, not Opus 5, I want to use Kimi, so I’ll exit out of this. And, I have a… let’s see… I made a few scripts here, and I’ll share some setup instructions for you all if you’re interested in For example, like, try out, let’s say, Kimmy in your own workflow, or any… it works for any, any, any other open models. It’s a very similar, idea. So, here, what I’m gonna do is… I have some scripts here, where it will help me sort of, like, swap out, swap out the model, and I’ll show you what that looks like, just generally.

So, I have this, Kimi environment, file here, where it basically… all it does is it swaps the engine to Kimi K3, and it’s basically pointing Anthropic’s, quad code kind of, like, path to the Kimi, API that I already have. So it is, you know, like, it’s pretty painless. None of us needs to know any of these. Quad helped write… write all of these. But the idea is, I just wanted to show you how easy it is. So, I’m gonna reset my demo, actually. Okay. So I reset up my demo, and I can just do, just do this. source, demo, Kimi, environment, And then, basically, what I did was, I just, ran this, ran this script, and then it says, okay, the engine is now KimiK3, And now, when I do plot.

I am back in plot code. But notice that, I’m now running on Kimi, Kimi K3, instead of Instead of, Opus 5, or whatever it is. And so, I’m gonna run a real analysis on here. Right now. Hang on one sec. So, I have this prompt. That we can read together. let’s see… which is, using the workflow in this repo. Again, this repo is help… is, an Agentic system to help answer business questions with data, similar to how a product data scientist would do it. How has the checkout conversion rates, trended over time? So we have a dataset here that is a fictional e-commerce company called Nova Mart.

Think of Amazon, so it has all the components that you’re familiar with if you shop online, which is, you know, you have a checkout screen, you can start payments, you can complete payments, and all that kind of stuff. So the idea here is, we want an analysis on how is checkout conversion rate trended over time, and what opportunity do we have to optimize it? output the analysis in an HTML, report, and then pop it open for me, because I don’t want to find it, for a VP of product. So that’s my audience. And I also have a… script here, where it will help log how much input and output token this analysis, used for, you know, like, end-to-end.

So this is the receipt.sh file here, based on a model that you’re using, which is Kimi, and what that translates to in terms of dollar amount. So what I’m trying to do here is to understand When it finishes running this analysis. How much this whole thing costs, and then compare it against, let’s say, like, other models and stuff. Alright, so… I am gonna fire it off… Alright, so… It is… So this now sends, kind of like, this question to, to… Kimmy? And we’re not gonna sit through the whole analysis. The analysis is gonna run somewhere between 15 to 20 minutes.

We’re gonna look at what It does, just high level, and then I’ll show you the output from another run, so that we have an idea, and then, and then we’ll finish up and let it finish running, and we’ll kind of, like, close it out with questions and stuff. But I think the idea here is to, get an understanding of, sort of like, hey, if we run, Kimi in this, in this repo. How… what is this… what is its thought process, and how is it tackling this problem? And what its plans are, for example, and what is it doing behind the scene? So we’ll take a look at it for just, you know, a couple minutes, as it turns through.

So you can, you can always follow, kind of, like, its, its chain of thought reasoning. I’ll run this through the repos analysis workflow. So, if first… is trying to get oriented with what the dataset is, what this script that I asked it to log Tokens is about, and then all the different things that it has access to. Also loading the dataset context and, stuff like that. And then it’s connecting to the, the data itself, and, and then figuring out, kind of like, hey, has this been run before? What other available information can I get? Off so that it knows what to do. And so, this is sort of like the multi-step, the multi-step analysis, kind of, like, steps that it’s, that it’s come up with.

And this is, again, gonna take probably, like, 20 minutes or so. But these are all different. skills and instructions that we have in the AI Analyst repo, so… You know, high level, what it’s trying to do is go frame the question, and then also define this convergent metric. like, how do you measure that? It’ll run the trend, the funnel, segmentation analysis. It will then validate its findings. In our skills and agents, we have a few layer of cross-validations and checks on the final outputs before it’s being compiled into kind of like a deck or a doc or something. It will size the opportunity based on the insights, so that’s also something, based on this repo.

It will then build HTML reports and, and open it, and so on and so forth. So, oops. So these are all kind of, like, the workflow steps that we… that we have, and so one way to really understand as you use new models, in your own workflow is to really understand, kind of, like, the thought process here. Like, does it call the right tools? Is it able to figure out The, like, how to navigate, what is its reasoning for doing certain things, and, and so on and so forth. back in May, when the models aren’t, aren’t good enough. What they’ve been doing was, they constantly run into walls, like, oh, I encountered this error message, then they try to fix it themselves, and then it goes nowhere.

Or they’re like, hey, let me reframe this question, but then it has no idea what a good, you know, good, good, good, good reframed question or hypothesis is. And then it’ll run in circles, so on and so forth. as models get more smart, as things get more intelligent, those things are kind of like a thing of the past, if you will, and at least supposedly. So, in your workflow, the idea really is to make sure that that’s actually happening for you. Cool. So, it’s gonna… it’s gonna take a while to run. Again, it is going to take a few… In, like, 20 minutes or so, so we’re not gonna wait for it, we’re just gonna let it do its thing, and and we’ll take a look at what it looks like.

Okay, so this is the previous run that, that I’ve run. It’s the same exact question. And and then I did the, sort of, like, the entire full analysis on that data set. The run that I’m currently kicking off, or in between any of these runs. There is a reset step where we wiped out all the… all the… all the things that it’s done in the previous session, so that it can’t just peek at the answers, for example, or what has been looked at, or, you know, like, what… what it does. like, what the data set… what the data says, kind of thing, without, you know, actually doing the analysis itself. So each one of these are, independent.

And so the… this… this guy is the report, supposedly for the VP that Kimi, ran, just, just earlier this morning. So it gave kind of, like, the, the TLDR, like a summary of, of what’s, what’s happening based on the question. Remember, the question was, how has checkout conversion rate, trended over time, and, what, how can we optimize it? And so, first. question in our first title sentence here literally answers that.

And the thing that I’m looking for, for example, because I’m, like, we’re intimately familiar with this dataset, because we work with it for so much, and we taught it for so long, is that there is a component of, kind of like a Simpsons paradox baked in to this, to this data, which is, hey, high level, the checkout completion rate fell from 68% to 30%, while traffic grew to 27x. default is a mixed shifts, so not a broken checkout. So this is the Simpsons paradox that we’re expected We’re expecting the model to be able to pick up. If it’s, for example, like, smart enough, or, you know, like, it follows instructions or, or something.

This is a… the real fixable gap is mobile, so it figured out that, the checkout The checkout page to payment start has a dramatically different, different, conversion percentage, between those steps. And then also other levers that it found based on, kind of, like, the, the different experiments that are available, kind of like the old experiment, that, that we have, baked into this, fictional, dataset. And so, going down, it’s basically just, it’s, you know, highlighting all the supporting evidence, all the different, different visualizations to, to help illustrate a story. So, you know, what happened, completion rate went down.

But this is mostly a makeshift issue, again, the Simpsons Paradox check, so… It is not just a, it is not just a straight-up decline that you have to be, like, super alarmed, but then, you know, like, the new users, they convert much better. Returning users, not so much, but then the share of them have gotten a lot more. And the gap is mobile checkout. So, by the way, these are the correct answers, since… since we’ve used it for, for… for so long. And then it highlights, kind of like, the steps between, the steps in the conversion rate funnel. So, from checkout to… from add to cart to checkout start, from checkout start to payments attempted, from payment attempted to purchase.

between, I think this is between web and, and mobile, and so it’s highlighting, hey, this, this guy, there’s a bunch of, opportunity here. Everywhere else seems okay. And then… It looked at a previous experiment. Again, we would expect a good model to be able to pick that up, because that’s part of a part of the… part of the setup in… in the… in the repo. And then it does the opportunity sizing. If we fix certain things under certain assumptions, what does it net us in terms of, in terms of opportunity, and whether, we should pursue it, kind of thing. so on and so forth, and it’s got a recommendation here, and all the receipts on validation methods and caveats, is under here.

So that’s kind of, like, the output that, that it’s able to give us, and, again, because we know this data and this story pretty well, that, This is pretty satisfactory in terms of, what it came up with.

Shane Butler: I mean, it’s kind of crazy, I don’t know if you want to scroll through some of those, visualizations, but it’s crazy that, like, we’re now at a point where like… Agentic analytics systems that we build ourselves, that, like, you know, we built this, but anyone in this room could build this yourself. This isn’t some, like, crazy enterprise tool you had to buy, is able to deliver this level of rigor in analysis from end to end, but then it’s also even crazier that this didn’t come from, like, open AI or anthropic models.

Like, this came from an open waste model at a fraction of the price, and… I don’t know, it is, for me, I’m like, I don’t… I don’t understand, like, how these, like, Frontier Labs are gonna compete, because if it’s already here, it’s just gonna get better and better and better. there’s some threshold of model capability where, like, you don’t need a better model. Like, for us at the lab, like, we found kind of, like, the analytics capability kind of hit that threshold in February with Opus 4.6, and so any models that are as good as Opus 4.6 can do most of the analytics workflow.

The Frontier Labs are going to continue to push out newer, bigger, better, better quote-unquote models based on some benchmarks, but that doesn’t mean it’s better at getting your job done, doing your job faster, doing whatever task you need to complete. So, there’s just, there’s just cheaper, more secure, more controllable, opportunities for how you can start leveraging AI to get where you actually need to be get done. You don’t need to be on the frickin’ most expensive, newest thing that these companies are gonna automatically default to you and their products, or try to sell you.

A lot of times, when I do analytics, even before I was getting into open models, I was just always trying to push back to Opus 4.6, because I don’t want the newer Opus models, because they take too long, and they’re more expensive, and I can get the same answer with the earlier ones. Also, Opus 5 talks so stupid, I hate… it just actually infuriates me how it talks. I’m like… and I find myself cursing at it, and I’m like, well, I shouldn’t do that. Like, you know, I gotta be nice. Like, like I was in the Opus 4.6 days.

Hai Guan: It is very AI, that’s all I can tell. Okay, cool. Let’s, wrap up this, this demo here. So the last thing I wanted to show you guys is how much it costs for these runs that, that you see. So, I ran the same thing on Opus 5, for example. And it was something like $16 for that run, if it were to be billed at the API, pricing list. Obviously. we have subscriptions here, so it’s, like, we’re not paying on a per run or per analysis basis. So, roughly, like, $16. It came up with the exact same, or very similar, kind of, like, Outputs, by the way. And then KimmyK3, which was the run that produced the other one that I… that you guys just saw, was roughly 5.60.

And I actually have a usage counter myself that I monitor. It is less than that. It was more like $2, $2.50, $3, something like that. And I have another scratch run here that was, like, $6.50. I think this was, like, also, like, $3. You know, so there’s some discrepancies from how it does the token counting, and kind of, like, the charges. take it with a grain of salt. The directionality is what’s, what’s, what’s, what’s interesting here. Okay, cool. So let’s get back to the slides here. Alright, gonna wrap up really quickly. Let’s see… cool.

So, if you are interested in, let’s say, like, an open model, to learn more about it, or, or anything, or specifically with Kimi, for example, since that’s the latest and greatest, you can go to kimi.ai, that is the… like, the chatbot equivalent, that you can see. It’s no different than from the things that you’re familiar with, like Claude or ChatGPT. Very similar interface. If you want to build on it, use the API or a subscription with them. You don’t have to subscribe with them, you can subscribe with other, sort of, like, hosted model, hosted providers as well, like Olama.

We don’t have time to demo it today, but, like, we’d love to show you guys at some point, for those who are not super familiar with, with that. And you can, you know, even download the weights. I think they just literally published it maybe a week ago, a week or two ago. And when is an open model the right call?

So if you have a lot of volume, so, for example, like, hey, I have to do a ton of analysis, I have to design a ton of Agentic systems, or whatever it is, volume, typically there is a… there’s a break-even point where it makes no sense to continue to pay for Sort of like, you know, certain, certain models, if you can move it over as when the capability is already there, to save on cost, kind of like, as your volume scales. We talked about control, and the third one, which we teach in one of our courses, is around using different models to check each other’s work, especially now that Frontier models are getting pretty smart. They can actually do this pretty reliably.

Like, you have more options, not just let’s say, like, one model, then you have to trust it, or, you know, figure out ways to have it be as trustworthy as possible. But you can actually have a symphony of different model families to check each other’s work, and that’s a powerful validation tool as well. All right, so this is the last slide. We have more free lessons coming up for those who are interested. Next week, we have a free lesson around building semantic layer, so AI can, so AI can define your metrics. None of these is possible if you have not even defined your metric yet.

AI is just gonna guess, so it doesn’t matter how capable the models are, if the very basics of Doing the foundational work of what does checkout conversion rate even mean is not defined, it’s not gonna save you. So, pretty important topic, and, excited to, to talk about it. Shane’s gonna be the instructor on that one. And then we have two… we have a couple paid courses. Folks who’ve been following us are pretty familiar with these. The first one is on Agentic Analytics, building the AI analyst system itself, so building the Agentic Analytic system that can… that tailors towards your workflows, your, your needs, and, and your data, stuff like that.

five-week course, that starts in September, where we go through quite a bit of, you know, how do you… how do you build one from scratch, to how do you bake in evaluations, context. We go deeper into open source models as well, in those, in those 5 weeks. So, if you scan it or… use the link or use the code here, you get 20% off for the upcoming cohorts. And then, in parallel, we have another course around, just the foundational knowledge of, analytics. So, how do you, ask really good questions, and then let AI execute the rest?

And the idea is to help everyone to be, you know, like, kind of, like, develop the human skills that are still needed in the age of AI, where you understand, sort of, like, what, what does good look like, in terms of In terms of doing really good analysis, in terms of what the, you know, the, the in-between components are, to make a question leading to decision, a insight leading to action, and everything else in between, but leave the rest of the execution pieces to AI, to help you get to… to help you actually get at what matters for a business.

Shane Butler: Yeah, maybe a little bit more on the Ag Analytics one that’s coming up in a few weeks. It’s really freaking good. It’s, so we’ve ran this 5 times now, we’ve had over 250 people go through this, it’s got, like, a 4.9. on Maven, there’s no other course like this right now that teaches you to go from literally never used Cloud Code or Codex or anything before, to by the end you’re using multiple models, Cloud Code.

Codex, open models, you’re going all the way from, like, building your own system from scratch, learning all the foundations of Agentic systems, to learning around how to build AI evals, which is just an extremely important skill for product development in general, but we go specifically into the use case of AI evals for, analytics cases, which is very different from other products, by the way.

To context engineering and management, where, like, we’re finding this is, like, this is what makes, the, like, open models and, like, more affordable models, like, extremely capable and reliable, is, like, being able to have this, like, really, systematic way of approaching how you manage the context of your business and your data in such a way that agents, can read it. And it’s very bespoke to every person’s individual company, so we go through that. A lot of our free stuff we do, like, every week, you know, we have these free lessons, and you know they’re usually more, like, of, like, us talking and doing a demo. These courses are extremely hands-on, so you’re building the entire time.

We actually have a rule in our courses when we develop them, that we won’t talk for more than 10 minutes in a row at you before we start building as a group. So… Highly recommend that. If you have any questions at all, email us. Or reach out to us on LinkedIn, or we’re happy to hop on a call with you as well. But that’s a really cool course, hope to see some of you there. I know we had a lot of questions in the chat, but I think I got to… Most of them But if anyone has questions on open models. For the last few minutes here.

Hai Guan: Throw… throw them… throw them out. Looks like the, the other run finished as well. This one is a little off. Yeah, I’m not gonna read it, but it seems like it did figure out the thing. Anyway. Oh, If nobody has questions, the one thing that I wanted to… Actually, we’re not gonna have time for it, it’s fine.

Shane Butler: Oh, man, you’re just gonna just tease… tease us all like that. What are you gonna do? Oh, we had a question here from Attendee. Want to build layers and tune an open model?

Hai Guan: Built layers and tuning. Do you guys also cover that? What do you mean by layers, Attendee? If you were talking about, like, fine-tuning an open model, we don’t actually do that, in… in our course. There’s… there’s a bunch of, different, different resources online as well. We… Try to just cover the, kind of, like, the practical things around how to use it for analytics workflow, so not so much on tuning the model itself. To be sure this session was recorded, yes, I… yes, we… I don’t know how to tell, but yes.

Shane Butler: Yeah, it’s recorded, it’ll be sent out, maven should automatically send it out on Friday, actually. It should automatically be emailed to you. And then I’ll try and get it up on YouTube tonight or something, though. If you look up AI Analyst Lab on YouTube, you can also find all of our previous, workshops. They’re on Maven, too, but it’s a little easier on YouTube.

Hai Guan: Hey, your Attendee what’s up?

Attendee: I was just gonna ask on that same question about, fine-tuning. Have you guys, in your usage, had to, like. had to ever go through and, like, fine-tune a model, because I know it’s… it can be expensive and time-consuming and stuff like that, so… just curious if you’ve had to, and, like, what was the situation around that?

Hai Guan: Yeah, no, we… we’ve not… we’ve not had to fine-tune anything. Fine-tuning, I think, would be really powerful if you have very… much, like, proprietary data, and you have a lot of it, for example, that would materially differ from the base capability and what, for example, building your harness would allow it to do. So, I think, you know, I don’t think we’ve had that thing encountered yet, because we’re more general purpose, in the lap here.

Shane Butler: Yeah, I think, like, I’ve used some fine-tuned models for, in the legal space when I worked there, and I did find, like. You know, things like… you know, if you read a contract from a lawyer, like, to me, it’s always like, I don’t know, this looks like a legal contract. I do find that, like, pulling out the, like, semantic, like, different discrepancies is done a lot better by, like, a, say, like, legal bird or something, like, a model that’s fine-tuned on a bunch of contracts. But… Those are… those are fine-tuned and other, like, public legal information. We didn’t… we’ve never, like, done it ourself.

Attendee: Okay, no, that’s helpful, those very specific domains that a general model wouldn’t have all the data to train on. So, that’s really helpful.

Hai Guan: Yup. Yeah, totally. I think health… healthcare is one… one aspect, and then not. And then some stuff, for example, for companies that really don’t want, let’s call it, like, Anthropic to know this, the trade secrets, I think those are… those are very well suited for fine-tuning. Cool, okay. I will recap all of these in an email, including… I saw… someone wanted the token counter thing, I’ll send that out as well. And, we’ll also send out the instructions for how you could perhaps swap your clock code or whatever, with open models. And yeah, and hope to see you again in the future, Attendee.

Shane Butler: Thanks, everyone!

Hai Guan: Thanks, Al.

Shane Butler: See ya.

Free, every week

The next one is this Wednesday.

10 AM Pacific, live on Maven. One topic a week. Bring a question from your own work.

WED SEP 30
Ace Analytics Interviews with AI
Register
WED OCT 7
Metrics 101: Define a North Star with AI
Register
WED OCT 14
Build a Semantic Layer So AI Defines Your Metrics
Register
WED OCT 21
Experimentation 101: Run an A/B Test with AI
Register
WED OCT 28
Trust Your AI Analytics: Know When the Number Is Right
Register

Next cohorts start Oct 19 and Nov 2.

AI Analytics for Everyone
$1,800 · Oct 19 · ★ 4.9/5
Enroll on Maven
Agentic Analytics: Build an AI Analyst
$2,500 · Nov 2 · ★ 4.9/5
Enroll on Maven
Or come to a free workshop this Wednesday. Register free