Sravya Madipalli: Awesome. I think people are coming in! Hello, everyone! Hello, hello!
Shane Butler: Hello!
Sravya Madipalli: This is awesome. Great to see you all today.
Hai Guan: Good morning, good afternoon, and good evening.
Sravya Madipalli: Yeah. We also had midnights at the… in the last, sessions. It would be great today, before we get started, to know where are you all coming from? Do you want to, like, simply message? Where, where are you all joining us today from?
Hai Guan: What is that?
Sravya Madipalli: Put in the chat. Attendee, I said, California.
Shane Butler: Vancouver…
Hai Guan: Love Vancouver. New York City.
Sravya Madipalli: Awesome! Argentina, nice! It’s pretty cool. Bye! Oh, it’s late in the bye right now. Wow.
Hai Guan: Denmark. Wow.
Sravya Madipalli: We do have wide variety.
Hai Guan: Buenos Aires, that’s cool.
Sravya Madipalli: Awesome, awesome. So while everyone’s entering, I wanted to give you all a short intro about myself. I’m Stravya Madipali, I’m leading today’s session. I am a Senior Manager of Data Science at Superhuman. It was previously called Grammarly. I have been in the space for 14 plus years, starting at companies like Microsoft, I was there for a long time, then moved to eBay and Nextdoor. That’s where Shane Hai, and I met, actually. And when we were leaving Nextdoor, we wanted to do something together. So here we are, one of our initiatives. We also have a podcast called Data Enabled Podcast.
And now, we have a few courses on Maven that we teach, and here we are sharing one of, our, like, you know, learnings in the space of cloud code and AI. Shane or hi, do you want to, like, a quick intro yourself?
Hai Guan: Yeah, hey everyone, my name is Hai, and I lead the data team at a company called ENTRE. It is a legal tech AI company. For private markets, and met these folks at our previous job at Nextdoor, and then prior to that, I spent a lot of time at big tech across LinkedIn, Meta, Pinterest, all the big consumer tech companies. So, yeah, excited to… Excited to… to be here, and thank you all for taking the time in your busy day today. Over to you, Shane.
Shane Butler: Yeah, hey everyone, I’m Shane, the principal data scientist at the same, company Hierworks at in Legal AI. Yeah, been in data science for about a decade. Spent the last, probably, year and a half, almost two years now, in… more of the AI evals and agentic analytics space, so… Excited to, see y’all here today and share a bit about analysis design.
Sravya Madipalli: Yeah, awesome. Okay, everyone, let’s just get started. Let me start sharing my screen. Can you all see my screen? Okay, so hi and Shane, please, look at the chat. I’m, I can’t look at the chat. If there’s anything, please raise your hand. Hi, Shane, and I can pause, okay? Okay, so, today we are covering something that costs teams more wasted time than a bad SQL or a broken dashboard. This is basically designing a wrong analysis. You probably know the feeling, right? You spend a week building charts and cohorts, you present probably to your product leader or a VP, and then all they say is, interesting, but that’s not what I’m looking for.
So today, I’m going to show you how a 7-line framework prevents that, and an AI system that makes the whole process automatic. Okay? Let’s get started. Before we get started, we start with a typical quiz. Can you basically give your answers in a quick poll? Hi, I’m Shane, we look at, Please let me know once it’s done what is the most popular choice. So, this is your question. If your VP says, conversion dropped last week, and you have one hour, almost like an SLA, that you reply in an hour, because it’s your VP, what would you do? Do you do A, B, C, or D?
One thing I want to let you all know, no judging here at all, I think all options are pretty common, like, we all actually go through all these options, but I want to make sure, what would you do first? Can you please start giving your answers? Are we… yeah, looks like we’re getting some messages.
Shane Butler: A lot of C’s got some A’s.
Hai Guan: Yeah.
Sravya Madipalli: Okay.
Shane Butler: That’d be…
Sravya Madipalli: Okay? That’s awesome. Okay, so… is the popular ones between ABC?
Shane Butler: I think C is the most popular, but we got a lot of A’s, also.
Sravya Madipalli: Awesome.
Hai Guan: 2 comes at A and D, or no, A then D.
Sravya Madipalli: You know what? So, jumping… so, here’s the thing, right? I think… we have some good set of people, if you’re starting with C, but there is a lot of people that I’ve worked with in my career that, we start with A and B. I myself have done that. I myself have started with A and B, right? But, you know, I understand that that’s the instinct, to just jump in and understand what it is, but we probably… need clarifying questions. To understand what does VP even mean when he means by, you know, conversion rate drop. Does he mean week over week, year over year? what does it… what conversion is he talking about? Is he talking about, like, sign up to an Activation conversion rate?
Is it talking about, like, people free user to upgrading conversion rate? So, knowing these basic questions, and also, hey, why are you even asking this? Do you have a vo-tech that you have going for to report on these things? Or is it just something that you’re curious about? All of these things would help you like… understand how do you want to go about answering that question. So, most of the time, especially when I started out, so much of time was spent in just assuming answers for these questions and going ahead and, you know, doing these analysis, but you would benefit A lot, and a lot of cost, like, you know, time is reduced if you ask those clarifying questions.
One thing I want to say before we go into the next slides is that this, design analysis, like, workshop today would also help you for interviews. If you’re, you know, applying for interviews, if you’ve done data interviews and data science stuff, you definitely see yourself actually having at least a couple of interviews. For example, in companies like Meta, and any tech-in-tech company in Bay Area, for sure, has a… how they have a question, like, literally, this dropped last week, what do you do? It’s literally this question, and they are looking for you to come up with a framework and a set of things.
So, byproduct, along with learning how to use Cloud Code to learn it in your work today, you would also probably learn how to do it the best way in an interview. Cool. So, let’s go to the next slide. I want to say something, what will you have at the end of probably… optimistic, by minute 32, you’ll basically see a structured analysis plan document, the kind, basically, that you’d share with your pod, the pod in the sense, the team that you’re part of, your PM, your tech lead, before you spend a week running queries.
Basic… I’ve seen a lot of junior DS, if they have a question, they jump right into analysis and start putting things in a doc on the analysis, but agreeing with your stakeholders on what you’re going to look at for the analysis is the first key step, and there are… there is a framework on how we want to do it, and this is what we’re going to cover today, okay? And you would basically see, the final output as… at the end of the session. Okay, before we do the big demo, I want to show you something quick. This has been a style that I’ve been doing in my last three workshops. I want to show you something super quick.
I’m in the cloud code, and we can jump in into… the theory behind it, and then… wait, where’s my clock code? Okay. Cool. So, let’s get started. This is, like, a quick demo I want to… oops. I was using Whisper Flow to get started, and I clicked on it quite early, so let me just do this. My VP of Product just sent me this. Conversion rate dropped last week. Can you look into it? So I used WhisperFlow to get me this, and Cloud Code, as you see, we have a bunch of skills and agents that we’ve built that we’re going to cover in the bootcamp starting tomorrow. This is the initial set. This is not using the fancy skills and agents.
This is something our question framework and stuff helps us get here, right? So, what does it do? It asks which conversion rate. Recommendation, like, you know, what do you do, and where it breaks. And then, what does drop compared to what? Prior week? Trailing week? Same week last year? Like, so many things could be looked at. and the hypothesis tree ranked by likelihood. Drop is one funnel stage, drop is in one funnel is not at all. The category is product. First check conversion by stage, week over week. And then you have things like traffic mix shifted, device browser broke, like, all these wonderful reasons that are the hypothesis for something like this to get it started, right?
And then it says investigation priority has checked these three things first, and then draft, reply to your VP. Isn’t it pretty cool? It gives you a Slack message as well. So thanks, looking into this today. Quick questions. Which conversion rate are you seeing a drop? Any known changes? changes last week, so it also helps you with that as well. So, this, I would say, is the most minimum claw code framework that it uses, based on all the data that we shared with it. And I also have a fancy demo later. After we cover the… Theory behind what we’re going for today. Okay? Cool. So, what the system produced, basically, in the one sentence, it’s not even a detailed prompt, right?
And I didn’t give it even any instructions, because this is… All of that knowledge, it gathered from the CloudMD files, it gathered from the, existing skills agents in the context we built in our repo. And what did it… was the output? The system clarified what conversion means, because there are four different conversion rates in most products, right? And it clarified what dropped means, and dropped compared to what. It generated 7 possible causes and ranked them, causes the hypothesis, right? And it told me where to get started.
Now, those likelihood rankings are based on general patterns, but in your company, you basically give it enough context that it picks up the context that you’re working with in your company, and it’ll help you with, like, what are the hypotheses for the context that you gave. This is something that it picked. Today in demo, that we are showing, like, generic context and generic hypothesis, okay? So the framework gives you, basically, the structure in which that you want to plan your analysis, and the type of questions you need to ask the person who basically, you know, asked you to go do the analysis.
I’m going to give you a free prompt today, basically that you could use in any of the web, tools in ChatGPT, Cloud, or Gemini, because, not every one of you would have access to the complete… some of the agents and skills that we’re building. That would be… we will be sharing them in the bootcamp. We have the bootcamp tomorrow, but we’ll also, share with… I’ll also share with you this free prompt. Along with this, you also have access to the free, oh, give me a minute. You also have access to the free air analyst repo that we have, that has a scale for analysis design as well. The complicated analysis design and tools and skills for it, that would be part of the bootcamp. Okay.
So, what I showed you is, like, the system working, right? One sentence in and structured plan out. The free prompt I just show you gets you the most value, and also the free AI analyst report that we have in GitHub, that could also give you most of it. But, if you want to run this investigation, in, like, a repeatable fashion, and you want to build those skills and agents yourself, we are running an analyst bootcamp. Tomorrow. is when we are starting. So today is the last day that you get into this cohort, guys. So here we have, like, a cloud code, a code for you to get at least 20% off. We also have a 5-week… oh, it says async, but it is also a sync course, sorry for that.
But we also have a 5-week course where we go through the entire detail of how… what we built in these agents. How do you even come up with these questions? What are the metric trees? How do we do the root cause decomposition? Like, as a data scientist in tech, what are some workflows you go through? We have a course coming for that called AI Analytics for Builders. Okay, now let’s go to the next step. Which is the 7-step framework that I was talking about that’s, built in into our cloud code skills and agents. So, before you run any analysis, you should be able to fill in 7 lines, right? I call this Analysis Design Brief. It’s not a checklist, it is, like, a thinking tool.
Each line forces a decision that most people, like, maybe At least the most people that don’t have context about how things work in data science, they skip. So, question forces you to pin down the metric, pin down the platform, pin down the timeframe. So, and this comes from clarifying what the initial question came out to be, right? And another very, very important thing a lot of people, especially younger, junior folks, miss is the decision part. So, especially if VPs and, you know, like, the head of something with big titles come up, they do not ask, hey. what if I give you this wonderful analysis that you’re looking for, answering this question? What would you do with it?
If they’re… I’ll… you’ll be surprised with some of the answers. Some of them would just say, hey, yeah, I just thought, you know, this could just help us, get more clarity about who… what a user wants. That, to me, is not a great decision. Imagine if they’d say, we want to run this experiment, and we want to prioritize this framework and assign engineers to go after this feature. That’s the type of decision you’d need. Anything that says something like, oh, I was just curious about this, is not a great decision that you want to spend cycles answering those questions with that analysis, right?
So ensuring that you have a decision ready before you go start exploring the analysis is very important so that you don’t end up spending a lot of time on it. And another, like, a bunch of things, like, you would need are the hypothesis. The hypothesis basically makes you commit to what to expect. So you can, you know, actually be wrong. Maybe there’s a hypothesis people say, hey, there is something happening with our new feature that just caused this, so that you could do the analysis and prove it wrong or privilege right. And comparison is the next one.
Comparison is where people default to compared to last month, without even thinking about seasonality, natural experiments, what’s happening out in the wild. That’s not related to your company, but, you know, something that’s happening, like, outside, like, in the ecosystem. And then the next one is Sigmunds. Sigmunds is where Simpson’s Paradox and other things hide. You might see something at a higher level, but when you start seeing Sigmunds, then you see total… sometimes you would see entirely contradicting things. New users are absolutely doing great, but it’s, something like, like.
existing users or older users, account age with more than 2 years, are where we are struggling for this metric. You might come up with something like that if you start looking at segments. And another very important thing is confounds. It is, like, confounds is the one everyone, to be honest, kind of skipped as well. So, three things that changed it once. And which one caused it? There could be something, like, happening in the ecosystem, happening, like, there’s an experiment, there’s a marketing campaign. How do you define which is the cause behind it? Having them listed out is absolutely going to help you get started.
And then criteria, is what separates investigation from rationalization, right? We’ll go through, basically, each of these, re… this with, like, real examples, and also possibly do this with cloud code in, like, one prompt. Okay. Bad versus good. Before I go, hi and Shane, anything in chat that I want to know about? Or pause? Because I’ve been talking too much.
Hai Guan: You can go ahead.
Sravya Madipalli: I can go ahead? Okay, awesome. Cool. Question. So the left version looks… reasonable, right? I mean, you’ve got a metric, you’ve got a direction, you’re even breaking it down by channel, device. But which conversion? Browse to cart, cart to checkout, checkout to purchase? Like, you know, like, last week, compared to what? The week before, same week last year, a trailing average. What do we mean by, you know, all of these things? Knowing these could mean so many things, right? So the right version paints down the exact metric, the exact platform, the exact comparison. And you can design an investigation for that.
And the decision on the left, if you see that, it’s like, we need to understand what’s happening with conversion so the team has visibility. That’s a pretty, you know, I would say something that Accounts do, like, are not a great decision that you would… spend a lot of time working on. So the right version names, like, two specific actions, like reward the checkout flow, or brief marketing, and have a deadline, you know? So… so that the changes, that you propose as part of the analysis, the recommendations, actually have, like, an outcome. And, it also helps you understand how deep you want to go in the analysis. or how fast you want to turn around the analysis.
So the decision is the one that’s… another one that’s pretty important. Okay, and let’s go through these quickly, as well, the hypothesis, comparison, and segments. The left version has a hypothesis, which is… I mean, better than nothing, right? But pulling the data and confirming is confirmation bias baked in. So, the right version is falsifiable. It predicts a specific magnitude, specific categories, specific customer type. You can actually be wrong, because you’re going to test in the analysis to go after, you know, verifying the hypothesis, and there’s a chance, a good chance that your data returns that, hey, this hypothesis was not true. And then conversion.
Week over week and month over month sound like solid, analytics, but last month might have had a promotion. The week before might have had a holiday. So the right version finds a natural experiment, also, along with all of these things, like a mobile removed the widget first, or desktop still has it, but mobile removed it, basically. Same customers, but different timing. So all of these, are something that you want to look at in detail when you’re talking about comparison. And the similar thing for segments as well. Trying to move on fast so that we have time, for the demo, okay? Cool. So, the next one is the confounds and the criteria.
So, we’ll control for the seasonality and check if anything else has changed. That’s… Sounds reasonable, right? That’s one on the left. So it’s not necessarily bad, but we can definitely, you know, enhance it so that it actually becomes useful. So something on the right that you see is that loyalty program Restructured same week. Marketing also ran a 15% off new customer promo. And guess what? You might also find something with the data pipelines, too. Like, data engineering changed the mobile tracking pixel mid-month. So, in reality, I’m sure all of you deal with this as well. I deal with this a bunch in my, like, in my experience as well, is that you… have multiple confounds.
happening at the same time, and that’s what makes this entire, you know, designing the analysis plan and also doing the analysis quite tricky. So having that written down helps you, like, go through the data and actually see attributions of these multiple styles, to ensure that You covered for them in your analysis. Now, what about the criteria? So, on the left, something that looks reasonable but not too great is that if the data supports the widget theory, we’ll recommend restoring it. I’m not sure if that’s, like, you know, the detail that we would want.
So, the criteria on the right that we’d want to be slightly more detailed is that, except if consumables dropped greater than 3 percentage points on mobile, but not desktop, because the loyalty change doesn’t explain the timing, right? And reject if the drop is uniform across all categories and platforms, then that would mean that it’s not a mobile versus desktop thing. So having the criteria locked down is also pretty important. Okay, so, one thing, I think I’ve said it multiple times already, but to make sure that I actually nail it down, is every failed analysis violates at least one line, you know?
So, the scoping black hole, someone says, look into conversion, and you spend 3 days on the wrong metric. The fishing expedition, right? No hypothesis means that you’ll find a pattern in anything. I remember, even at our previous place, Shane used to be very good at forcing stakeholders to, like, guys. Give me a bunch of hypotheses. You want me to go after this question? I have this hypothesis, but what about, like, if you have this question, ensure that you know the product enough that you give me these hypotheses.
I remember the staff conversations from our last experience with Shane and others, but yeah, make sure that not just you, you ensure that you work with your stakeholders and your teams to get these hypotheses written down, and also keep your, you know, stakeholders accountable. So that they have a list of hypotheses that we… they want us to look into as well. And, yeah, another thing that you’ll see is that, like. Simpson’s paradox is something that people miss as well, if we… the overall number lies and the segments tell the truth. So, ensuring that you have a bunch of segments already look into, and especially, not just any segment.
I know geo comes up very easily, and, you know, things like, device comes up very easily, but I’ll tell you something, in your company, in your context, there are a bunch of segments that are not as generic. In my company, specifically, if people were, like, a paid person before or not, is a segment that you wouldn’t probably see as, like, a generic segment that Helps for every type of product, but it is absolutely something that shows always having Simpson’s part of effect in my, like, you know, context. So what are the segments that, in your context, in your company, that you would see a difference, and you ensure that you look through them, as part of this analysis, right?
And then the attribution error is the other thing. Three things, changed at once, the confounds, basically, and you picked the one that confirmed your theory. But guess what? There could be so many other things that you missed out on if you did not write them at the get-go in your design analysis plan. And then 2 minutes of design prevents 2 weeks of wasted work. In the sense, 2 minutes is… I’m not talking about the entire design plan coming up with it, but actually the framework check. did I get the, question clarified? Did I get the hypothesis, all the hypotheses listed out?
You know, that little, ensuring that you check those boxes, would help you definitely, like, you know, get to a place where you don’t end up two weeks from now as, like, oh, no, I need a V2 of this. My V1 was entirely, like, not useful. So, ensuring that you go through these would definitely stop you from having, like, a failed analysis. Okay, so now a little sneak peek into the system that we do today. We basically do a kickoff analysis design scale. I heard myself for some reason. This basically kicks off all these three, like, you know, this entire system, and… in… you might have seen this in your place as well.
You basically have a question, and you come up with a V1 analysis, and then you share the V1 analysis. You have this meeting for it, you share it with your stakeholders, and then you get a bunch of questions, or stakeholder feedback. And then, you take all those feedback, and you take all those questions, and you run it again. come up with a new analysis plan that incorporates the feedback and also your VIN analysis, and then the V2 analysis pops out. This is the system that we do today, because this is the regular cycle that’s something that I’ve seen at my place, and I’m sure most of you see this as well. Question, V1, feedback, V2 design, and then V2 output.
So, I basically have this entire loop. I’m not sure if I’ll have time to go through the entire thing, but I can definitely share with you, the first one, the design for the V1, and the Weaven analysis plan by itself, and the findings, and what we two came up with. The last two things, I probably not have time to demo with you, but I can share with you the output of, what’s something I have already created that I could share with you. Okay. So, one thing that I want to, again, talk about is, in the bootcamp, you don’t just watch the system, like how you’re doing today. you basically will build it.
So the skills, the agents, the document pipeline, all of it becomes part of your personal AI analyst repo. You could already have this, because we do have a free AI analyst repo, but This bootcamp would cover a bunch of new agents and a bunch of new skills, along with helping you create them yourself, so that You wouldn’t need others to train you, or, like, you don’t have to watch this, but you could actually create one, and maybe you do this in your company, wherever you are, but you train other people on how to do this. So yeah, today is the last day to get into the bootcamp. Tomorrow, this weekend is when we are going to have the bootcamp.
Okay, so today’s analysis, we basically… I think, jumping into the demo, let’s… I’ll try to do the demo in the next 10-15 minutes, so that I have the last 10-15 minutes for the questions. Let’s see how we’ll do. So what is the demo that I’m going for? Today, we’ll be doing Cot Loop, online marketplace for home goods. This is something that you probably might have seen it in your companies if you’re a data person, if you’re a PM, you probably do this to your data people as well, or maybe your own analysis that you want to kick off. Basically, there is something with repeat purchase rate. Let’s say a PEM has a hunch, right?
I think our repeat purchase rate dropped because we removed the post-purchase recommendations widget last month. Customer support tickets mentioning, I used to see suggestions after buying have increased. This is absolutely a hunch, currently. It doesn’t have any data that’s backing up, it doesn’t have any, metrics that say, because, you know, they probably, maybe don’t have an experiment as well that looked at this removal and looked at the metrics change. Ideally. When these changes happen, you want to put that behind an experiment. But there are a bunch of reasons why, in companies this doesn’t happen today.
Every change is not behind an experiment, so the data team gets asked these questions, and… let’s get into understanding how do we do this. Okay. What is this demo tool going to be? We are going to basically take the hunch and basically go through the PM’s hypothesis. How do the agents sharpen these hypothesis, scan through the whole thing, and do the terminal? And then we’ll have this for you, the Weaven analysis plan, the Weaven findings combined with feedback, and then V2 Analysis Plan. So the terminal is basically agents thinking, and documents is basically what you’d share with your team. So, let me jump in into the… Cloud Code. Anything, quickly, Shane or hi, before I jump in?
Hai Guan: Attendee says, I’ve been the source of these hunches many times.
Attendee: Transparency.
Hai Guan: No, no, you know, no judgment here, Attendee.
Shane Butler: Well, you know, I kind of posted this on LinkedIn, I think, like, the whole analysis design thing is, like. a translation kind of process when… an analyst gets hunches or questions from… or comments from stakeholders, but I also think, like, it doesn’t have to be the… the analysts doing that can be, like, anyone, right? Like, if, like, stakeholders… non-analysts have… have their hunches, like, they can go through this analysis design Process to kind of, like. get it a little more into the language of the analyst.
Although, I don’t know, if I was a… if I also had a stakeholder that was like, hey, do this… do this very specific type of analysis, Claude Codes said you should do it, then I don’t know how to react to that either, but… I like… I mean, I just like the translation piece here on either end.
Sravya Madipalli: Okay, so what I’m doing, I’m calling an analysis design skill here, which in turn will call a bunch of skills and agents to get this started. So all I’m doing is whatever I just told you, I am, copy-pasting here, I’m giving it the data as well. Let’s see… What’ll come up. Okay, so analysis design pipeline. If you see what it’s coming up with. Give me a minute… I didn’t give it… So if you see this, it basically is talking about, analysis design pipeline. I’ll investigate this in three stages. Hypothesis sharpener, because there’s already a hunch, right? How do we make sure this hunch is actually a testable flame? So that’s the first one. And then the second one is Confon Scanner.
Find everything that could make this wrong. And the third thing is the investigation plan. Prioritize what to check first. So these are the three stages that it’s going around this, okay? So now, stage one, hypothesis sharpener. It basically is looking at the stated cause, stated effect, implied metric. implied timeframe, because it was a pretty detailed, like, you know, hunch with details. And also the assumptions. The assumptions was widget was the primary driver, the drop is uniform, like, no other changes occur simultaneously. Those are the assumptions, obviously, because if that’s something that you have a hunch. And then, what is missing?
It basically talked about everything that’s missing here. So, it has a testable hypothesis and also a secondary hypothesis. It also has the metric definition, it has what is, like, you know, gives you notes, like, the nuance as well. Like, repeat purchase here means any subsequent purchase, not same category purchase. Like, and also talks about, like, this is going to be different if it’s going to be that. And then it talks about comparison groups. Natural experiment, like, what this versus this, and widget users versus non-widget users, before versus after, the same population and stuff. And then it also has accept or reject criteria.
Accept if widget users show greater than 2X, or, you know, things like that. Reject if, you know, so and so. So these are, like, a bunch of things. Like, this is the analysis design brief that we have, and investigation priority. what are the multiple steps that we have, and it prioritized it based on why we want to prioritize it, right? So… and then it has a stage one summary. So what it says is, testable hypothesis, best comparison, and what is the key insight. Now, it starts at stage 2, so it gave you the debrief of what it’s going through fully, and then goes through the stage one, gives you a debrief of what’s your hypothesis sharpener, and then goes to Stage 2.
All of these demos… so this is a demo, right? All I’m trying to show you is what is the capability of claw code here. How do you do this at your work is something that we could discuss in the bootcamp, and also something that you basically give it the context, give it the memory and skills that your company and your business would you know, have the nuances of your area, which It would take into consideration while coming up with this. Different stages and, you know, the confounds and other things. Okay, now what is the stage 2? It starts in the Stage 2, and there’s a confound scanner. In the Stage 2, it’s basically going through The claim.
Removing the post-purchase recommendation widget led to a 3.77 BP decline, and it comes up with what are the other concurrent changes it found? So, how does it, in an ideal world in your company, how could you do this? What I do is I have MCPs attached, to my experimentation platform. I have MCPs attached to my Slack. most important Slack channels, where we keep sharing our experimentation updates. It… basically, my system, my cloud code, has skills and agents that grasp these contexts from these multiple places that you gave, early on. And this is to demo you, that we kind of already… this already picks it up from the existing dataset, but that’s how you would do in your, like, workplace.
Okay, so it goes to the confirmed scanner, it found, hey, product has these changes, marketing has these changes, technically these things happen, and there are also data quality threats, the severity of them. And what are the selection biases? Like, a bunch of things, a bunch of alternators and their plausibilities. And then it gives you an entire confound scan report. Right? And then this is stage, this is the stage 2 summary as well. And now at the end, we get stage 3. So this is the final analysis brief.
It basically goes through the question, removing, the widget, and then the decision, the hypothesis, so this is, like, the seven, like, the framework flow that we had, and this is what it gives. Let me give Claude Ward. Okay, cool. So, review the plan, or should I proceed with Vivin execution on the data? I will ask for it something. Let’s see if this is going to do this live well or not. If not, I’m going to share with you the Weaven doc. Okay, can you… yeah, I’ll talk to this. Can you move this design plan into a Google Doc that I could share, with my stakeholders? Use the Google Doc, Creator, and export sales. Let’s see if this comes up. So, until then, I… Let me read.
So, it’s basically what it’s doing currently, is it’s checking a bunch of skills and agents that I have to ensure that my Google Auth is correct, and, you know, it has the Google Export and the review skill set. Okay, while this happens, anything that’s happening in the chat that we want to chat about, this might take… A few minutes. Shane or hi?
Hai Guan: Yeah, Attendee has a question, and this is probably… prevalent, I think, from a couple of folks as well. How much do you need to explain your company slash product to the tool? Or does Claude figure it out somehow?
Sravya Madipalli: That’s a great question. like, if you’re… I don’t know if you all tried Cloud Code in the before version, when compared to that, it needs a lot less information, for sure. To me, in my experience, if, if you have… oh, looks like my token expired, sorry. Okay, send it to… let me stop sharing my screen so that I give it my auth, and then I can get back. So, I was kind of winging it to try this online, like, live, but, I mean, this is not, this is something that you’ll definitely run it. Hi, and Shane, do you want to answer that while I do the auth setup for Google?
Shane Butler: Yeah, I think so… I think it’s pretty good at reasoning, and it’s pretty, like, intuitive in terms of, like, even if you didn’t provide that much detail, but obviously it’s gonna be… way better if you do, but it doesn’t necessarily have to be, like, you sitting down and explaining everything. Like, I will literally… so I have it hooked up… hooked up to, like, Notion MCP and Google MCP, and I also have repos with a bunch of, like, my past analyses. And, I sometimes go as broad as I’m just, like. go search Notion and Google, or… and I have it connected to my Slack, too. And my Slack for, like.
this topic around this product and past things I’ve done in this area, or if I want it to be a little more focused, I’ll just copy and paste a bunch of Notion links or Google Doc links, like, hey, read through all this relevant information documentation. Also, something extremely… I mean… like, cloud code is ORM, right? So it’s, like, really good at reading through codebases. So, if you can just give it the repo to your product’s codebase, it can make… connect a lot of the dots in terms of, like, how things flow through as well. So yeah, MCPs, other sources you can connect to make this thing, like, way more powerful. So… I don’t know, I’d say the more information you give it.
The better, but, like, even, like, no information’s gonna get pretty good.
Hai Guan: Yeah, if you… there’s a… there’s also a really fun exercise to do, is that if you have it just crawl your company’s website, it could… it will be better than if it didn’t have that context. So, any incremental information you provided, it gets that much better.
Shane Butler: And then Attendee had a question for you, Sravya. Could you remind which, information sources Cloud Code accesses here? Like, what does it interrogate to check? For other, concurrent changes to the target metric.
Sravya Madipalli: So you’re talking about this demo, per se, or generally it worked? Like, you know, where it could get this information from?
Attendee: Either one, are there the specific case here, or… but I think Shane also answered it generally, and the rule of thumb is the more information, the better.
Sravya Madipalli: Yeah, yeah, exactly. So, Claude is so darn confident, it… if you don’t give it information, it’s going to take information from what it knows, obviously, from the web, about your specific area or context, and the more information and the nuance you give at your work. the better it gets. And the first time you work with it, the feedback you give it, and the next time, and the next time. And by the time you work on a week or two weeks with it, it’s super smart already, and it does work, like, way faster than you would, if you had, like, a bunch of these, you know, skills in place already on how to talk to it, how to create these things from the get-go. So, guess what? You see this?
Google Doc created and formatted, B1 analysis plan is in this deck, okay? Let me… oh, I’m sharing my cloud code, so let me share with you what it produced. Okay. This is what it produced, guys. It basically gave the Vivan analysis plan repeat purchase, the context, the questions, the approach, the known risks, the investigation priority, what would change your conclusion, and the deliverable. This is something that you saw, that… purely, like, was generated by Cloud Code. And, once we have this, it basically… would we… we could use this and start the findings.
So, ideally, if I had time, because we only have 15 minutes, I could share this analysis plan, feeding that to Claude Gold, and Claude Gold generating the analysis itself as well. And then, we could come up with, hey, this is the feedback on that analysis, can you come up with a V2 design tool? So, I have all of that for you that I kind of already did, and I’d like to share that. Give me a minute… Okay, So, what it did, is I fed it, I basically did this on the background. I got, the Weaven analysis plan, and I got it to execute on the analysis plan because it had the dataset handy. And then it… I gave it feedback on it, and then it gave me the V2 analysis plan. So that’s what you see.
This is the second one. I couldn’t do that in today’s demo, because it’s too long, but that’s what it says. It’s Analysis Plan V2, revised after we even find it some stakeholder review. background. In V1, we investigated so-and-so. It basically gives you all the context to it, and it gives what V1 established. There is a real material drop, and you know, all of that. But it also talks about what V1 could not resolve. Because we generally don’t answer all the questions on the earth about, like, you know, something in the first iteration.
So you go in the second iteration, and then we talked about what was not handled, and then it talks about the stakeholder feedback incorporated, it takes the stakeholder, this was the concern, and this is how the V2 addresses it. So it talks about the revised questions, it talks about approach, and you know, and a bunch of things. Yeah. I mean, each, company does it differently. I just shared one approach. The other approach is a set of questions, a sort of segments, and all of those written down as well. So yeah, we have 14 more minutes, but I want to share something pretty quick with you all. This generally, I always really like it, at the end of any session.
We could ask Claude, hey, can you share the… architecture of how… You’re solved. the above. Huh. So, it basically gives you everything that it did. This… every time you work on a complicated problem, you want to understand, like, how it went about and did it. We’ll see, basically, the entire loop. You could ask it to be in more detail than what it is. What it’s giving it right now. So, this is what it is, slash analysis design. I would repeat purchase rate dropped, architecture preview, shows the user three-stage plan, data profiling. coming up with this, and then the stage 1 came in, and then the stage 2, and then the stage 3, and then the auth pre-flight, we had an issue with auth.
Ideally, if you had auth, it looks like it’s been a week, more than a week I did my auth, so it triggered my auth, but if not, you wouldn’t have the stage. Let’s say you have this, and then Google Doc Creator. We also have agents and skills for Google Doc Creator and Export. And yeah, this is how the entire thing… it talks about how many skills were used, how many agents were involved, and what tools were used, and what MCP was used. Okay, this is where my demo ends, and we have 12 minutes.
Hai Guan: Attendee have a question. Do we cover any data visualization slash dashboarding in the bootcamp or 6-week async courses?
Sravya Madipalli: Absolutely! That’s literally, one of the very important, aspects, of, the courses. So we basically cover even more than that, Attendee. We cover visualization, dashboarding, and the most important thing, storytelling, and also what you want to visualize and dashboard, basically question framing as well, because all of those are very important into understanding that. So, yes, we do cover that in the, 6K course, and it’s not async. It’s totally sync, like, it is basically live bootcamp. It’s a live course that, hi, Shane and I… we run, yeah. I’m sorry for that typo in the slide, but yes, we do cover that.
Hai Guan: Did I miss any questions from anybody? Anyone else have any questions? Feel free to come on… come off mute as well. Yeah. Attendee, do you want to speak to your question? Attendee says, take my money, okay.
Attendee: Yeah, absolutely. Hi, thanks for the nice presentations. So, I just want to highlight one of the problem statements I’m getting on a regular basis, like, I have the semantic layer, I have the visualization part ready for the AI part. What I’m struggling is how to connect the particular semantic layer to the visualization part. Like, I have the data in my database, like, data warehouse like Snowblake. And I want to visualize, using GPT or Claude, anything, anything that is working perfectly fine. So… how to connect that particular piece, that particular Snowflake data to cloud, and that… that should be refreshed on a daily basis.
So that is some… that is something I am facing a major challenge in between. Like, so, are we covering that particular piece? In incoming… upcoming segment, like, how do you plan… are you planning to do that?
Shane Butler: I think we’ll go more into, like, the full, like, how do you implement this into, like, a… production environment, more in our 5-week course. So, in our 5-week course, we actually will work with Cloud Code, we’ll also work with some other third-party tools, like enterprise tools that do some of this for you. this weekend, what we’ll talk more about is, like, how we, like, build an agentic system from scratch, and I think we could work with you in it, actually, to be like, hey, let’s try and build… an agent and scale that… refreshes some, like, HTML dashboard on a regular basis. We go into MCPs a bit as well, too, so I know that there’s, like. depends on, like, what your company has, right?
Like, we have, like, Sigma at our company, and they have an MCP where you can create dashboards directly from Claude code, so, like, could… potentially play with something like that, probably more in the 5-week course, but I don’t know, maybe we could try and mess around with it tomorrow, or… and Sunday. We’re basically gonna have, like. How tomorrow’s structured, we’ll have some, like, lecture time, example, demo time, and then we’ll have big blocks both days for about 90 minutes, where we’re gonna do breakout groups. From, like, beginner intermediate to advanced, and everyone will be kind of, like, building their own thing with one of the instructors involved.
So if that’s, like, a project you wanted to work on during one of those sessions, I could definitely, like, try and play around with that with you. But… Yeah, obviously, production pipeline’s more involved, so probably, like, more of the 5-week course. We do run a… a kind of two-for-one thing with the 5-week course, so if anyone takes the boot camp and wants to do the full five-week course afterwards, we just deduct whatever you paid for the boot camp from the 5-week course. So you basically get two-for-one, without having to commit to the 5-week course up front. Did that answer your question?
Attendee: Yeah, that answers my questions. So, there are, like, a particular two problems I’m facing, like, one which I’ve mentioned, another one is the development part, development, the particular dashboard into some… somewhere so that, everyone can… I can… I can say with any stakeholders on a regular basis, or they… they can refer that particular link so that they can get to know about the numbers without depending on me or my team. So, those are the two Sweet one.
Shane Butler: One way… when we have dinner right now is so, So, like, if you were to create, like, a customized dashboard, right, you have to host that somewhere internally at your company, which is totally feasible. We’re not necessarily going to go into that, because every company is completely different with, like, how they want to host things internally. So there’s, like, the MCP versus, like. with, like, like, your… whatever your BI tool is route, if it offers it. Another thing I actually do is I have a skill… Where, you run this skill, and it refreshes, like, a Google Doc or a Notion page for, like, a very particular stakeholder that has, like, a full-on report of everything I know they want.
That’s, like, a kind of nice, like, hacky middle ground way to do it as well, and that’s definitely something we’ll cover, this weekend.
Attendee: Yeah, and about the BI tools, we are as a team, trying to eliminate the BI tools, because it’s a pain. It’s a pain to maintain and build using BI tools in the era of, like, AI. So that’s what… that’s what I’m trying to, save some time and eliminate the BI tools. But happy to know more about that. But I think the major problem I’m facing is about connecting the data warehouse to the AI part, where I’m… so, yeah, happy to know about that particular piece.
Shane Butler: Yeah, I’m on your side about eliminating BI tools. I don’t know how everyone is, but I never want to maintain a dashboard again, so I like this idea of eliminating them, and even getting stakeholders, self-serve empowered enough where they can be building their own dashboards themselves, because if you think about it, like. they know what they want, they know the question of what they want, but they need, like, the trust and the data, so if we can create a system where they’re like, hey, I want to add a filter to this, I want to add another dimension, but I don’t have to go to an analyst, I can just do it myself, that’d be pretty cool.
There’s obviously trust stuff with the data there, but… that’s kind of where I see the future going.
Attendee: Boom.
Hai Guan: Yeah.
Attendee: Thanks, Anthony.
Hai Guan: And the line of sight is pretty clear. If you have one of these, like… I mean, even if it’s not one of these, kind of, like, big-name data warehouses, like, you probably have some, you know, Snowflake, Databricks, or something like that, there’s a very clear path to doing that.
Sravya Madipalli: Yeah, I mean, Databricks, Snowflake, I’m sure all of the big-name ones have MCP set up, but for other things, that probably your company doesn’t have an MCP or don’t have one, you know what? What… something you could do is show them what good looks like. hey, if I had access to this. you put a CSV there, and show them, guys, this is what I want, and I’m going to help you just, you know, get so much more money for free if you had this, you know, like, these type of systems set up for us, because in some companies, I know security is a big thing.
People are worried, about that, so… We don’t, there are a bunch of ways to go around it, like, to ask these tough questions, and you know, management, when you show what is possible, and what is stopping you from getting there, they would… we could… people will find some ways to get you there.
Attendee: Yeah, absolutely. I agree with that point. Yeah.
Sravya Madipalli: Huh?
Shane Butler: What other questions? Anything else?
Attendee: Hello, I have one. Thank you first for all of this presentation. I think now… now is the time where all the data should be self-service, right? Because… Actually, I’m pointing to get that in my company. But… what happens if I apply this, for example, with the after that? Where… where can we… Set up some kind of feedback loops, or framework of… Improving data analytics. For example, we have something that we want to decrease, which is the retention in my company. And we are preparing with the… with code, like, something to prevent that, okay? So, in your business area, or in your companies, how do you manage?
to, Possibly propose the plan, but actually, get some feedback in feedback loops to re… to make this work, actually something functional, right? Because you are providing the insight, you are telling where we have to go, and how that happened with that insight after. Are we going to cover that maybe in that session, or…
Shane Butler: So I think you’re… so, just to… to come up to, kind of. summarize. So you’re asking, like, after we get to, like, the insight from the analysis, how do you then, like. Implement. That afterwards, and then get, like, a feedback loop going, so, like, Consistently is being executed.
Attendee: Exactly, exactly. It’s… how do you manage that cycle? Because this is potentially… it’s very help… For… for make the insight and validate hypothesis. Yeah. But in your framework, how do you manage to, okay, how can we continue improving that, or seeing if the people are applying the action plans that maybe we can propose, right?
Shane Butler: Yeah, we do, so… So, in our 5-week course of the last week, we spend on, basically, like, insight to action. It’s more about, like, how do you, like. It’s more like an organizational change, how do you, like, share out your information? But I think, so there’s that piece of it, where it’s just, like, presenting and… Getting buy-in, etc. But I think there’s also, like, a question there around, like, okay, I… armed with AI, like, I have these insights. can I do it myself?
Or can… so I’ve been doing some of that at my job, which is… this is kind of what I think is the future of, like, the data role, where it’s like, you start going more end-to-end, where instead of just us doing analysis, we start going up funnel and building data foundations, but also start being like. alright, we have some insights, like, can I now… whether I work in go-to-market or in-product, can I actually go into some code and leverage cloud code to, like, try and test and proof of concept some of these things out? So I’ve done a bit of that at my day job. And, like, I’ve found it pretty… compelling to be like, I’ve got the insight, and now I can, like.
kind of have, like, a sandbox of, like, the production code where I’m gonna try to apply it myself. And then I work with, like, software engineers to be like, hey, like, I’m actually seeing, like, increases in retention, or decreases, or in the predicted retention, or whatever, based on what came out here, can, like. I kind of pass this off to you to, like. implement that, or, like, work on it some more. There’s also this whole thing the past, like, couple weeks, Auto-research, like, from… Carpathy, where it’s, like. An actual, like, autonomous loop where the code Is, like, if you have some type of metric to optimize the code, like.
The agents change the code itself, then see if that, like… brings your metric up and down, does an analysis, comes up with new hypotheses, tests it again. I think there’s, like, a whole field of software engineering that goes that direction. We’re not going to cover that stuff in this boot camp. And we’re probably just gonna cover, like, more of, like, the share-out stuff in the 5-week, but… Yeah, I’m playing with an idea about doing, like, a two-week course on that called… it’s called Automate AI Evals. And maybe we’ll create something else, too, but… I think definitely, like, coming out of this. you should feel empowered to be like, hey, can I… can I go the next step forward?
Like, the insight doesn’t have to stop at me passing it off to a stakeholder.
Attendee: Excellent. Perfect, Shane. Thank you. Thank you very much.
Shane Butler: No problem Cool. Hi, I think you had a question.
Hai Guan: Yeah, do we want to answer this one last question here?
Shane Butler: Yeah.
Hai Guan: Okay, cool. So… Actually, Attendee, if you’re still on… on the line, do you mind, sort of, like, reading out your question to the group?
Attendee: Hi, thank you so much, Pai, for giving me the opportunity. I’m not quite sure if my background is quite noisy here. So I’m currently looking for entry-level position as a new grad. And in some… to some degree, that I find it even more competitive than, like, a junior or mid-level positions. But I’m also in the position that I don’t… feel comfortable validating all the uploads or codes generated by large language models on my own, which is different from this senior data science position.
So I’m not quite sure at this stage of my life if you would suggest someone like me to first focus on getting myself… In the board first, by drilling those interview… style questions, like, time to put back those, instead of… focus too much on, integrating AI workflow by, Pardon.
Hai Guan: Got it, yeah, I think that makes sense. Shane, do you want to take it, or I can take it? Pete away.
Shane Butler: Why don’t you take it, and I can add… I can add some color, or whatever.
Hai Guan: Yeah, I’ll take a crack at it. And so, you know, like, my thought is, the… sort of, like, AI is gonna help almost, like, accelerate your learning, if you will. Like, if you… If you actually look at the open source repo, like, it’s got a lot of best practices for, you know, like, the many different components of being a data scientist, effectively, in the analytics space. And, you know, like, being able to leverage that from a, for example, like, you know, like, hey, I, I know what good looks like for, you know, like, framing questions.
I know what good looks like for designing and analysis, I know what good looks like for validating and stuff like that, and I know how to use it with… or I know how to, I know how to do it with the help of AI. in this new day and age is actually pretty powerful to, like, you know, like, be, like, your learning companion, along the way. So, if I were understanding your question correctly, the advice I would give is to sort of, like, use it to supercharge you.
I mean, like, there’s the component of, like, you know, hey, we want, like, you know, should we understand the fundamentals and, like, you know, do it really, really well, first, before kind of, like, you know, the math before the calculator, right? Like, that analogy. I actually think doing both at the same time Would actually give you an unfair advantage, over everyone else who’s kind of, like, you know, just I don’t know, like, grinding the traditional way. I don’t know if, Shane, you agree with that.
Shane Butler: Yeah, I think so. I think… I think it’s, like, It’s like a… it’s like a challenging time on the other side of the spectrum, where it’s like, if you’re coming in… totally new, now there’s this, like, very new… AI thing that is able to do a lot of the stuff that, like, was more kind of entry-level analysis work for another side of the spectrum. Like, a totally new way of working is… is being created, and so people who have been doing this work for a decade or two or three decades are having a hard time adjusting to basically unlearning things that used to be very good habits that now become bad habits.
So… I… I think… the only thing I would add to it, I agree with everything Hai says, and the only thing I would add to it is I think, like, the job of, like, most people are changing, is changing. No one really knows what it’s gonna be, but, like, at least right now, like, my bet is, and where I’m investing, is that The job’s changing from, like, me doing analysis to me building agentic systems that do analysis for me. right now, I have to validate a lot of that. Who knows what happens later on as these get better and better, you’re gonna need less validation, less validation, just like… with, like, you know, GPT 3.5 a few years ago, it was like, oh, it hallucinates all the time.
And now it’s like, yeah, it still hallucinates, but, I mean, like, you’re not checking for hallucinations as much as you used to, and that’s probably gonna happen, like… Industry ride with other validation stuff, too. So either way, like, I think my, just, TLDR there is, like, investing in learning how to, like, onboard and nurture and build out Agentic systems, I think, is gonna be kind of, like. a new… Valuable skill, rather than, like. Just doing kind of, like, the traditional way of… Learning analysis. I don’t know, anything else high on that?
Hai Guan: No, you nailed it. I mean, definitely a tough time, and it is our belief that if anybody, not just experienced or not just entry-level, who master and are comfortable using, for example, right now, Cloud Code, is probably gonna have, again, back to the same spiel, unfair advantage over others. Okay, cool, Thank you, guys. Unless there’s any question, you can catch us in the Slack community, and yeah, hopefully, I think we’ll see a lot of you, in the bootcamp tomorrow, if you want to join. Please feel free to join us as well. Thanks, everyone.
Attendee: Thank you very much. Good stuff. Hi, Nathan.