Transcript Copy transcript 0:01 Welcome to our session on Evaluating AI for Social Impacts, the four-level framework. 0:09 My name is Kelly Zhang. I'm a researcher and data scientist with the Agency Fund. 0:13 And this is... 0:14 I'm Edmund. I'm a software engineer with the Agency Fund. 0:21 So the structure of today's presentation is first, we're going to give you some background on the four-level evaluation framework that we have. 0:28 This is just to give you context for the demos that we're going to be going through. 0:31 First, we'll be doing a demo on Level 1. 0:33 This is doing model evaluation on a platform that we built called Calibrate. 0:37 This is something you'll be interacting with directly. 0:39 And in the second part, we'll be doing a Level 2-3 demo that's basically A/B testing on Evidential. 0:44 You'll all be a part of this experiment. 0:46 So this is the part I wish we were able to ask more questions. 0:49 But I think for the background, we'll kind of just go through the slides first just so we make sure you actually have time to go through the demos. 0:54 And then finally, we'll have a feedback form in Q&A at the end, so there'll be plenty of time for questions at the end as well. 1:01 Okay, so we're at a conference, Africa Evidence Summit, and we know impact evaluation, right? 1:06 This is typically what researchers are focused on. 1:08 But we want to also emphasize that there are multiple levels to consider before evaluating the impact of an AI product. 1:14 And our four-level framework, impact evaluation, is Level 4. 1:18 So it's kind of the last level, but also one of the most important ones, right? 1:21 Whether or not product usage of an AI product will actually improve development outcomes. 1:25 This is usually what people here are specialized in. 1:28 But we really want to emphasize that the first three levels are also important. 1:32 First, if the AI model is actually working, as you think it's working. 1:35 And actually, there's a huge role for domain experts like you to play in helping to ensure this, checking for model accuracy. 1:41 I think over 90% of you said that you've used AI tools before. 1:44 I'm sure you guys have all encountered the fact that AI sometimes hallucinates and is not always giving you the correct answers. 1:50 And this is where, especially when you're deploying things in health, education, economic types of things, 1:56 you want to make sure that the model is actually accurate. 1:59 Then we also have product evaluation, whether or not the product is actually engaging and retaining users. 2:04 Basically, if no one's using the product, you're probably not going to have an impact. 2:07 And this is in the realm of tech, where usually product managers and data scientists optimize. 2:11 This is where, like, Facebook, Google, WhatsApp, they're really optimizing to make sure you actually engage with their products. 2:17 And we should also be thinking about this in any impact evaluation that we run. 2:20 Are people actually using the AI product that you're deploying? 2:23 And then the third level is user evaluation, which is very closely tied to impact evaluation. 2:28 I would say often we do this in piloting for RCTs. 2:31 You're already trying to understand how people are actually using the products. 2:34 And I think even outside of AI, we can think about products that are being deployed where people aren't necessarily using them the way that you think that they might be using it. 2:41 The important thing is when you build an AI product, that they're actually using it in a way that is moving them towards a development outcome. 2:47 So, for example, I think we all know the story of malaria bed nets, where you gave people malaria bed nets, but they used them as fishing nets instead of actually using it to prevent malaria. 2:56 And that's an example of where a product may not be used in the way that you intend. 3:00 And that's also where level three is really important before you even get to level four of the impact evaluation. 3:05 So we're not going to talk as much about impact evaluation here, because I'm kind of assuming this is something you guys are all already familiar with. 3:12 But we will talk a little bit more about it in the first group. 3:17 Yes, so let's start with level one, model evaluation. 3:21 How do we ensure the AI model is working as intended? 3:24 All right, so a bit about how large language models actually work. 3:35 So fundamentally, they are predicting the next word, or technically the next token. 3:42 And based on the training data they've been trained on. 3:46 So in this example, if a piece of text says, agency means, it could be a government body. 3:53 So the scope is that it might be a government body. 3:55 But for a piece of text that says human agency means, you know, it's more likely to be personal entitlement. 4:00 So when dealing with a model whose inputs are based on probabilities, we need to invest extra effort in evaluation to make sure those outputs are what we want them to be. 4:13 All right, but, you know, in a more production-like AI chatbot, it's not just predicting the next token. 4:22 There's some extra layers around this that we add. 4:27 So here's a chatbot you'll be interacting with in the interactive portion. 4:31 At the agency fund, we get in-kind advice, you know, software engineering, data science, behavioral science advice. 4:37 So we've included that in a demo chatbot today for you to put on. 4:41 So let's talk a bit about how an AI chatbot like this actually works on the web. 4:54 Okay, so the LLM that we just talked about is just one piece in a typical architecture for an AI chatbot. 5:01 You know, that's the foundational model that I referred to here. 5:05 So if you're building a voice chatbot, you want a good processing stage to turn the audio into text. 5:12 Or, you know, if you're building a chatbot to support, you know, non-Western languages, you know, like Amharic and others, you often want a translation set. 5:23 You know, to translate the Amharic to English because these models perform best in Western languages. 5:28 Those are languages that have a lot of training there. 5:31 And then there's a context layer. 5:33 So if you want your model to, you know, benefit from some purpose information, you know, like a health manual or something like that, you will encode that information in a knowledge base. 5:44 If you want the model to have access to tools, the ability to search the web or generate a report, you can get the model access to that in a context layer. 5:52 And then there's a post-processing step. 5:54 So if you're building a voice chatbot to turn it from text to speech or, you know, if you're supporting Amharic to turn it from English to Amharic, et cetera. 6:03 And this post-processing step is often where the safety layer comes in. 6:08 So if there are certain kind of responses you want to prevent your chatbot from having, you usually enforce that as a post-processing step. 6:14 So, yeah, that's what a real ballistic chatbot architecture looks like. 6:25 Anyway, so that's the architecture that kind of powers responses like this. 6:30 So we asked, hey, can Kenyan mothers generally eat avocados while pregnant? 6:34 And you see the response refers to the further knowledge base, avocados are generally safer for mothers to eat. 6:45 All right, so now I have a question for you. 6:47 I would like to hear from one or two of you. 6:49 Where do you see AI fail in your work or your personal use? 6:55 So, yeah, you. 6:57 Yeah, I would ask AI to tell you what a helmet is. 7:03 Tell you what a helmet is, as in a motorcycle helmet. 7:06 And it elucidated a different answer. 7:08 Yeah, it glows on and off. 7:10 Yeah, yeah. 7:12 Yeah, no, that's a very good one. 7:14 Because of the nature of predicting where it comes next, it can put together something that sounds like it makes sense, but it's actually wrong. 7:22 So it can confidently give you the right answer. 7:24 It's a great example. 7:26 How about one more? 7:28 Yeah. 7:29 Yeah. 7:30 Yeah. 7:31 Yeah. 7:32 Yeah. 7:37 Yeah. 7:51 Yeah. 7:57 Yeah. 8:00 Yeah. 8:01 Sorry, I didn't mean to interrupt. 8:03 Your voice was kind of low. 8:05 Did I hear it right? 8:06 That AI is kind of, can it cite sources that don't exist? 8:10 No, that exists, but the authority or the password written in it is not relatable or present in data. 8:20 So if you pick it anyhow, you can mislead yourself. 8:25 Yeah, yeah, that's a great one. 8:28 So I should say that AI will cite those sources, but those sources are not necessarily the most authoritative source. 8:38 They're just the one that happens to be indexed by Google and these other tools. 8:44 So, yeah, that's a really good point. 8:46 So one more. 8:49 So sometimes I feel like when you input a lot of information, it kind of spasms. 8:56 So there are times where I'm feeding it, like, maybe to my face, and it's like, I can't read all this. 9:02 I'm afraid to, like, use this information. 9:05 Like, what's the point of the AI? 9:09 That's another good one. 9:11 Great. 9:12 Yeah, those are all great. 9:15 And just to categorize some of those failure modes, you know, biased training data, misinformation, unverified claims is a very common one, 9:23 poor performance in low-resource languages. 9:26 So just a quick second. 9:28 By low-resource language, we just mean that it's a language where there's often a lot of information on the Internet. 9:36 So these model companies, they train these models by scraping information on the Internet. 9:42 And, you know, for example, there's way more English on the Internet than Amharic. 9:48 So we call English a high-resource language because it's all over the Internet. 9:54 Whereas Amharic, we would call it a medium or low-resource language. 10:01 All right. 10:02 So now we understand the need for structured evaluation. 10:06 These models are somewhat, when we say stochastic, you know, 10:10 they output responses kind of randomly. 10:13 So we need to guide evaluation with some guiding questions. 10:17 And, yeah, here are some. 10:19 So are the AI responses grounded in credible sources? 10:23 Are the failures we see in the model because of the model itself? 10:27 Or how are we going to prop up the model? 10:29 This is important. 10:30 And again, are users in a literacy context getting accurate, safe, culturally-tuned responses? 10:37 Can we understand why AI would have produced different outputs for the same inputs? 10:42 And if we're building an AI tool, can we identify when to discard the AI response completely 10:49 and fall back to the actual data? 10:52 And, you know, in the social sector, we're often dealing with vulnerable populations. 10:56 Yeah, it's very important to take evaluation seriously. 10:59 We can't skip this. 11:02 All right. 11:04 So let's talk about how to actually do the evaluation. 11:08 You know, the re-evaluation. 11:10 So when you think about an intervention that kind of relies on humans, let's just say like front-line workers, 11:17 it can be helpful to think about how we would evaluate those humans. 11:21 So one is how much, you know, the throughput for every front-line worker, 11:27 how many people they're going to serve. 11:29 So, you know, it's a fixed number of humans, but in reality, with AI, we're able to serve much more people. 11:36 And, you know, after a front-line worker serves a beneficiary, we can deploy and use a service to that beneficiary. 11:43 And I said, okay, you know, how was this experience? 11:46 How good was it? 11:47 So this rich information, as well as outcomes tracking. 11:49 You know, if we're building an agricultural advisory track, can we measure things like agricultural yields 11:54 to see, like, is this advisory actually helping? 11:58 And now when we're talking about AI interventions, there's some things to consider. 12:04 The first is that because it's AI, we have the opportunity to really test things out 12:12 before deploying it to one real human, one real beneficiary. 12:16 And this is a very important ticket out of that, which we'll get into. 12:19 And it's also more important than ever to define what good looks like up front. 12:25 Because with humans, we can rely on common sense a little bit, even if you don't write up front 12:31 what all the front-line workers are supposed to do. 12:34 With humans, we benefit from a bit of common sense that we can improvise at the moment. 12:38 But with AI, it's very important to define what good looks like first. 12:47 Alright, so here's some examples for our model evaluation criteria from one of our grantees, Norah Health. 12:53 They operate a medical health line in Southeast Asia. 12:57 And, you know, one we focus on, this is an emphasis for both humans and AI. 13:03 But with AI, we've had to get more crisp about, like, hey, what kind of secret response is medical health? 13:09 One is factual accuracy. Is every claim medically correct? 13:14 Is every claim easily understandable? Is it grounded in Norah's internal protocols? 13:21 And, yeah, does it have human empathy to the responses? 13:28 Alright, so given that, okay, if you're building an AI product, you have this opportunity to test it out 13:37 before deploying it to one user. 13:40 There's a common trap that we saw early in the organization deploying AI into, 13:44 which is that it has to be absolutely perfect before you deploy it to one user. 13:50 That, like, you would have to spend months testing out these models 13:54 before you feel confident enough that it's safe to deploy it to one human user. 13:58 And that's something that we want to break a little bit, because you will learn more from real users. 14:04 Exposing this AI product to a small group of privileged owners and facilities 14:09 is something that will help you get your product out the door faster while still being safe. 14:16 And we have created this minimal, viable evaluation checklist to describe just that. 14:24 Like, hey, before you deploy your AI checkout to one user, 14:28 here's a checklist to help you feel more confident that you have evaluation in place 14:34 for both your organization and external partners. 14:37 What is a golden data set? So what is a golden? 14:40 A golden is a representative query to your AI chatbot. 14:45 Like, earlier I asked, hey, can Kenyan mothers eat avocados? 14:49 This is an example of a golden. 14:52 I have an idea of what the ideal answer should be, 14:54 so I want to keep that in a sort of spreadsheet 14:57 so that we can track the performance of the AI model on these representative test queries. 15:02 Number two is bake in the input of the domain expert. 15:08 So what do we mean by the domain expert? 15:10 It's the person who understands what a good response looks like. 15:14 And that's not always the person who's building the model. 15:17 Developers can build the model and get something working, 15:20 but in the context of a medical healthcare, it's the nurses and the doctors 15:25 who should be the final authority on what a good response looks like. 15:28 In the agricultural advisory chatbot, it's the agronomists. 15:33 In the education chatbot, it's the math teacher and English teacher. 15:36 So this minimized evaluation process, you want to make sure you have their input. 15:43 And another is a safety metric. 15:46 Like, define for your intervention how it can go wrong. 15:51 You know, in the case of medical chatbot prescribing a prescription 15:57 that somebody goes and takes that prescription and possibly has some adverse effects. 16:02 You can define the rate at which the chatbot is doing this as your safety metric. 16:09 And you can define this safety metric for other kinds of interventions as well. 16:15 Yes, the final part I want to talk about that kind of holds this in 16:22 is the evaluation we do before we deploy to one user is the offline evaluation. 16:28 You know, the grading of goldens, the safety metric, et cetera. 16:31 And now when we deploy to real users, we're still doing L1 evaluation. 16:37 We can still do model evaluation of it, which we'll demonstrate. 16:41 So in one case, you can use an LLM to rate your LLM's response. 16:48 You know, since you understand what it looks like, you can find these LLM as a judge criteria. 16:54 Like, hey, on a scale of one to five, how correct is this response? 16:57 And you can be tracking that for real users all the time. 17:00 And, you know, it's a little different. 17:04 Great. Now we'll move on to product evaluation. 17:07 And this is whether or not the overall product is actually engaging and retaining users. 17:11 And this is just to remind you, even if you built a perfect model, 17:14 if no one uses it, it will still have zero impact, right? 17:17 So I think using model evaluation in the first stage is extremely important to make sure it's accurate. 17:21 But it's also important to make sure that people are actually using whatever it is that you built. 17:26 And so we like to think of this as kind of one metric per stage within product evaluation. 17:30 There tend to be acquisition stages, activation, engagement, and also retention metrics that you want to consider. 17:38 So an example from one of our portfolio organizations, Digital Green, 17:42 they have a farmer chat, like a digital chatbot that you can interact with. 17:46 So their acquisition would be recruitment into the app. 17:48 This is where a farmer creates an account on the app. 17:50 This is something you log in your platform data. 17:53 Activation would be when the farmer logs in and asks the first question on the app. 17:57 So it's kind of like they've logged in, they can see what the value of the application is. 18:00 And then we want to think about engagement. 18:02 Like is this person logging in daily to ask questions? 18:05 This is something, you know, tech companies need to think about in terms of monthly active users, 18:09 daily active users, just how much people are actually engaging with and using the app. 18:13 Like how often are you on WhatsApp texting those people? 18:15 How often are you logging into Instagram looking at posts? 18:18 That's a form of engagement. 18:20 And then retention is, you know, whether or not people continue to do this over a longer period of time. 18:24 Right? 18:25 It's not just that they use it for one day or one week, but are they using it three months later, 18:29 six months later, 12 months later? 18:32 So the importance of having metrics for each of these stages is because it allows you to actually define things 18:38 from going from a gut feeling to being more data-driven. 18:41 So you can have something like, oh, well, I think users dropped off after onboarding, 18:45 or maybe the interface is confusing and that's why people aren't using it. 18:48 But if you have metrics, you can actually track this. 18:50 And I think as researchers and a lot of people in this room would understand, 18:53 it's important to have the data to actually evaluate why it is that people are not using the product. 18:57 And so, you know, user metrics can actually say, only 10% of users are retained on the app after day three. 19:03 What's happening in days one and two? 19:05 Or you can say, oh, the activation rate drops from 71% to 34% on low-connectivity devices. 19:10 Is that a segment that we should look into? 19:12 Maybe that's something we should think more about in terms of how we're deploying the app. 19:18 Now, I know that people here love experiments. 19:20 And, you know, the nice thing with product evaluation is they also love experimentation. 19:24 We just had to experiment on different things. 19:26 For example, product features, testing new features. 19:30 So, you know, in model evaluation, this will tend to be built in. 19:33 Your model, hopefully, should be improving over time. 19:35 So as you improve the model, you can roll out a lot of updates. 19:37 You'll see this with ChatGPT. 19:39 You'll see this with lots of version updates on these NetFlips or any of these type platforms. 19:43 You're often getting newer versions. 19:45 So they're always A-Testing these versions before they roll them out to everybody. 19:48 Another is just changes to user interface. 19:51 This is where you move buttons. 19:53 You make colors different. 19:54 You make things more appealing. 19:55 This is like if they change where the login button is. 19:57 It's just kind of changes to the interface that kind of make it easier. 20:01 You can also vary program activity, so this would be something like, I'll show you in a little bit, 20:06 where you can have, you already have messaging on the platform, sometimes you make it weekly, 20:10 or biography reminders, or you can build in behavioral nudges. 20:13 These are all things you can also implement as a part of an A/B test to try to increase engagement and retention. 20:17 And then there's basic things like bug fixes, like sometimes things just don't work, 20:21 so it's also important that you fix those because that will improve your product. 20:25 So one example of A/B testing is also from one of our organizations that we work with, Youth Impact. 20:31 There's actually a paper by Ingris that I have at the bottom, Ingris colon, in my gut, from 2025. 20:36 I'm not sure if I pronounce my name right, but what's really cool is they ran 12 A/B tests. 20:41 And what's also cool about A/B testing, just to let you know, is when you have outcome data, 20:45 you're getting continuous outcome data, right? 20:47 And so to the extent to which you can measure the outcome that you care about, 20:50 that's pretty much how rapidly you can experiment. 20:52 So here, they basically, at the end of each full term, collected learning outcomes. 20:56 And what they were looking at was iterative A/B testing of a phone-based tutoring service. 21:01 So in this phone-based tutoring service, they had a couple of different innovations. 21:04 They had tests that they were aimed at reducing costs for their intervention, 21:09 which is important for any organization. 21:11 And they also had a set of interventions or innovations that they believed 21:14 that could increase the efficacy of their tutoring. 21:16 So just to give you a sense, I won't go over this in too much detail, 21:19 but it's for dosage distribution. 21:21 You know, you basically have for A/B testing, version A is a status quo model. 21:25 This would be like a typical control group. 21:27 And then version B is the slight modification you're making to your program, right? 21:31 And so this is kind of a more rapid way of iterating on your intervention at a smaller scale 21:35 before you get to measuring potentially the main development outcome level for it, right? 21:39 It's a good way for you to kind of set up your experiment or your RCT in the future for success 21:43 and kind of be iterating on this. 21:45 And you can see they tested weekly calls versus bi-weekly. 21:48 They did same tutor versus different tutor. 21:50 They had different ways of scheduling. 21:52 They also included SMS messages or WhatsApp videos to help improve or boost student performance. 21:56 They had caregiver involvement nudges. 21:58 They sent videos on WhatsApp of, like, testimonials. 22:01 We also had tutoring session changes where they actually encouraged parents 22:05 to become more involved later on in the learning session. 22:07 But what's really cool here is you can see the variety of ways you can do A/B testing 22:11 and kind of iterate and build off each A/B test. 22:14 And then the third part is user evaluation. 22:18 This is very closely linked to product evaluation. 22:21 And often some of the changes that you make, even what I was saying before, 22:24 can lead to downstream outcomes on how users are thinking, feeling, and behaving. 22:28 And in particular, in a way that you want to be moving towards a development outcome. 22:32 We tend to think about level three outcomes as intermediate outcomes. 22:35 These are the outcomes that are necessary to achieve the development outcome of interest. 22:39 So, for example, if you care about infant mortality, 22:41 like an intermediate outcome could be mothers asking about where a clinic is 22:45 or mothers visiting a clinic, for example. 22:49 In particular, for intermediate outcomes, we want to know if a user is thinking, 22:53 feeling, or acting differently in ways that might lead to the development outcome of interest. 22:57 Some example dimensions that you might want to measure of change. 23:01 Cognitive. 23:02 This would be, like, for a learning application, 23:04 like comprehension, knowledge acquisition, complexity of reasoning. 23:08 You might also want to measure affective. 23:10 Like for health, for example, you know, do people feel more confident and empowered 23:13 and have a sense of agency? 23:14 Or are they feeling frustrated, right? 23:16 If somebody's feeling frustrated by health information, 23:18 or if a student's learning and they feel really frustrated and confused by the application, 23:21 that's likely to be detrimental to the development outcome. 23:23 So that's also something you want to track. 23:25 And then there's just behavioral outcomes you might want to look at as well. 23:28 So if you care about learning outcomes, you might want to look at shifts in platform behavior. 23:32 Like, are they downloading more tools? 23:34 Are they integrating the information that they've been learning into asking new questions? 23:38 Or are they asking for additional information in your resources? 23:41 So these are all kind of dimensions you can think about as potential community outcomes, 23:44 whatever it is that you want to measure mobile for. 23:48 So three ways that you can measure this. 23:50 What's really great is if you have application or platform data or chat data, 23:54 is there's multiple ways you can look at it. 23:56 So first is just log data. 23:57 If you have a platform, you know, look at the number of follow-up questions by user. 24:00 If you have a chat bot, for example, engage it with a specific set of features. 24:04 This could be like downloading tools or engaging more learning resources. 24:08 Then you can also use survey data, 24:10 which I feel like a lot of people here probably already have a lot of familiarity with. 24:14 You can do short in-chat surveys. 24:16 If it's a chat bot, you just embed it into the chat bot. 24:18 You can also do longer off-platform surveys where you actually go out into the field 24:21 and you interview users. 24:22 You can also have knowledge quizzes baked into whatever chat bot you might have 24:25 if it's a learning course. 24:27 And then you can also just analyze conversation logs and mind the text itself, 24:30 looking at sentiment analysis and topic modeling. 24:32 So these are all examples of things that you can potentially do 24:34 to engage with your media outcomes. 24:37 Also, I mean, we're at the agency fund, 24:39 and as I think there's been a lot of discussion, 24:41 user agency is also critical as a guardrail in this section. 24:44 Because, you know, there's a lot of discussion right now 24:46 about whether AI products are building capacity or creating dependency. 24:50 And especially in the spaces that we're working in, 24:51 in terms of healthcare or education, 24:53 this is something that we think is really important. 24:55 We tend to think about it in two ways. 24:57 One is subjective agency. 24:59 This is whether or not you're empowering users to believe 25:02 that they can solve their own problem now. 25:04 The second one is external agency, 25:06 which is can they actually plan to execute on the goal without the help of AI? 25:09 And this is kind of a metric for dependency as well, right? 25:12 It's like if you're giving a student a chat bot to help them with math 25:15 or to help them solve their math problems, 25:16 are they actually learning how to solve their math problems on their own? 25:19 Or is AI just substituting whatever knowledge they should be acquiring 25:22 within the course? 25:23 So these are kind of also a set of guardrail metrics 25:25 you want to consider in user evaluation 25:27 when you're deploying an AI product 25:29 to make sure that the AI product is actually improving and beneficial 25:32 as opposed to potentially creating dependencies. 25:36 And then finally, impact evaluation, 25:38 which I think most people here are already familiar with. 25:40 I'm not going to go that much into detail here, 25:42 but it is one that's very important. 25:45 As we know, a traditional randomized controlled trial 25:47 is an experiment measuring development outcomes. 25:49 At the end of the day, 25:50 this is the stuff we actually care the most about. 25:52 I think our emphasis here is just that there's three other levels 25:56 that are very important in shaping Level 4. 25:59 So if your model is not accurate, 26:01 it's probably not going to have a development impact outcome. 26:03 Or an impact on the development outcome. 26:06 If people aren't using your product, 26:08 it's also probably not likely to have an impact. 26:10 In user research, if people aren't using your tool 26:12 in the way that you believe that it is, 26:14 it's also not likely to have an impact. 26:16 So really, Levels 1, 2, and 3 are really important 26:18 just setting you up for success when you do run across a team. 26:23 Yeah, I think this is just to reiterate. 26:24 If your product is inaccurate, 26:25 if no one's using your product as intended, 26:27 there's unlikely there will be an impact. 26:29 So now, we may move to the more exciting section 26:31 of the technical demos. 26:41 Yeah, we're going to go through technical levels for Level 1 26:46 doing model evaluation. 26:47 You're going to see how it actually happens. 26:51 And then, we're going to show how to run an online A/B test. 26:57 Okay, so bear with me for a second 26:59 before we come back to this Q&A. 27:02 So here's the chatbot I talked about before. 27:04 When I say chatbot, again, 27:06 the agency kind of gives advice to non-profits. 27:08 So it's kind of equivalent to all the other devices. 27:10 It's a demo chatbot for that. 27:13 Alright, so when you log into this chatbot, 27:17 which I'm going to give you a link in a second, 27:19 what you all are going to do 27:21 is ask questions of this chatbot based on your expertise. 27:24 If you have questions like, 27:25 hey, how would I design an experiment for this intervention? 27:28 Ask questions of this chatbot. 27:30 If the response is not to your liking, 27:32 you're going to mark it for review 27:34 so that we add it to the L1 catalog. 27:37 Alright, so I'm going to demonstrate this test query. 27:40 Alright, so the question is, 27:43 for the fictional 2-leaf-leaf-light-dance pilot, 27:47 who is eligible and what transference approval 27:50 will participants receive? 27:52 If you ask this question on ChatWBT, 27:54 it's not going to know what this pilot is. 27:57 It's not going to give you the right answer. 28:00 So we're going to demonstrate how we can make 28:02 our chatbot give an informed answer to this. 28:05 Alright, so let's get started. 28:09 Alright, and real quick, 28:11 for simplicity purposes, 28:13 we've made the knowledge base a Google Doc, 28:15 you know, a flat Google Doc like this. 28:18 So this is what the chatbot looks to 28:21 to answer every question. 28:22 So, yeah, as you can imagine, 28:24 this is a complicated kind of domain 28:27 for the chatbot to speak over. 28:29 Alright, so we're going to ask the question. 28:38 . . . 28:54 Yeah, so per the knowledge base, 28:58 there's nothing in this knowledge base 29:00 about the light-dance pilot. 29:02 So it's saying, like, hey, 29:04 I did not see a 2-leaf-leaf-light-dance pilot 29:06 in the knowledge base provided, 29:07 so I can't give you eligibility rules 29:09 or a transport-disability test. 29:11 This is what we expect. 29:12 It doesn't have to be what this is. 29:14 Alright, so, yeah. 29:17 Now this is what you're going to do. 29:19 When it doesn't respond the way you think it should, 29:22 you've got to click this mark problem button here. 29:27 And first you're going to describe, 29:30 like, hey, you know, how is this wrong? 29:32 I kind of already have this prepared. 29:36 And it's like, hey, yeah, 29:38 it says the program information is missing, 29:41 so it can't answer the question. 29:44 Then, you know, based on, 29:46 here's a draft of what it actually answered. 29:48 You can either edit this or just, you know, get rid of it. 29:51 But you've got to put, like, 29:52 what the ideal answer should be. 29:53 It doesn't have to be perfect, 29:54 but just some sense of what the ideal answer should be. 29:59 So, I have this. 30:03 Oh, okay. 30:07 And, yeah, you know, 30:09 so I'm putting here the criteria of my ideal answer. 30:13 The response should say that participants are eligible 30:16 if they are between 18 and 24, 30:18 have been out of education, 30:19 and live in Alabama. 30:22 And, yeah, it should say the salary 30:24 and this is what it's mentioning. 30:26 So, then we're just going to hit mark problem. 30:33 All right, so we've added to, 30:36 when we mark this as a problem, 30:38 we want the chatbot to improve at this response. 30:41 So, here is a view that I'll be showing on the screen 30:46 of, you know, all the problems you all have marked. 30:50 And then, you know, at the end of the 20 minutes, 30:53 I'll give you to ask these questions. 30:55 We're going to demonstrate improving the chatbot live. 30:59 So, I'll do the first one here just so you can see. 31:10 Yeah, yeah, okay. 31:11 So, if you don't know the first one, 31:12 feel free to ask the question. 31:23 Yes. 31:37 Yeah, yeah, exactly. 31:38 No, it's going to be added to the Google Doc. 31:41 But, you know, we're doing this 31:43 so that we can capture the email manager 31:45 so we can read from that Google Doc. 31:54 Um, okay. 32:03 It'll show in a second so that we can add it to categories. 32:10 Okay. 32:22 Okay. 32:24 Okay, as disclosed, please scan this QR code. 32:30 Let me make this a bit bigger. 32:40 Yeah. 32:41 Scan this QR code, and I'll continue to demonstrate. 32:44 Now, feel free to mark a problem, 32:45 and to create a new, basically, or continue to add an answer, 32:48 and we'll explode it on the screen. 32:50 The internet's just a bit slow with everything we're doing, 32:52 so... 33:09 Okay. 33:39 Has everyone scanned the QR code? 33:45 Sorry? 33:46 Oh, you marked problems? 33:48 Yeah. 33:49 So, you will just talk to the bot, 33:51 and mark problems. 33:54 So, we'll demonstrate that here. 33:58 Yeah, ask the things that you know about. 34:00 You know, if you're an M&E expert, 34:02 like, ask the question, 34:04 like, hey, how would I design and maybe test for this information? 34:07 Ask the questions where you're an M&E expert. 34:09 You know, you all come from varied backgrounds, 34:11 with research and implementation. 34:14 So, ask the questions where you can judge 34:16 whether these questions differ or not. 34:23 Yeah, it can be anything. 34:24 And then, you know... 34:31 Yeah, and then you'll go through the mark-problem flow, 34:33 and I'll show you how we improve the chatbot, 34:35 you know, based on this... 34:45 The QR code? 34:46 Yeah. 34:59 Okay. 35:00 Let me walk through the mark-problem flow one more time. 35:04 And, yeah, let's ask a different question. 35:09 Let's say... 35:15 Where is Africa? 35:19 Evidence. 35:22 Summit. 35:24 Post. 35:26 2026. 35:27 Show all this. 35:28 You know, this is past the honest training date. 35:33 Okay. 35:59 And here's an example of, like, you know, 36:01 it's calling tools. 36:02 It's trying to search the web for that information. 36:04 So, this one actually... 36:11 Okay, but let's just say, hey, you know, 36:12 this should have been shorter. 36:14 This is an example. 36:16 Should have... 36:19 been... 36:21 shorter. 36:23 And then we can just, like, you know, 36:26 filter this to say... 36:32 Okay. 37:03 Yes. 37:04 So, while we're waiting for the answer, 37:06 any questions about the presentation so far? 37:10 Yeah. 37:12 My name is Lisa. 37:14 Earlier, you asked, 37:16 where was the 17th case, 37:18 which was doing some searching on the web? 37:21 Is it the case that this is on site, 37:23 the 17th, 37:24 17th day, 37:25 16th day, 37:26 17th day, 37:27 17th day, 37:28 and it's not? 37:29 Is it the case, 37:30 is it not? 37:31 So, I'm just trying to find a place 37:33 where this could be placed. 37:34 So, my name is Lisa. 37:35 So, my question is, 37:36 where was the 17th case, 37:37 which was doing some searching on the web? 37:39 Is it the case that this is on site, 37:40 the 17th day, 37:41 16th day, 37:42 17th day, 37:43 16th day, 37:44 and it's not? 37:46 So, since it's serving the Internet, 37:48 why, 37:49 does AI do the site anyway? 37:51 In addition to that, 37:53 is it a case where 37:54 some sites 37:55 or domains 37:56 don't allow AI 37:57 to come into them 37:58 and look through them? 37:59 Is that why 38:00 AI makes all of those things? 38:03 Yeah, that's a great question. 38:04 Should we take two more? 38:05 Yeah, sure. 38:06 Any other questions? 38:07 And yeah, 38:11 questions? 38:13 No? 38:14 Oh, yeah. 38:15 Great. 38:18 Okay, 38:19 thank you so much 38:20 for answering my question. 38:22 I have a question 38:24 about the jumbles. 38:25 Does it have a limit? 38:29 As in, 38:30 what limit? 38:31 Like text? 38:32 The size of text? 38:33 It's 38:34 considering the data. 38:35 Ah, yeah. 38:38 Anyone's questions? 38:40 Okay. 38:41 Yeah, yeah. 38:45 Thank you. 38:46 I'm Mofadi 38:47 from Madagascar 38:48 and 38:49 I would like to 38:50 ask some 38:51 tips from you 38:52 because 38:53 based on my 38:54 experience, 38:55 we have 38:56 conducted some 38:58 analysis 38:59 using 39:00 closed AI 39:01 or 39:02 CROC 39:03 and based on my experience, 39:04 I perceive that 39:05 sometimes 39:06 the AI 39:08 expects to 39:09 follow 39:10 my instructions 39:11 even if the instructions 39:12 are not right. 39:13 So, 39:14 can you suggest 39:15 that it's 39:16 recommended 39:17 to use 39:18 a longer route 39:19 over a short 39:20 and concise one? 39:21 Thank you. 39:24 Great question. 39:25 So, 39:26 let me start 39:27 in order. 39:32 Yeah, 39:33 let me start 39:34 with you actually. 39:36 So, 39:37 in AI, 39:38 there's a phenomenon 39:39 called 39:40 context rots, 39:42 meaning that 39:44 if you feed it 39:45 a bunch of instructions, 39:46 like a lot 39:47 of instructions, 39:48 it's going to pay 39:49 attention to the, 39:50 more attention 39:51 to the beginning 39:52 and then the end 39:53 and like ignore 39:54 everything in the middle. 39:55 So, 39:56 that's why it's really good 39:57 to have a short 39:58 approach. 39:59 You know, 40:00 if you have 40:00 It's not such a task that you feel like, you know, justifies such a big prompt. 40:04 It's a sign that you need to break it up into several separate prompts and stitch them together in what's called a workflow. 40:11 You know, it's a sign that you're trying to do two ambitious things in one prompt. 40:18 So, yeah, I would suggest breaking up your big prompt into several smaller prompts in order to get the AI to do what you say, essentially. 40:29 And then not ignore a whole swath of the question. 40:33 And on the limits of the chatbot, yeah, the chatbot does have some limits. 40:39 So there's what the foundational model companies have limits. 40:46 If you try to ask the chatbot, how do I make a bomb, it's not going to answer you. 40:50 You know, there's certain topics you have about it, especially medical advice. 40:56 So whichever you try to give medical advice, but it will always avoid telling you what to do or financial advice, you know, 41:03 because these things come with a big compliance risk, you know, become a big compliance risk. 41:09 I'm sorry, can you say that one more time? 41:13 I'm sorry, I actually forgot. No, it's okay. 41:16 I just asked you what you need. We asked both questions. 41:20 You just asked me what the subject is. 41:23 So, yeah, I just asked you what the subject is. 41:26 What I did was, it went back to what it started with and came back to. 41:29 Okay, yeah, definitely. 41:30 So by default, the model has a bias towards well-known sources. 41:40 If it sees, if it uses the web search tool and it sees that this is coming from the BBC.com, it's going to wait that long. 41:47 But, you know, if it's like a random site, it will often copy up that in the response. 41:53 And yes, it is true that today, some sites have what's called a robot-style text. 41:59 It's just a file just for elements. 42:02 There's an element-style text, but there's also an older robot-style text, 42:06 where you tell, like, hey, it's a more, like, the web is meant for humans to consume, 42:13 but these elements-style texts and robot-style texts are meant for AIs and machines to consume. 42:19 So some sites, like the New York Times and stuff, they add a robot-style text that says, 42:25 you know, you don't want to use my information in your response, 42:29 because they often have a separate deal with OpenAI and the environment, like, hey, 42:35 you need us so that our models can use your website. 42:39 So if you have a site that you don't want elements to just try to summarize at all, 42:44 you just need to add something called an elements-style text with your information on it, 42:48 and a robot-style text as well. 42:53 Yeah, so while you guys are asking questions, I'm going to demonstrate the evolve of the calendar. 43:00 So thanks, I see you guys, you know, adding questions in here. 43:07 Yeah, so this one is a good fun one. 43:09 Who won the World Cup? 43:11 He said, I don't have information about the World Cup in my knowledge base. 43:16 Yeah, my search tool is limited to websites of NGOs. 43:21 All right, so I'm going to add this to Calibrate. 43:24 This is Calibrate. 43:28 All right, so Calibrate is a platform we put together with Art Park. 43:35 And here's ours. 43:44 All right, so now let's run the evolve in this platform. 43:49 Calibrate is a platform that makes it easy to run these kind of model evaluations. 43:53 And we should expect to see a fail, right? 43:55 Because, yeah, it doesn't have information about the World Cup. 44:15 So let me explain what's happening here. 44:17 So we have our AI shopper. 44:19 Calibrate is using another AI to evaluate if our AI is correct. 44:26 It's a bit confusing, but just imagine Calibrate as kind of like an auditor, an auditor AI. 44:33 So you define what good looks like. 44:36 And, you know, this is what good looks like, staying on the World Cup. 44:43 And the Calibrate AI will run your AI on that same question and then see, okay, based on this, is this AI correct? 44:51 So you can see there's a section in here called reasoning that, you know, illustrates this state. 44:57 So the auditor AI said the response does not answer who won the World Cup. 45:03 And instead, it reflects the unrelated information. 45:05 So, you know, this is the reasoning output of the auditor AI. 45:10 All right, so let's go ahead and add this to our knowledge base. 45:17 So, you know, like we said, this is just a flag Google file. 45:21 So all it takes to update the knowledge base is going to be adding this here. 45:26 Okay. 45:34 Yeah, we should always manually update the AI ourselves. 45:41 What you're talking about is just like autonomously, the agent is autonomously improving itself. 45:47 That is an active area of research that, you know, Andropic is working on. 45:53 Currently, we always make humans make a change to the agent and then they remember. 46:01 Sorry for the text size. 46:06 So all we did here is just, you know, we just added the information to our knowledge base. 46:12 So now the chatbot can answer this question correctly. 46:23 Okay, now we should be able to answer the question. 46:27 So I'll rerun the eval and calibrate. 46:33 And we'll see the result. 46:42 All right, so now it's answering the question the way we want it. 46:47 It's correct now. 46:50 So as you can imagine, you know, when you're building a chatbot, you're going to build, 46:56 it's the work of the product team that's building the chatbot and the domain expert to build up this library of goals. 47:03 You know, we recommend at least 50 before you deploy to a real user. 47:08 But, you know, we have organizations like Digital Green that have over 100,000 goals. 47:14 Because this is so important to aligning the AI chatbot to behave the way you want. 47:19 You know, we can't just hope and pray. 47:21 Or, you know, some orgs, they have a spreadsheet and whenever they make a change, 47:26 they like ask each question in the spreadsheet. 47:29 A human is asking each question in the spreadsheet, which is just not sustainable. 47:33 You know, you can't do that. 47:35 So this is where the value of a platform that calibrates is useful. 47:41 Cool. 47:42 It looks like folks are adding email items. 47:45 So, yeah, let's go ahead and do a couple more. 47:48 But at this time, I want to lean on you to help me update the knowledge base. 47:52 So let me pick a, let's just survey these. 47:55 Propose an A/B test for a potential PhD candidate in finance who is trying to recruit a research partner. 48:04 This is, yeah, it's really long. 48:06 Let's try to find one with a row right next to it. 48:10 Okay, this one is good. 48:12 What econometric model do I use to tell me what I think my rank is? 48:17 Yeah, so it basically fails to answer. 48:20 So let's, yeah, this is good. 48:22 I think this is the right answer. 48:24 Yeah, okay. 48:25 It should be an object-defined model. 48:28 So we're going to go through that same loop, adding it to calibrates. 48:35 Please continue to ask questions. 48:43 So let's always just run it just to see it fail. 49:05 Can I ask a question? 49:14 Yes. 49:15 So this is quite laborious. 49:17 Yes. 49:18 I assume if you have 20 key questions, you can do this all automatically and run through all the questions. 49:28 Yeah. 49:29 And check the stochastic responses because sometimes the responses can vary. 49:32 Am I right? 49:33 Yeah. 49:34 You're exactly right. 49:35 So there's a lot of hard work in creating the goldens, but the benefit, the leverage really comes when you make changes, 49:44 you want to make changes to the chatbot because without this, right, if you make changes, you have to wait until the user complains that, 49:53 you know, this response is not what I expect. 49:56 And often the user will just stop using the chatbot. 49:59 They will not complain. 50:00 But with this, you can just rerun the e-mail and say, hey, it still behaves the way I expect, 50:05 even though I swapped out an expensive model for a cheaper model. 50:09 Yeah. 50:10 Yeah, you see what we see in practice with our partner organizations as well is they'll make a spreadsheet, 50:15 the goldens of like 50 or 100, and then they'll have the prompt, like, you know, what the question is, 50:20 and then the ideal response. 50:21 And then they'll actually, like, score their models. 50:23 So they'll run it through. 50:24 So, yeah, so we don't do it, like, line item by line item here. 50:26 This is just for demonstration purposes for you guys. 50:29 But in practice, what actually ends up happening is people make a spreadsheet, 50:32 and then they continuously rate the performance of their model. 50:35 And each model improvement that they make, they can score on things like accuracy, elucidation. 50:41 Some people have respectfulness, how encouraging things are. 50:44 So, you know, this is a very simplified version just so that you guys understand how it works. 50:48 But, yes, in theory, like, we actually have spreadsheets that tell you how it works. 50:52 It's not like I have an e-mail or anything. 50:54 All right. 50:55 So something interesting happened here. 50:57 So here was the original, you know, eval line item. 51:02 What the actual checkbox was not able to reach the server, 51:07 maybe because of a software bug or something, and they put it in the ideal answer. 51:11 But then when we ran it in Calibri for the first time, it was actually correct. 51:15 So here's what why, you know, here's a bit of the stochastic nature of this model. 51:23 So running it here allows you to just be able to run it over and over. 51:27 If the response runtime is not what you expect, because you built it in software 51:31 versus a manual spreadsheet that you have to go through over and over again, 51:35 you can just rerun it a couple of times to see, like, what the output is. 51:38 So because of that, let me give you another one. 51:41 I want to lean on somebody for updating the line. 51:47 This one looks right. 51:49 Okay, this one looks a good one. 51:51 Yeah, this one is great. 51:53 So I hope, you know, more of an app answer this, be prepared to help me put this in. 51:59 Okay, so this one is a great one. 52:01 I have multiple binary questions measuring gender attitudes. 52:05 I want to construct an index. 52:07 What are the best approaches available, and which one do you recommend? 52:10 Okay, this one is perfect. 52:12 All right, so please, can the person who added this… 52:16 Or gender experts in the room. 52:19 What is the right answer? 52:25 Let everybody come back here. 52:32 As a fallback, I would read my text and answer it. 52:37 Okay, yeah. 52:39 Oh, question? Okay. 52:48 My question is that before you are able to make a nice presentation, 52:53 I have one more question I would like to ask you. 52:56 How do you get the kids to use the source code? 53:00 Because that is the source code you already have. 53:02 And there is another question. 53:04 How do you get the company from Africa directly to Africa, for instance, 53:09 in terms of technology? 53:14 Yes, yes, definitely. 53:15 So to this. 53:17 We will make this source code public. 53:19 So just fill out the feedback form and put your email, 53:21 so that we can follow up with the open source code. 53:24 And yeah, adapt it into… 53:27 Let's say I'm Arabic. Let's say you want to adapt it to Amharic. 53:30 So things to think about. 53:33 If you just wrote knowledge-based Amharic, 53:37 the bot will perform worse than if you wrote the knowledge-based English. 53:42 So what you want to do is… 53:46 What you want to do is translate all user queries, 53:49 all user questions to your chatbot to English, 53:53 then run it through this chatbot, 53:55 and then the answer will be in English, 53:57 and then you translate that English to Amharic. 53:59 So that's the workaround we have to do right now, 54:01 because if you did everything in Amharic, 54:03 it won't perform as well as in English. 54:06 That's just another way I think we can work it out. 54:08 Why is that? 54:10 Because these models were trained on text-operated 54:14 and English… 54:16 We need to move on to the next section. 54:19 Okay, yeah, we need to move on to the next section. 54:25 You've seen the loop. 54:27 I would just add… 54:30 I don't even know the answer to this question. 54:33 Should we use PTA? 54:35 PTA? 54:37 Okay. 54:39 But I also have a follow-up question. 54:45 You can just say, 54:47 statistically, you can combine several biases. 54:50 You can say, 54:52 principle of analysis, 54:55 which is a method to combine several… 54:59 Yeah, okay. 55:00 Yeah, that's it. 55:01 Okay, so I'm not going to demonstrate the full loop, 55:03 but I hope you guys got this. 55:06 You know, you're adding it to this evaluation harness, 55:10 which just allows you to quickly run it over and over again. 55:13 And as you make changes to the model, 55:15 you can confidently ensure the behavior stays the same. 55:18 This is the crux of level one model evaluation. 55:23 All right, so… 55:25 Yeah, thanks for that, guys. 55:28 I want you to keep interacting with the chatbot, 55:31 but this time, half of you… 55:33 You guys have all been in an A/B test already, actually. 55:37 Half of you are being served by one model, 55:40 and the other half of you are being served by the other model. 55:44 Now, I want you guys to keep talking to the bot, 55:48 but this time, for every response in the conversation, 55:51 rate it as helpful or not helpful. 55:57 Yeah, so this one, I'm going to say it was helpful, 55:59 because the answer is correct. 56:02 I can ask another question, like, 56:05 who created the conference? 56:18 So if you guys just want to do the chat, 56:20 do thumbs up, thumbs down, 56:21 you can see the outcomes populate for the A/B test. 56:23 So just ask whatever questions you want. 56:26 You don't need to mark the problem this time. 56:27 Just enter thumbs up, thumbs down, 56:29 and then you'll see it pop up for the A/B test. 56:45 Okay, so as you guys are talking to the chatbot, 56:48 just keep rating it thumbs up, thumbs down. 56:50 This is what we'll use to understand 56:53 what variants in the A/B test is winning. 57:00 And in turn, we'll keep answering questions. 57:07 Sorry, so I have a question on calibrate. 57:10 Because there is no calibrate to check 57:13 or verify the answer, right? 57:16 What is calibrate based on? 57:19 And if you already have data base, 57:22 I guess there is some data that exists to respond, right? 57:29 So I'm just trying to understand why there is two pairs 57:32 that you're using. 57:34 And why does this make use of this API? 57:38 Okay, so a simpler way for us to do this, right, 57:44 is just have a spreadsheet. 57:46 Like we could just put our test queries in a spreadsheet, 57:49 put our goals in a spreadsheet. 57:51 But then someone has to copy over each... 57:59 Someone has to copy over each question, 58:02 ask the question, see if it looks right, 58:05 then come back to the spreadsheet like, 58:07 okay, this is correct, right? 58:11 This would be a manual way to do what calibrate is doing. 58:15 Just appoint somebody to go through the spreadsheet, 58:18 copy over the questions, does it look right? 58:20 Yeah, it looks right. 58:22 In software, we call it the right test. 58:26 So basically, this is just a test to make sure 58:28 that the model is behaving the way we want it to. 58:30 But like we said, this is too laborious. 58:32 Especially if you have up to hundreds of goldens 58:36 or in the case of Digital Green, 58:37 hundreds of thousands of goldens, 58:39 it is impossible for a human to do this 58:42 and it would be too expensive to have humans do this. 58:45 So calibrate is a platform to have AI do this for you. 58:49 So to have a second AI ask your questions, 58:52 check if it's right, and then it will update 58:55 whether it's passing or not passing. 58:57 Does that make sense? 58:59 Sorry, so is this calibrating, checking, 59:01 or verifying the response? 59:03 Yes, yes. 59:05 Or just the formulation? 59:07 No, no, it's checking the question, 59:10 whether it's correct or not. 59:13 Yeah, yeah, so that ideal answer that we had 59:17 in the Mark-Pramul flow is looking at the generated answer 59:21 to check if it's the same as the ideal answer. 59:24 What a human does, you know, if a human is like 59:26 pasting it in the chatbot, they're reading the response 59:28 to see like, okay, does this match 59:30 to what the ideal answer should be? 59:32 That's what calibrating AI is doing. 59:34 Okay, and you're using the AI because it's just so much 59:37 that you have to verify. 59:39 When you're building a production chatbot, 59:41 you can get to the tens of thousands, 59:43 hundreds of thousands of goldens 59:45 because you want to verify that the behavior 59:47 stays the same no matter what changes you make. 59:49 Because your architecture is not going to stay the same. 59:53 Models are improving, so when a cheaper, better model 59:56 comes out, I want to be able to release that model 59:58 while having confidence in the chatbot. 1:00:00 So every chatbot doesn't have to rely on some sort of algorithm? 1:00:30 Not as much as a human, but, you know, because you're using an LLM to check your LLM, so it doesn't incur some cost, but way cheaper than what it would be in a video room. 1:00:44 Yeah, A/B test results. So, yeah, let's take a look at the A/B tests. 1:01:00 All right, so the two models that half of you were exposed to was Claude Simon by Anthropic, and the other model is the open-source Chinese model, Kimi 2.6. 1:01:18 And based on your readings, Kimi 2.6 has more satisfactory responses than Anthropic's science. 1:01:26 You know, and it's probably harder in two of them, right, because we probably feel like the one that is closed-source is better, but here we've been able to... 1:01:35 Okay, so here's the caveat with this. To run an A/B test, we need appropriate sample size to say with confidence that this variant won't have it. 1:01:43 Because we're, what, like, you know, 40 people in the workshop, we don't have, like, the statistical power, but, you know, this... 1:01:50 If we're taking this direction, you know, like, this direction that Kimi 2.6 is probably better than Simon, you know. 1:01:57 This would also flip, like, if we double the sample size, for example, right, because we're only, what, 40 people in the room, this would also flip. 1:02:04 So, as you get a larger sample size, the results, like, we actually, if we're running this on an evidential, we have an A/B testing platform, and it's not even displaying the results because it says you don't have enough sample size. 1:02:14 So, these results could change because the sample size is relatively small, but it gives you a sense for the power of A/B testing and how rapidly you get results, right? 1:02:21 Because, basically, within this session, we've already shown you some data collection outcomes, and as the sample size grows larger, you can basically detect the more likely you're going to get that. 1:02:31 Yeah, exactly. And when we start thinking about, like, hey, what's the alternative to A/B testing? 1:02:39 Yeah, there really is none. It's a way to learn live on real users, you know, by running these controlled experiments. 1:02:47 Any questions on that? 1:02:52 So, what happens when we add some AI models to the model? 1:02:55 Say that again? 1:02:56 Where are some AI models, where are the models? 1:03:00 Yeah, good question. 1:03:02 Two factors go into how good a model is. 1:03:08 They're compute, so they need a kind of powerful computer called a GPU. 1:03:15 So, some models are trained on more computes than others. 1:03:20 So, if you're familiar with the Adapter models, they have Haiku, Sonnet, Opus, and now Fable. 1:03:27 They're increasingly trained on more and more computers, like millions of computers. 1:03:33 So, that's why they compute. 1:03:35 And then the data, the number of data, the amount of data they have access to. 1:03:39 So, right now, all the models, actually, they have access to all the same data. 1:03:46 But where they're different is this step called reinforcement learning with human feedback, 1:03:52 where they take experts in many different fields, in coding, in math, in biology, 1:03:57 and they get those experts to sit in front of the model and correct the model's responses on test questions, 1:04:03 and train the model to become better. 1:04:05 So, some companies, you know, hire more of these domain experts than others to make them better. 1:04:11 But, you know, there's a third factor called distillation. 1:04:15 Basically, it means that if you have a powerful model like Fable, 1:04:20 instead of investing in all that training data and human experts to help it, 1:04:25 you can just train your model on the other model, you know. 1:04:28 Like, ask the powerful model a bunch of questions, and then use that as training data for your weaker model. 1:04:35 So, that's actually what the community folks did. 1:04:38 They used Entropiq's model. 1:04:40 They distilled Entropiq's model in it. 1:04:43 Now, it's even better, at least in our experiment, than Entropiq's model. 1:04:47 So, those are the three main ways to do that. 1:04:51 Distillation, there's not a lot about distillation. 1:04:54 I guess it falls under copyright, but the model, because they're all trained with copyright data. 1:05:02 That's a truth. 1:05:03 So, there is no legal precedence that is illegal. 1:05:10 Yeah, it's probably wrong. 1:05:12 I don't know what that means. 1:05:14 All right. 1:05:16 So, yeah. 1:05:17 No, but rather than session, please scan this feedback QR code. 1:05:24 Put your email up. 1:05:26 You can follow up with more information in future workshops and, you know, the Open Source Conference. 1:05:30 Thank you. 1:05:36 My question is, is there a different kind of interest in different countries? 1:05:42 Do you have different interests? 1:05:45 Yeah, great question. 1:05:47 So, right now, AI is a private sector enterprise. 1:05:58 It's companies who are raising money to train models. 1:06:01 So, most of those companies are in America and now increasingly in China. 1:06:08 But there is more interest in sovereign AI. 1:06:14 Like, countries don't want to rely on these private companies to be the only ones that create the AI models. 1:06:21 You know, Nigeria just created, like, a fund to create their own sovereign AI. 1:06:27 So, not right now. 1:06:29 Right now, countries don't create their own AI models. 1:06:32 But there is more interest in this going forward. 1:06:35 Countries want to start creating their own models so they don't have to rely on these big companies to be the ones that create the models. 1:06:44 That's a good question. 1:06:50 Okay. 1:06:51 So, how can we be confident in using the operators? 1:06:58 You ask questions, and you will get some answers. 1:07:03 You make that to personal use, I mean, humanized. 1:07:08 To what extent that is humanized? 1:07:13 Evidence is, you are competent for personal use. 1:07:18 As if you are the one associating with it. 1:07:25 Okay. 1:07:26 So, if I understood your question right, as you're customizing these models to both be accurate and then also, like, empathetic, you know, 1:07:39 to respond in a way that, you know, seems human or, you know, can be digested by humans, yeah, what gives you that confidence? 1:07:47 Because an AI, reviewing another AI's output, there is bias. 1:07:54 There's bias there. 1:07:56 This has been shown by research that when an AI is presented with an AI-generated content and a human-generated content, it reads the AI-generated content as more favorable, as more competent. 1:08:10 And, you know, in some cases that might be true, but that doesn't reveal the inherent bias that AIs have for AI-generated responses versus human responses. 1:08:23 So, if you want to put in the effort to make, so, to make the AI content more manageable by humans, it is up to you, you know, when you create these goals, to make sure the ideal answers are the flavor of, like, writing that you want. 1:08:41 Because otherwise, you know, the bias is towards AI-generated content. 1:08:45 And this is why it's good to not, like, to not use AI to generate the ideal answer. 1:08:51 Use the domain expert. 1:08:52 Don't, because, yeah, it can be easy to say, like, you know, just generate with AI what the ideal response should be, but then, obviously, you're putting into that bias that AI has for each other. 1:09:04 This is where you should have the domain expert, like, be the one that creates it from scratch. 1:09:10 Don't even, like, show them a draft. 1:09:12 Have the domain expert create the ideal answer from scratch without looking, you know, from the other eyes. 1:09:22 Thank you. 1:09:23 Just a quick one. 1:09:25 You sent some of us to Cloud, Sonic, and the other group. 1:09:32 When I ask the same questions directly on Cloud, they give responses. 1:09:39 But when we use Latin, some of those questions will not get responses. 1:09:47 So, I was just curious why Latin is not using directly. 1:09:52 Yeah, right. 1:09:57 So, this question was that, you know, when he asked the same question to Cloud.AI versus Cloud.AI actually answering the question, whereas Latin, he was going to give the wrong answer. 1:10:13 So, for the purpose of the workshop, we made it all the answers from Google, just to make it easy. 1:10:19 But what you bring up is that the financial model labs, they search the internet, you know, they search the internet and try to make a response. 1:10:28 So, you know, they build their answer that tries as best as possible to answer the question. 1:10:34 But for educational purposes here, we just made it all from Google Maps, even though it theoretically, you know, both of us are featuring knowledge or whatever. 1:10:44 So, if I'm using Latin, let's say, for practical purposes, it can be borrowed directly from, let's say, AI, you know, from the Cloud. 1:10:56 Yeah, Latin is not meant to be production ready yet. 1:10:59 You know, it's just like a demo chatbot we built for the purposes of this workshop. 1:11:04 But, yeah, if we want to make Latin more production ready, we would make it, you know, both from the internet and try as best as possible to answer the question. 1:11:18 Yeah, this is an example of what I said, so that they're posting a lot of this in the general internet. 1:11:25 Thank you so much. 1:11:26 My name is Joseph. 1:11:27 I want to go back to the presentation itself. 1:11:29 Thanks for the presentation. 1:11:31 I think with this kind of work and all the other initiatives, we're excited to be able to do it. 1:11:36 Usually, when we want to use the data, we're excited to explore it and to be able to perform it. 1:11:40 But my question is, how do we establish causality at the end of the day? 1:11:45 You know, in cases where, if, for example, the lack of data confidence increases, then we take it back to the data-driven product. 1:11:54 And then, when we talked about the user nodes, we added the organization nodes and implemented data source. 1:12:01 And I was wondering how we balance that with ethics and privacy in using them. 1:12:07 And then, finally, in organizations, we really have established the change. 1:12:12 I was wondering how we integrate this so that it's not separate, but we can make some of that really sustainable for us. 1:12:19 Excellent questions. 1:12:21 Okay. 1:12:22 So the first one on causality, which I would say the best, like the gold standard for user evaluation is similar to what it is for level four, 1:12:29 which is also if you build in some A/B tests. 1:12:31 Right? 1:12:32 So the example I showed in the English paper, they basically, you know, they're looking at engagement, obviously, 1:12:36 but they were looking at, you know, kind of intermediate learning outcomes, like using tests. 1:12:40 Right? 1:12:41 So I think there is something where the ideal for level three is you're building in measurement of intermediate outcomes into an A/B test that you do already. 1:12:48 That being said, similar to other constraints, sometimes that's not possible. 1:12:51 Right? 1:12:52 So having some user data is better than not having any user data at all. 1:12:55 Some of the organizations we work with, we encourage them to do surveys. 1:12:58 We encourage them to do qualitative research condition because you want to have some sort of user research. 1:13:02 I used to work at a large tech company called Meta, and they have quantitative researchers, qualitative researchers, 1:13:08 and they use both platform data where they just do observational data analysis, but they also do, like, regular A/B testing. 1:13:15 Right? 1:13:16 So ideally for causality, you have some sort of randomization or experimentation where it's possible, but sometimes it's not possible. 1:13:22 But causality is a great question, also important for level three. 1:13:25 You also asked about user privacy. 1:13:27 This is something that is actually very important. 1:13:29 I think for products, usually, you know, when you sign up to Entropiq, when you sign up to Chats with BT, they have a user privacy policy. 1:13:37 So if you're working with an organization, a lot of organizations you work with, like health organizations, they also, in theory, like, 1:13:43 we do actually, you know, for anonymization and data and things like that, user privacy is very important. 1:13:48 Like, usually, they don't share any of their anonymized data with us, to the extent to which you can mask AI. 1:13:53 I mean, it depends as a researcher in an organization. 1:13:55 As an organization, that's the point of products. 1:13:57 You should be very careful about user privacy in some ways. 1:13:59 Right? 1:14:00 So this is something that I think whenever you're building AI products, you do have to think more like a tech company or a startup, in a sense, 1:14:05 about ensuring the privacy of your users, your personal data on users, who you're going to share it with. 1:14:10 I mean, we all work with IRB human subjects, very similar to researchers. 1:14:13 So on the tech product side, it would be a user privacy policy. 1:14:16 On the researcher side, it would be like IRB. 1:14:18 Right? 1:14:19 So user privacy is very important here. 1:14:20 You should not be sharing user data with everybody. 1:14:22 And then third, the theory of change is a very good question. 1:14:25 In some ways, I think the way the four-level framework is set up is very much integrated into a theory of change, because level four is a development outcome. 1:14:32 And usually, you have some sort of theory of change about how you're going to achieve your development outcome. 1:14:37 And the intermediate outcome should really fall from your theory of change. 1:14:40 So whatever you think is going to change in the child health outcome, whatever you think is going to change in education outcomes, 1:14:45 there should be a part of your theory of change that is saying, what is the intermediate outcome to reducing infant mortality? 1:14:51 What is the intermediate outcome to increasing student learning? 1:14:53 And how can I track that? 1:14:54 And that will be very specific to each organization. 1:14:56 But that's usually what we try to do with our partner orgs, is we work with the theory of change. 1:15:00 If the outcome is business revenue, then the intermediate outcome might be, oh, they've actually learned the business training course, for example. 1:15:07 So I think it's up to each organization to determine that intermediate outcome. 1:15:10 And also the job of researchers like you to work with organizations to increase these kinds of things. 1:15:14 Does that answer all of your questions? 1:15:16 Great. 1:15:17 And everyone got a chance to scan the feedback code. 1:15:20 Also, we would love to run more workshops like this. 1:15:22 And so for us, it would be really helpful to have feedback, whether it's useful or not useful. 1:15:26 And we're also going to sign up for slides as well, for whoever's interested. 1:15:30 So any other questions? 1:15:32 I know we're out of time, and people really want to get to lunch. 1:15:34 But any additional questions before we wrap up? 1:15:36 Yeah. 1:15:39 Any last questions? 1:15:40 Go ahead. 1:15:43 Thank you. 1:15:44 I think that was an excellent session. 1:15:46 All this seems quite expensive to go through all these different stages. 1:15:53 Would you have any recommendations for organizations working on a budget and capacity constraints? 1:16:00 I can speak to levels three and four, maybe more, and then I can speak more to the first two. 1:16:06 I mean, I think for impact valuation, these are definitely very expensive. 1:16:09 And I think the point is more to try to build this in as much as possible. 1:16:12 Usually people are already doing piloting for RCTs. 1:16:14 So as a part of your pilot for your RCTs, just do more user research specific to products, right? 1:16:19 Try to collect how many users you can. 1:16:21 So that's my recommendation for level three. 1:16:23 If you're already doing an RCT, try to look into piloting and try to look at product data. 1:16:27 Like with our partner organizations, a lot of them are doing products. 1:16:30 And I would say for each question that you have in RCT, think, 1:16:33 is there a way I can measure this in the platform data or in the chat data that I have? 1:16:36 And that's, you know, not to mention it costs additional money, 1:16:39 but it's something that just requires extra thinking before you deploy a level four RCT. 1:16:43 I think for level two, I mean, A/B testing, yes, requires some tech support. 1:16:48 And then model over model testing costs are also quite expensive. 1:16:54 Yeah, for level one, the model evaluation. 1:16:58 Yeah, and it's costly in terms of you need some expertise, 1:17:03 like someone who can set up this model evaluation pipeline. 1:17:07 But if you're using an off-the-shelf platform like Calibrate, 1:17:11 you can actually bring your own API key so that, you know, 1:17:14 it's not incurring any cost other than, you know, the open API token cost, 1:17:18 which is very small. 1:17:20 Like for, you know, one organization, it was like $20,000. 1:17:24 And to run it once, it was like $50. 1:17:27 So, you know, I imagine, like, what it would cost in human time to do something like that. 1:17:33 You know, it probably costs like tens of thousands of dollars, 1:17:36 but $50 in comparison is good. 1:17:39 And also, like, you don't need to use the frontier models like OpenAI and Science. 1:17:45 When you have this model evaluation pipeline, 1:17:49 one of the first things it allows you to do is switch to a cheaper open-source model, you know, 1:17:56 so that you can save more cost overall. 1:17:59 But I would suggest that if you're scrapping, 1:18:02 go with an off-the-shelf platform like Calibrate where it takes care of a lot of this volume 1:18:07 and allows you to, you're not locked into Calibrate, you know, 1:18:10 it allows you to export results in a spreadsheet or something like that. 1:18:14 And A/B testing. 1:18:16 Yeah, A/B testing, if you're small, I'd actually say, like, just put that down the road 1:18:21 because you also need power. 1:18:23 You know, you need statistical power and thousands of users to even, like, 1:18:27 get meaningful results from A/B tests. 1:18:29 But if you're a smaller organization with a budget constraint, 1:18:33 you probably have the most leverage from offline model evaluations that I've shown you here. 1:18:40 Great. Thanks so much, guys, for all of your interaction and presentation. 1:18:43 Thank you. 1:18:44 Thanks, guys. 1:18:56 All right. 1:18:57 Thank you, both Kelly and Edmund.