Welcome to GiveWell’s podcast sharing the latest updates on our work. Tune in for conversations with GiveWell staff members discussing current priorities of our Research team and recent developments in the global health landscape.
Elie Hassenfeld: [00:00:00] Hey, everyone. This is Elie Hassenfeld, GiveWell's co-founder and CEO.
Today I'm talking to Alex Cohen, a Program Director at GiveWell, who leads our Cross-Cutting team. The Cross-Cutting team in our research team is responsible for assessing our methods and ensuring that we continue to improve as a research organization. It asks questions like, what's really happening in the programs we support? What do we know? What are our biggest open questions, and how can we resolve them? And this work that this team has led has been part of GiveWell improving the quality of our research substantially over the last few years.
You know, something that I often feel is, people wonder whether GiveWell has been able to maintain the quality of its research as we've grown. And in fact, we've, in my opinion, substantially improved the rigor, the quality, the level of analysis we're able to do as we've grown the amount of funding that we're [00:01:00] directing. This continued focus on the quality of the research that we're doing, and this commitment to our core values of truth-seeking and transparency is even more important now as we are preparing for increased growth in the future.
The Cross-Cutting team focuses on improving the quality of our research through a few different mechanisms. First, they're focused on recruiting more people to the team. One of the best mechanisms through which we can improve the quality of what we do is having more people at GiveWell who are asking critical questions and able to find answers. We've successfully grown the research team significantly over the last few years, and that's been a big part of what has led to the improvements in the work that we're delivering. We are hiring very actively now, and so if you are interested in what we're talking about, please do look at our jobs page. If you know people who might be a good fit for what we're looking for, please share the job postings with them. And this Cross-Cutting team is where many researchers start when they first [00:02:00] come to GiveWell, because it's a great training ground for all of the work that we do.
So one way in which we improve quality is through recruiting. A second way is by pressure-testing grants and gathering data that gives us more information about what's really happening on the ground. GiveWell's reached a size and a scale where we can support our own data collection to answer the questions that we have. And this data collection that we've supported on our own has helped us learn more and make better decisions about grantmaking.
And then finally, the team is focused on using artificial intelligence and its tools to make our work more efficient, and to allow us to gather and analyze more data than we've been able to in the past. We talked about an example of this in a recent episode focused on looking back at a malaria net distribution in Democratic Republic of Congo, but we are also using AI tools to do things like look at the spreadsheets that we use to create cost effectiveness estimates and find errors. And so, this team as a whole is [00:03:00] responsible for ensuring that the quality of our research and grantmaking is on an upwards trajectory by doing all of these different types of work to make sure we keep getting better.
So Alex, thanks so much for joining me today and having this conversation with me. Can you introduce yourself and then just share some of the big picture priorities that Cross-Cutting is working on that you're most excited to talk about?
Alex Cohen: Yeah, thanks Elie. My name's Alex Cohen. I'm a Program Director at GiveWell. I lead our Cross-Cutting team. Broadly, I'd say the goal of Cross-Cutting team is, look for ways to maintain and improve the quality of our research over time. Three things that we've been focused on toward that are: first, looking for ways to check our work against ground truth data. So we funded a net campaign or a seasonal malaria chemoprevention campaign. Did those programs get delivered, and did they reduce mortality rates or prevalence? So trying to collect that data and make sure that it's high quality.
[00:04:00] Second piece is looking for ways to use AI in our work either to speed up what we do, or make our work better—look for errors. And then third piece I've been focused on is recruiting and hiring. As we look to scale up our grantmaking, we want to make sure we're adding more great researchers and program officers to the team, and I've been leading the charge on that.
Elie Hassenfeld: Yeah, so let's just dive in. Let's start with checking our work with more ground truth data. Tell me a little bit about what you've been doing to work on that priority.
Alex Cohen: Yeah, so typically, GiveWell, when we fund a program, we collect data to measure whether the program got delivered. But we did a reanalysis of the data that we've collected to see, is this data reliable? Is it high quality? This was prompted partially by something we talked about on a previous podcast, which is this finding with Dispensers for Safe Water, where we got some data from the grantee that we [00:05:00] funded showing much higher chlorination rates than an independent check. And that was one of the things that prompted us to go back and look at the data that we collect on program coverage and other monitoring indicators to see, are we collecting data that we should actually trust?
And, you know, I think this sort of thing is really important because we're making these grant decisions. We're estimating cost effectiveness. We're saying we think these programs are gonna be impactful, but I think one way we test that is expose it to reality. Say, you know, when we go and follow up with people, are they actually using these things? Making sure that those data are collected well seems fundamentally important.
Elie Hassenfeld: And so something that's been true about the groups we recommend for a long time is, they collect data at a high level relative to most organizations working in development. We review that data. That plays a big role in our thinking. And the improvement we're making here is that we're investing more in additional independent checks [00:06:00] and collection of different kinds of data that the organizations themselves aren't collecting. And so maybe we could talk through a couple examples of what the potential issues are with that data that comes directly from the organizations themselves, what questions we're left with, and then what we're doing to try to fill in the gaps and get a better sense of what's really happening.
Alex Cohen: Yeah. Yeah, that's great. I'll give a couple examples here.
Start with one about seasonal malaria chemoprevention programs. So, seasonal malaria chemoprevention, or SMC, is a program that provides anti-malarial medications to children under five years living in parts of Africa where malaria is seasonal.Typically, this is provided over several months, but in a three-day regimen. And the first day, there's a distributor that gives the SMC medication to a child, makes sure they consume it. But then they send the next two days' doses home with their caregivers. The idea is over the next couple of days, the caregiver, mom or [00:07:00] dad, gives them their SMC doses. And obviously, key question we want to know is, are those doses getting delivered? Typically, we've relied on surveys, where people go door-to-door to these communities and ask, "Did you give those follow-up doses?"
Intuitively, we can think of reasons why this might be overestimated. Kids may refuse to take the second and third doses. This medicine does not taste great—I've actually tasted it, I can attest to that. Sometimes there's vomiting because the taste is so bad. I personally didn't experience that, but this is commonly reported among kids. You also hear stories that caregivers might think, "Well, actually, let me save this medication for when my kid gets malaria." It's intended to be preventative, but parents may think, "Let me save it until we really need it." There are those kind of plausible reasons the uptake could be lower and caregivers may report it being higher due to social desirability. They wanna tell surveyors, "Yes, I did the [00:08:00] thing that they asked me to do."
Elie Hassenfeld: And so in the scheme of things, this is like a pretty straightforward program. It's the organization goes door to door, delivers this medicine, but then there are doses that need to be taken over the course of several days. And you're describing the reasons, like even in a program like this one, where kids just need to take some medicine, there are ways in which this program can fail to have the expected impact.
That's why the organization itself is going back and monitoring and trying to determine whether the doses took place. But, yeah, it's not even clear that that information would necessarily be accurate. So yeah, what would lead you to question the organization's survey information? You know, in the case of Dispensers for Safe Water and the chlorination data that you mentioned—and if anyone wants to learn more, you can see that in a previous conversation that we had here—you know, that was hard-to-interpret data, and surveyors may have been selecting individuals to survey through non-random methods. Like, what are the reasons [00:09:00] that you'd be less likely to trust the organization's data here?
Alex Cohen: Yeah. So I think in this particular case, the concern is, we're asking caregivers to report whether they gave their child SMC doses over the following two days after the first one that the distributor observes. And there's social desirability bias in the sense of caregivers may be inclined to say they gave the kid the medication, even though there may be reasons why they might not.
The other maybe more compelling thing that's making us raise this question is when we were going back and digging deep on the monitoring evaluation data we collected, we found a study in Niger that has this interesting test where they compare caregiver self-reports, so the data that we typically collect for SMC campaigns, to a more objective measure from dried blood spots. So this involves taking a sample of the child's blood to see whether they have the markers of the SMC drug that would be consistent with [00:10:00] getting these second and third day's doses. And they found a pretty big gap. Something like 70% of kids were reported to have taken the second and third doses due to self-report, whereas the blood test showed less than 20%. So this gap between 70% and 20%. Of course, this is one study, but given the kind of qualitative reasons why we think this could be the case. This is prompting us to do some of these dried blood spot testing ourselves. So we decided to fund a study in Burkina Faso to measure this alongside campaigns we're funding.
Elie Hassenfeld: And that study in Niger, that was not from a program we funded. That was another…was that from another study that just prompted this question for us?
Alex Cohen: Yeah, that's right. This wasn't one of our programs. This was a study that we found in the literature online.
Elie Hassenfeld: What do you expect to come from this work we're supporting in Burkina Faso? How could it influence our grantmaking?
Alex Cohen: Yeah. So I think one way is we get their survey results, we get the dried blood spot [00:11:00] test, we can compare them to caregiver self-report, and they align pretty closely. In which case, yeah, we don't do anything differently. I think if we were to find the same sort of gap that we found in this study in Niger, that might prompt us to say, "Okay, what can we do to boost the uptake of SMC on day two and day three?" If we're seeing that the dried blood spot tests show much lower coverage, is there additional outreach that Malaria Consortium can do to encourage the uptake of the second and third doses? It kind of prompts a lot of like, “Can we improve the program?” type questions.
Elie Hassenfeld: And when do we expect to get those results back?
Alex Cohen: I believe the study wraps the end of 2027. So, yeah, this is a long-term study. We're exploring adding these dried blood spot tests in a couple other areas too, so possible we get some results before then. But that's the plan for this Burkina study.
Elie Hassenfeld: I mean, I think this is such an interesting case [00:12:00] because when you think about, like, if you zoom out and think about charitable giving as a whole, a program like delivering seasonal malaria chemoprevention is one of the more straightforward, evidence-based, relatively simple programs that exist. And I guess it just helps illustrate some of where we're coming from in our work at GiveWell that we're like, there are still really important ways in which the program could fail to have impact. And then we want to find those possibilities and gather the data that can help us understand what's happening, so we can either be more confident in the direction we're taking, or find ways to improve the program, or I guess in the most extreme outcomes, decide to exit the program. But even when you start from this, like, very strong case for impact, and that's pretty unique, there are still ways in which the program can fail to have impact, and that's what drives this work to get more data.
Alex Cohen: Yeah. Yeah, I think that's right. Even for a fairly straightforward program, things are complicated.
Elie Hassenfeld: Yeah, things are complicated. Do you have another example [00:13:00] of this ground truthing that you want to talk through?
Alex Cohen: Yeah. Another example that came up when we were doing this deep dive into our monitoring and evaluation comes from malnutrition treatment. So we funded a lot of malnutrition treatment programs. This involves providing ready-to-use therapeutic food and other interventions to children that have severe acute malnutrition, SAM, or moderate acute malnutrition, MAM. Typical way that we have tested for whether kids are actually receiving this treatment is to do a survey of households, to go to caregivers and say, "Do you have children here who are acutely malnourished?" Or they do assessments to see if there are children that are malnourished, and then asking the caregivers, "Did you receive ready-to-use therapeutic food? Did you receive other malnutrition treatment?" And this was a story where we found that the coverage estimates were maybe too low, relative to the truth.
So what happened in this case is that the organization that we [00:14:00] funded, ALIMA, went through and cross-checked what they saw from survey data of households with their actual clinical records of who came in and received malnutrition treatment. And they found that there were a fair number of kids whose caregivers said they did not receive malnutrition treatment, but they were at least logged in the records as receiving treatment. So in this case, it was something like the surveys suggested 20% of kids that were malnourished received treatment. When we add in these kids that were in the clinic records but counted as not receiving treatment in the surveys, it goes up to 30%.
We're not totally sure what's going on here. One story that we've heard kind of anecdotally is households are concerned that if they say they've received malnutrition treatment, they may not be eligible in the future. You know, they've already received the treatment, they're not gonna be able to get it again. But we kind of learned from this that it's important to do that sort of cross-reference between the survey data, and in this case the clinic data, because there are reasons that the [00:15:00] caregiver self-reports in this case may be underestimating coverage.
Elie Hassenfeld: And it's interesting because normally when I think about these additional data collection exercises, I do tend to think about it from the pessimistic perspective. Like, what do we need to learn about what might be happening where the programs we support fail to have impact? And of course, you can also learn that the data we're relying on is an underestimate. We do a lot of work in our modeling where I mean, we're trying to take our best guesses based on the data we have about the number of people who'll be served, the amount of leakage in a program, et cetera. And the more that we can home in on an accurate estimate, the better decision-making we're able to make with the funds that we're responsible for. And so better data collection can tell us something's gone wrong with the program, but it can also tell us the program is serving more people per dollar than we expect. And then that could cause us to direct more funding into that program than we were [00:16:00] previously.
Alex Cohen: Yeah, I think that's right. I think at GiveWell we all kind of skew more toward the looking for problems. Yeah, I think that's important. But, yeah, sometimes when you cross-check these data, the picture is more favorable than we thought originally.
Elie Hassenfeld: And so in this case, I mean, I know we're funding more work, and we don't have the results yet of what we believe to be true about the proportion of children who are served. But I mean, a 50% increase in the kids served sounds like this really meaningful update on the cost-effectiveness of the program. If you're serving half as many children again per dollar, then that should have a big effect on the bottom line. Is that the right way of thinking about it? Like, is it that material of an update?
Alex Cohen: If we're updating coverage by 50%, so going from 20% to 30%, yeah, that does mean 50% more kids got treated. In this case, this was a relatively small pilot that ALIMA did to do this sort of cross-reference. So we're talking about something like 60 additional [00:17:00] kids, which, you know, 60 additional kids that we were missing. But this is still small-scale data, something that we'd want to double-check.
But I think maybe, for me the bigger update is these data sources that we're relying on are fragile in many ways. They've got errors. That's why it's good to triangulate against multiple sources. And in this case, looking at the household surveys and the clinic data on people that actually came in, I think that's a broader takeaway for us because I think the same thing I imagine could apply to vaccination programs and others, and it means that we're…we may be way off on coverage. But don't want to anchor too much on this one because this is still small scale.
Elie Hassenfeld: Right. And so maybe like 50% as a magnitude seems unlikely. You know, if the magnitude were that large, you would know, like, how much of the basic food that you were providing. You'd have to buy a lot more of that food, so presumably you'd have another way of knowing. But it's an illustration of how fragile some of the data sources are and, like you said, the [00:18:00] benefits of triangulation to really hone in on what's true.
Other examples you want to share about ways in which we're ground truthing more to help us better understand reality and make better decisions?
Alex Cohen: Yeah, I guess the other one here is we're trying to collect more data, so asking people in Nigeria that run survey firms or people that run these programs to collect more information. We're trying to get out to the field a little bit ourselves too. And as we've been doing this, we've seen examples or maybe corroboration for having questions about the quality of some of the surveys that we're getting.
So one other thing that we noticed when we were digging into the monitoring and evaluation data we collect, is that in a lot of cases, surveyors are really pressed for time. They've got to do a lot of surveys per day. You could see how there could be incentives to skip households or to prioritize the households that are closest to the main road or easiest to get to you know, for very natural [00:19:00] reasons like, “I've only got a certain amount of time,” or yeah, “It's really difficult to get to these harder to reach places.” That's something anecdotally we've seen, we've heard it from other people, but we've also seen it ourselves.
So, last year I went to Uganda to follow surveyors that were doing post-distribution surveys for bed net campaigns. So going to households and asking, "Did you receive a bed net? Are you still using it? Can we see it hanging?" I thought the surveyors were very hardworking, seemed committed to getting the right answers. But at the same time, it was like a 10-hour day, no stops for lunch or for drinks of water. These people are working hard. And then a couple of my colleagues, Mark Walsh and Steven Brownstone, were in Abuja recently following surveyors, this time in more urban settings, but saw something similar where they might miss households that were in alleyways and places that were physically strenuous to get to. Again, very natural things. But kind of, again, first-hand corroboration of [00:20:00] how there are challenges in doing these surveys, and we should be scrutinizing the results, looking at other sources. Also considering ways…should we be paying more for surveys so that surveyors are able to take the extra day? Obviously that takes more money, because you have to pay people for an extra day of their time, but is that worth it to get more reliable results?
Yeah, that prompted a lot of questions for us.
Elie Hassenfeld: And I'm sure a lot of people listening to this will make this connection, but it's a real problem if surveyors are skipping the hardest to reach places, because it's also possible that the people distributing the goods are missing the hardest to reach places. And so if the people doing the distribution, and then the surveyors are missing the hardest to reach places, all of a sudden you get data back that says a higher proportion of people that you're trying to reach are being reached because the group of people that you think could have been considered was off the table.
And in many cases, I think it's often the hardest to reach people who have the greatest needs. It's not always the case, but that is not an unreasonable starting [00:21:00] point. And so, therefore, finding ways to ensure that the monitoring is assessing as broad a scope as possible of the folks who we were trying to reach, that's really critical.
Alex Cohen: No, that's right. And I…you know, we keep saying, reality is complicated, so maybe this is trite at this point. But even doing a survey, I think when I'm sitting at my desk in the US, I'm thinking, "Okay, do a survey of people to see if they got nets. How hard could that be? You just go, you knock on the door, and you ask the question." But of course, when you're there, like when you're seeing a program delivered, it's complicated. You have to get the household listing, you have to make sure to find all the households. You have to fit it all in within one day. These things are complicated, which means there's chances to compromise quality.
Elie Hassenfeld: In some cases, it's fairly intuitive how the independent monitoring improves the outcome relative to the organization-led monitoring. So especially in cases where the data is somewhat subjective and the surveyor needs to evaluate the data. [00:22:00] If they're from the organization, they might have an instinct to lean, more towards the organization's…like the data that's better for the organization than a view that could go either way. But when it's more challenging to reach certain places, how are we solving the problem that, at the end of the day, the surveyors, they're going to make decisions that, without the right incentives or the right oversight, are easier for them. And so just, why should we expect that the independent monitoring is better than the organizational monitoring in this respect?
Alex Cohen: Yeah. So I think that the independence of the monitoring is maybe one characteristic of ideal surveys, so you don't have the incentive to report more favorable results. But it's a necessary but not sufficient condition. I think independence is one thing. We also want to make sure that surveyors have enough time. We want to make sure that there are back checks. I think this is surveyors that follow up to a random subset of households to see [00:23:00] whether they were visited. See when they ask the same questions, do they get the same results? There are things like GPS tracking to see, okay, did this person fill out their surveys, like 30 surveys in this one spot, or can we see them moving along geographically to the areas that we expect them to go to?
So there are checks like this. None of it is foolproof. You could ask, well, who's back checking the back checkers? But part of this process, we developed a set of survey guidelines that include some of these features. We're talking to our grantees now about these to understand what's feasible and what's not. But, yeah, definitely independence is not the cure-all for this.
Elie Hassenfeld: It's interesting because you'd think that in many cases, the grantees themselves should also be interested, or are interested, presumably, in higher quality evaluation of whether their programs are working. And I'm wondering if there are cases where they take on board the ideas that we're using with our independent monitoring to support their own programmatic monitoring.
I mean, the idea would be, [00:24:00] when it's independent monitoring that GiveWell is funding, we can require whatever we wanna require in order...you know, we make the grant to the survey group that does the things that we want. And I should say they are a research group, so they're positioned to take on requests for gathering data in a different way.
But then I would also imagine there should be cases where the organizations themselves want to do better, and in many ways that's a more leveraged outcome, because then they're getting the data sooner, they've built better internal mechanisms. And so how do you think about supporting better evaluation within organizations as part of this, do you think any are taking that on themselves?
Alex Cohen: Yeah. Yes, a lot of the organizations that we fund do want this data because they want to know, okay, were there these villages where people weren't getting nets, and do we need to go back to those and figure out what's going on there? And I think, as we're developing these survey guidelines to say these are the characteristics that we wanna see, we want to hear from the organizations that we fund, not only do you think [00:25:00] this is feasible, but would this be useful to you, and how would you use this information?
I think the other part of it too is a lot of these things may cost more money. If you're doing more back checks, you have to pay more people to do those back checks. If you're doing the GPS tracking, that requires the digital devices to do that. Or if we want surveyors to do fewer per day so that there's less of a risk that they skip places, that means you have to hire more surveyors or more days per surveyor. And I think if we're an organization that's focused on cost effectiveness, it can be easy for a grantee to say, "Well, we'll just cut that off the budget because our cost effectiveness will look better." And yeah, we want to have those conversations with grantees and hear from them. I think in a lot of places, it might be worth the extra cost to have this better data.
Elie Hassenfeld: Yeah. It seems like over time, the ideal outcome is to push more of the evaluation to the grantees, to the extent they can take it on and it makes sense. And I [00:26:00] think we believe that better data helps better decision-making, and it would be great to have that inculcated among grantees, not just in context where GiveWell is the funder, but also in other contexts too.
Obviously, there are other challenges. You know, we're probably one of the few groups that says to groups doing research, "You know, we want to pay more to get better outcomes. We want you to ask for more money." So there's a challenge in organizations being able to find the funding they need, but does seem like the right direction to head.
Alex Cohen: Yeah. Yeah, I agree.
Elie Hassenfeld: So we talked a little bit about ground truthing data vis-a-vis programs. And then what else is on your mind where we're trying to do a better job getting these local insights and improving our own decision-making?
Alex Cohen: Yeah, another area that comes to mind where we're trying to get more information from people in countries where we're funding programs is on our moral weights. So we use our moral weights to quantify, how valuable is it to [00:27:00] double consumption for a year, double the amount of income that people have relative to averting a death or improving health? This is obviously a really challenging thing to estimate. We do it because we're considering funding programs that affect different outcomes. Some aim to raise incomes and make people richer. Some aim to provide programs that prevent diseases. And so we need a way to try to compare those.
We've tried a few different ways to get at this. One we've done more recently is asking individuals in low- and middle-income countries hypothetical questions like, “How much would you be willing to pay for a pill that lowered a child's chance of dying by 10%?” The idea is to try to elicit this trade-off, how much money are people willing to give up to lower mortality risk?
We think there are a lot of benefits to that sort of hypothetical question. But lately we've been trying to get more data on revealed preferences. So, instead of asking people to imagine a scenario, we look at cases [00:28:00] where they're making a decision to invest in something or buy something that actually does lower the risk of dying in real life.
So in the US, a common way to do this is to look at wage risk studies. So look at differences in wages between jobs based on their likelihood of dying on the job. An ice road trucker in Alaska makes a lot more because there's a risk of dying on the job. There's been a lot less of that sort of study in low- and middle-income countries. And so we decided to fund more. We opened a request for proposals to get more studies on this. These look at things like willingness to pay for a motorcycle helmet. Motorcycle helmets are a way to reduce risk of dying using motorcycles, which are very common in several low- and middle-income countries. And by looking at how much people are willing to pay for these, we can get at that trade-off income and health in a more high stakes, real world way.
Elie Hassenfeld: One of the things that has always struck me when, [00:29:00] I don't know, I've thought about this or spoken to people when I've traveled is—and this is very anecdotal—but it seems like the willingness to pay for prevention is much lower than the willingness to pay for treatment. So someone seems much less likely to say, spend the money needed and take the time required to go get preventative malaria medication. But in the event that their child is sick, is willing to do a lot to go and get medication for their child. I guess I have a broad question, which is like how does that fit into our thinking and your thinking about the way to set these moral weights and the kind of revealed preference data we could gather that would help us make better decisions?
Alex Cohen: Yeah, I think it's a good question. I think when we're thinking about the right trade-off between income and health, there's a question of when we survey people on this or use these types of revealed preferences studies, [00:30:00] what do they say versus how much stock should we put in that?
An example here is an initial version of this helmet study, which found something like a value of a statistical life of something like $500, which is much lower than we assume, obviously much lower than the value of statistical life that the United States EPA uses, for example. There are lots of reasons why that might be the case. One story is, people are poor, and so they would rather have the $500. But it does seem surprisingly low, and I think it could come from issues like that, people aren't fully grokking the trade-off that they're making here. And a helmet is like a preventive medicine in some cases. You know, it's lowering your chances, but it's not like you're sick and you're getting the trade-off for the treatment itself.
So you know, we'll use these as one point of triangulation for our moral weights. I don't think it's like we should take the results at face value, more we're adding this to the other [00:31:00] data points and benchmarks that we have for this.
Elie Hassenfeld: Are we funding any revealed preference in surveys that focus on the treatment side rather than the prevention side? Because I think that's what I'm really curious about, because in the treatment side, the condition is, like 100% likely to be present. Where in prevention, you don't know whether what the result is showing is the value of a statistical life or, I don't know, like treating a low probability of something bad as approaching zero.
Alex Cohen: Yeah. So for these studies, they are in this more prevention bucket. That's a good question, though. If we were to do something with treatment, how would that look? But yeah, a lot take the form of, choosing a less risky mode of transportation, or choosing a pesticide that differs in its mortality risk. And suffers from this, “can people interpret these probabilities correctly” type of a critique. So yeah, the treatment side would be interesting to explore to see if we could do something along those lines.
Elie Hassenfeld: I wonder if there's [00:32:00] any academic evidence from the US, I don't know, like looking at the ratio between the two. Just a thought that I'm curious about, because I'm always struck by the fact that we know there's this finding that, I mean, when GiveWell got started, I think the dominant malaria net program was asking people to pay for malaria nets. That was also part of deworming programs in the early days. And the consistent finding was that even a small fee massively reduced uptake. And on the other hand, whenever I've talked to people about their experiences with healthcare, like you know, last summer in Malawi, the summer before in Kenya, the stories they tell are going to just extraordinary lengths in trying to raise extraordinary sums to purchase medical treatment. And I mean, I guess it's possible that those two truths could square mathematically. Like, I haven't tried to do that, but I suspect that we would get a different finding looking at the treatment side versus the prevention side of the equation. [00:33:00]
Alex Cohen: Yeah. It definitely seems right.
Elie Hassenfeld: I'm glad we're doing more on this. I mean, we've just done so much on moral weights over time to think about them. Obviously, it's an area where we're not going to get true answers, you know, because there really isn't one. But the more that we're able to do, the more we have both to inform our own decision-making, but also help donors think this through on their own to the extent they're looking for help.
So we've talked a lot about the work the Cross-Cutting team is doing, and you're leading, to get better data on the ground. And then I know another part of what you're focused on is using AI tools more effectively in our work to help GiveWell do more, do it more efficiently, and do it better. Maybe talk us through a couple of the things on that side too.
Alex Cohen: Yeah. So, there's definitely a lot in the using AI to analyze data. So, quantitative data that we're getting from surveys, or qualitative data from interviews. But maybe the piece that would be more interesting to talk about is using AI to critique our work.
So we've done a couple things in this regard. One is using [00:34:00] AI to vet our spreadsheets and cost-effectiveness estimates. And we in the past have had a team of research analysts on the team that are vetting and looking for errors in calculations, looking for conceptual mistakes. And we thought this was a great use case for AI. Especially now that these models, especially the agentic ones, can go through spreadsheets, they can reason, we can give them a bunch of context, we can show them past work. Could they do some of this vetting work?
We developed an approach to do this using Claude Code, and we tested it out, and then we gave it a number of vets that human reviewers had done and compared how Claude Code did against those. This was kind of a holdout sample, this wasn't part of Claude's training. And we got to the point where it was catching 100% of the errors that the human vetters were catching. So we started using this as one of the main ways that we vet our spreadsheet [00:35:00] calculations. We've still got a human in the loop overseeing these as we go, doing some spot checks. But this is a spot where we've been able to offload quite a bit of our research critiquing to AI models.
Elie Hassenfeld: Yeah, I mean, how much have we learned from that? How good are they at critiquing our models?
Alex Cohen: For things like spreadsheet vetting where it is a little bit more mechanical, they do pretty well. The other use case here is we've been doing what we call, AI red-teaming. So looking at the full case for a grant, not just the cost-effectiveness spreadsheet, but the whole rationale for why we're making this grant, and asking it to critique it like a manager at GiveWell would. There, I think we've had a little bit less success. Definitely not none. The issue with those is the models will spit out 30 critiques or so, and a couple of them will be useful, will point out things that we've missed that seem important, which is good. [00:36:00] It just definitely requires more human effort to go through that 30 and find the two diamonds in the rough that are useful. We're still trying to iterate and get better at that, but as we kind of widen the type of review that we're asking it to do, it's maybe a little bit less useful.
Elie Hassenfeld: Yeah. Cool. Well, thanks so much, Alex. Really appreciated this conversation. Anything you want to add before we wrap?
Alex Cohen: Just one last thing. At the top I mentioned that another area that I've been working on is hiring. And just wanted to put a plug that we are looking for researchers, program officers to join the team. We've got a lot of roles available. My role on that is to help identify, screen applicants, and then new hires typically join Cross-Cutting team to do onboarding, learn how to do GiveWell research. But please share our job pages with folks you know that might be interested.
Elie Hassenfeld: Cool. Well, thanks so much for doing this.
Alex Cohen: Thanks, Elie.
--
Elie Hassenfeld: [00:37:00] Hey, everyone, it's Elie again. I hope you enjoyed that conversation.
I think it's really amazing to think about all of the ways in which we're focusing on gathering more data and using new tools to make better decisions in our grantmaking. One of the things that's been on my mind, and I've talked about it a little bit with folks internally, is the extent to which, as we embark on a time when GiveWell may grow substantially, whether we're trying to maintain the quality of our work or continue to improve it. And I think this is an interesting question because everything that I see shows me that we have continued to improve the rigor that we bring to the decision we're making, the quality of the data that we can rely on, and the level of analysis that goes into the grants that we're making.
And so even though today, you know, we expect in 2026 to direct significantly more funding than we ever had in the past, I think those funds will be directed with a greater degree of rigor and more clarity into the [00:38:00] impacts that they're having than at any point in the past. And that's very exciting to me because I think it sets the right foundation for the growth that we are embarking on now.
So thank you as always for your interest in what we do and for listening. We really appreciate it.