Three Structural Patterns Across Eleven AIM Cost-Effectiveness Analyses

CEAmethodologyeffective altruismresearchAIMcost-effectiveness

I read eleven cost-effectiveness analyses published by Ambitious Impact (formerly Charity Entrepreneurship) over the last few weeks. Seven were global health and development; four were animal welfare. The earliest was from 2022, the rest from 2024. What I was looking for was not whether any individual CEA had the right answer, but whether the same conceptual parameters were being treated consistently across analyses. They mostly weren’t, and three of the inconsistencies seem worth flagging in public.

The short version: AIM’s “probability of success” parameter is constructed three completely different ways across the corpus, with values ranging from 0.2 to 1.0 for what is supposed to be the same conceptual quantity. The 2024 GHD template’s internal and external validity adjustments are applied inconsistently from one CEA to the next, including one where they’re explicitly zeroed out. And template “suggested defaults” are often left at their default values rather than customized, most strikingly in Digital Pulmonary Rehabilitation 2024, which draws on far more template parameters than the other CEAs and ends up with 30 of 55 suggested values left at defaults across the live model. None of these are individual errors. They’re patterns about how the template gets used in practice, and I think they might be useful to the people who maintain it.

Some context first

This post is in the spirit of recent Forum work by Vasco Grilo and a few others on AIM’s methodology, but it’s doing something different. Vasco’s posts typically take one AIM CEA, re-derive its bottom-line cost-effectiveness number under different assumptions, and report the new number. Mine looks across eleven CEAs at the level of how parameters get constructed in the first place, and asks whether the same conceptual parameter is treated consistently from one analysis to the next. I think both kinds of work are useful, and the second one is hard to do without a corpus.

I’m an independent researcher with no affiliation to AIM. This work was unpaid and I have no commercial relationship with anyone named in the post. Most of my recent work has been an analogous methodology review of GiveWell’s CEAs.

What I looked at

Intervention Year Track File
Maternity / Kangaroo Care 2022 GHD 2022_LSGH_-_275_CEA_MF.xlsx
Kangaroo Care (replication) 2024 GHD 2024_-_Replication_-__25_Kangaroo_care__KC__-_S5_CEA.xlsx
Differentiated Learning (TaRL) 2024 GHD 2024_-_I_G_-__71_-_S5_CEA_Differential_Learning.xlsx
ULAB Regulation 2024 GHD 2024-_I_G_-__16_ULAB_regulation_-_S5_Cost_Effectiveness_Analysis__CEA_.xlsx
Digital Pulmonary Rehab 2024 GHD _94_Digital_Pulmonary_Rehabilitation_2024_GHD_-_CEA.xlsx
Education Information 2024 GHD 2024_SDG___CEA____48___Education_information.xlsx
Lesson Plans n/d GHD Lesson_plans_CEA.xlsx
Keel Bone Fractures 2024 AW 2024_AW_-__31_Interventions_to_reduce_KBF_-_S5_CEA.xlsx
Cage-Free / Broiler Welfare 2024 AW 2024_AW_-__45_Cage-free_and_broiler_welfare_campaigns_in_neglected_countries_-_S5_CEA.xlsx
Fish Welfare n/d AW Fish_welfare_-_CEA.xlsx
Alternative Protein Scale-up n/d AW _242_-_CEA_-_Securing_Scale-up_Funding_for_the_Alternative_Protein_Industry.xlsx

A note on framing before the findings. My goal here is to identify patterns in AIM’s methodology, not to evaluate the researchers who wrote any individual CEA. Where I name a specific CEA below, it’s because it’s the clearest illustration of a pattern that recurs elsewhere in the corpus. I tried to be precise enough about cell references that anyone reading this can verify it in a few minutes.

Finding 1: probability of success is constructed three different ways

Every CEA in the sample has a parameter labeled some variant of “probability of success.” It’s usually a single scalar that multiplies the expected impact of the intervention to account for the chance the charity fails to deliver. It is a high-leverage parameter. In the time-series rollups it scales every benefit estimate in every year, so the final cost-effectiveness number moves roughly proportionally to whatever value gets chosen.

Across the eleven CEAs I looked at, the same parameter is constructed in three structurally distinct ways. There’s no documented reason for the variation.

The most common approach is to pick a number based on the researcher’s best guess. KC 2022 (Main CEA!B34) uses 0.65. ULAB 2024 (Main model!B116) uses 0.2. DL 2024 (Main CEA!B98) uses 1.0, which the template note suggests should be used “for direct” interventions. KC 2024 (Main CEA!B105) uses 0.7. The Cage-Free 2024 CEA’s “Approach 1” uses 0.8 across all three country sheets (UAE, Saudi Arabia, Egypt at row 21). Lesson Plans uses 0.8 in both its Rwanda and SSA tabs. Fish Welfare uses 0.5.

The second approach is to leave the value at 0.9, which appears to be carried over from the 2023 template default. Five of the eleven CEAs preserve an “OLD 2023 CEA” sheet, and every one of those sheets uses 0.9. More notably, Digital Pulmonary Rehab 2024 uses 0.9 in every one of its five live country and model sheets: Nepal myCOPD partnership, Thailand myCOPD partnership, India myCOPD partnership, Nepal self-develop (video supervision), and the old in-person PR CEA. All at row 11. The PulmRehab CEA structurally adopts AIM’s 2024 template but seems to have inherited 0.9 from the 2023 default without modification.

The third approach is to construct the parameter as a literal AVERAGE() over heterogeneous external reference points. Two CEAs do this. KBF 2024 uses the same formula across all four of its sheets (KAT Germany, KAT Germany - 35%, RSPCA UK at row 21; UEP USA at row 23):

=AVERAGE(5/27, 7/10, 10/28, 3/10, 3/28, 2/5, 2/8, 2/10, 0.1, 0.15, 0.1) ≈ 0.26

The fractions look like a combination of Aquatic Life Institute’s certifier success rates from multiple counting methodologies, Open Philanthropy’s grant success rates, and three hardcoded guesses (0.1, 0.15, 0.1). Cage-Free 2024 Approach 2 uses =AVERAGE(29%, 30%, 15%) ≈ 0.247 at Cage-free (UAE)!B26, identical in the Saudi Arabia and Egypt sheets.

So the variance across the corpus runs from 0.2 (ULAB) to 1.0 (DL), with the averaged-reference-class CEAs landing around 0.25, the single-point judgments clustering between 0.5 and 0.8, and PulmRehab carrying 0.9 across all five of its model sheets. I cannot find a documented basis for the differences in any of the CEAs themselves.

I don’t think any of these methods is wrong in isolation. The pattern that stands out is that the same conceptually identical parameter is being constructed by completely different processes across the corpus, with no shared definition of what reference class is appropriate or when judgment should override the template default. The template recently started flagging this parameter with the suggestion [suggest: 90% for direct, 25% for ...], which is a step in the right direction, but the suggestion alone doesn’t seem to be producing convergence.

A useful direction here might be a short methodology note, even just a paragraph in the template, that defines what “probability of success” is meant to represent, what reference classes are appropriate for what intervention types, and when researchers should use a single judgment versus an averaged set. That alone would make the parameter’s role in the final cost-effectiveness number much more comparable across CEAs.

Finding 2: validity adjustments are applied inconsistently

The 2024 template introduces two parameters, “internal validity effect” and “external validity effect,” that adjust every effect-size estimate to account for the gap between the studies the estimate is drawn from and the conditions of the proposed intervention. The template suggests ranges of 0% to -80% for internal validity and +40% to -80% for external validity.

The four 2024 GHD CEAs that use these parameters apply them very differently. KC 2024 uses -0.2 internal and +0.3 external (Main CEA!B51 and B52), and the same values appear again in B59/B60 for the morbidity calculation. DL 2024 uses -0.3 and -0.1 (Main CEA!B56 and B57). ULAB 2024 sets both to 0.0 (Main model!B74 and B75), explicitly zeroing the adjustments out. PulmRehab 2024 customizes them per country and per estimate, ranging from -0.1 to -0.2 for internal and -0.1 to -0.2 for external across its various model sheets.

The animal welfare CEAs in the sample (KBF, Cage-Free) don’t have validity adjustments at all. Lesson Plans uses a single combined External validity adjustment = 0.8 instead of the two-parameter formulation.

The thing I notice here is less about any individual choice than about the absence of a shared norm. KC 2024 leaves both adjustments at values that match the template suggestions. ULAB zeroes them out. DL customizes both. PulmRehab customizes them but keeps them in similar ranges across its model sheets. There doesn’t seem to be guidance on when zeroing the adjustment is appropriate (versus, say, when a study is so closely matched to the proposed intervention that no adjustment is needed), and there’s no record of the researcher’s reasoning about why their chosen values are correct.

This is closely related to Finding 1. It’s another place where the template introduced a parameter without enforcing how to use it. It also matters, because these adjustments are multiplicative on the effect size, so the difference between -0.3 / -0.1 (DL) and 0.0 / 0.0 (ULAB) is a meaningful swing in the final number.

A useful direction here might be a shared decision rule or short worked example in the template. Something like “if your effect size is from a single RCT in a different country, use roughly X; if it’s from a meta-analysis covering similar contexts, use roughly Y.” Even a rough rubric would make the parameter’s use traceable.

Finding 3: template defaults are often left unchanged

AIM’s 2024 GHD template embeds suggested values for many parameters using the convention [Suggest: 130000] or [suggest: 90% for direct, 25% for ...]. These are clearly meant as starting points for the researcher to customize, not as final values.

Across the four 2024 GHD CEAs in the sample, the rate at which these suggestions get left unchanged varies dramatically:

CEA Suggested defaults left unchanged Total
ULAB 2024 2 10
KC 2024 8 14
DL 2024 8 14
PulmRehab 2024 30 55

The pattern is most pronounced in the largest CEAs. PulmRehab uses substantially more parameters from the template (55, against a typical 10–14 for the others), and the rate at which those parameters get left at suggested defaults is comparable to KC and DL — meaning the absolute count of unchanged defaults is dramatically larger. The unchanged defaults in PulmRehab include several core parameters: probability of success at 0.9 (which connects this finding to Finding 1 from a different angle), “Last year charity runs if unsuccessful” at 4, “Delay on impact starting from intervention date” at 0, and “Total variable costs, NPV, if not successful” at 0.

The pattern I want to draw out is this. The template appears to be working too well in one sense, in that it gives researchers structure and reasonable defaults to work from. And not well enough in another, in that it doesn’t currently surface to readers of the published CEA which parameters were actively customized versus which were left at defaults. From a reader’s perspective, every value in a published CEA carries the same epistemic weight, even though some represent considered judgments and others are template fill-in-the-blank. A researcher who actively customizes a parameter is making a different kind of epistemic claim than one who leaves the default in place. That distinction is invisible in the final spreadsheet.

There are a few directions this could go. One would be a visual convention in the template, perhaps a cell color or a comment marker, that distinguishes “researcher-customized” from “template default unchanged.” Another would be a reviewer checklist that asks “which template defaults were intentionally retained, and why?” before publication. A more ambitious version would be a sensitivity audit that flags any final cost-effectiveness number heavily dependent on unchanged defaults.

A few smaller things I noticed

Two observations that didn’t make it into the main findings but seem worth flagging.

First, the sample includes both pre-template CEAs (2022 KC, the animal welfare CEAs) and templated 2024 GHD CEAs, and the template adoption is uneven. The 2024 animal welfare CEAs in the sample use a different structure than the templated 2024 GHD CEAs. No Key sheet, different parameter naming, different validity-adjustment treatment. I’m not sure whether this is because the animal welfare team is on a different template version, because the template is GHD-specific by design, or for some other reason. Worth a question rather than a finding.

Second, by the way: the 2024 template’s Main CEA sheet has columns labeled “Reviewer notes” and “Replies” in the Key sheet documentation, but in every templated CEA I looked at, those columns (G and H) are repurposed to hold years 5 and 6 of the time-series rollup. The reviewer-dialogue infrastructure exists in the template but isn’t being used in any of the published versions I read. I assume reviews happen elsewhere (in comments, separate documents, internal Slack), but it might be worth knowing that this part of the template has effectively been deprecated through use.

What I’d value hearing back

A few things, in roughly the order I care about them.

For AIM staff who happen to read this: the most useful thing I could learn is whether any of these patterns resonate with concerns you’ve already noticed internally. I’d rather know I’m surfacing things that are already on your radar (which would tell me where the real unsolved problems probably live) than tell you things you already know. I’d also welcome correction on any place where I’ve misread the spreadsheets or missed context that would change the picture.

For other Forum readers working on or thinking about CEA methodology: I’d be interested to know whether the patterns I’ve described here (especially the probability-of-success construction and the template-defaults question) are things you’ve seen in other organizations’ work, or whether they look distinctive to AIM. The version of this question I find most interesting is whether the cross-CEA parameter consistency problem is a general feature of evaluator methodologies that try to compare across cause areas, or whether some orgs have solved it.

For Vasco Grilo and others doing related Forum work: thank you for the existing methodology dialogue. I read several of your posts while preparing this and found them very useful for understanding how AIM’s CEAs are received outside the org.

On the offer of a deeper analysis

The findings in this post come from manual inspection of the spreadsheets. For a closer look at any individual CEA, I have a multi-agent critique pipeline I’ve been developing for this kind of work, originally built against GiveWell’s CEAs. Beyond the kinds of structural patterns I’ve surfaced manually, the pipeline can do a few things a manual pass can’t.

It can quantify how individual parameter choices affect the bottom line. For instance, it can compute how much PulmRehab’s final cost-effectiveness number would shift if its probability of success were set to ULAB’s value, or to the KBF averaged-reference-class value. The same analysis runs on every parameter that drives the result, ranked by leverage.

It can do cell-by-cell uncertainty propagation. Rather than treating each parameter as a point estimate, it constructs plausible ranges and propagates them through the model to produce a distribution over the final cost-effectiveness number. This often reveals that the published point estimate sits in the tail of the distribution, or that two seemingly small parameter uncertainties compound into a much larger uncertainty in the result.

It can run automated cross-CEA parameter consistency checks. Running the same parameter against multiple CEAs at once produces a structured comparison that’s harder to do by eye across many files.

And it tends to surface structural critiques the manual pass missed. The pipeline typically produces 10–30 findings per CEA, not all of which would be visible from inspection alone.

If anyone at AIM, or anywhere else with public CEAs in a similar shape, would find a deeper run on a specific CEA useful, I’d be glad to do that work. The API costs run about $15-30 per intervention, which I’d cover. I’d particularly value AIM’s own judgment on which of their CEAs would benefit most, since you know which ones are most consequential to current decisions and which are already in line for revision.

— Todd aka Tsondo [email protected] tsondo.com