Choosing Your Evaluation Design
Whether you're commissioning the evaluation or running it, the same path looks different depending on where you start. Good design means choosing the route the terrain allows.
By Ben Jaques-Leslie
Did the program work?
This is the question at the start of almost every evaluation. But it's worth asking what answering it is actually for: is it to confirm the program worked so a box gets checked, or to produce something specific enough that someone — a funder deciding whether to renew, or a team deciding whether to continue— can use it to actually decide something?
That purpose shapes the harder questions-behind-the-question for evaluators: What would count as an answer? What would it take to get there? And given the realities of this program, in this place, what's genuinely possible? Put simply: what kind of evidence can we actually build here?
The answer is not just a methodological footnote. The kind of evidence you can build shapes everything that follows — what you can claim, what you have to qualify, and whether your findings hold up when scrutinized. And yet design decisions often get made quickly, under pressure, and with less deliberation than they warrant — particularly when the context is complicated. Get it wrong, and even a well-executed evaluation can produce findings too weak to support the decisions they're meant to inform.
Here’s what that looked like in one real evaluation. The choices we made, from the evaluator’s side, including where we had to adapt mid-course.
Design is not a checklist
There is a tendency to treat evaluation design as a hierarchy. Randomized controlled trials (RCT) sit at the top. Everything else is treated as a compromise. That assumption is both narrow and can make evaluators passive. Designs long treated as second-best can offer insights that RCTs can't, and treating randomization as the presumptive gold standard can itself limit what an evaluation is able to ask.
Something different is needed: an approach that takes evaluation conditions seriously without treating them as constraints. That means being creative about what the available evidence base can support, expansive about what questions a design might answer, realistic about what is actually feasible, and honest — particularly honest — about what any given design cannot tell you.
When the program is already over
In 2026, Apricity was commissioned to evaluate the long-term sustainability of a nutrition programming initiative in Nepal. The program had concluded over a year earlier. Program participants had been selected, implementation was finished, and there was no possibility of randomizing anything.
The straightforward response would have been a survey: find the people who participated, ask them what they remember and still practice, and report the results. This approach is common, and in some circumstances it is reasonable. But it cannot tell you whether any sustained practices are attributable to the program, and it misses an important question entirely: did what program participants learn spread to anyone else?
We designed the evaluation differently. Rather than surveying program participants alone, we built a quasi-experimental comparison across three groups: direct program participants, non-participant community members in program areas, and residents of matched communities with no exposure to the initiative. This allowed us to estimate not just whether program participants maintained new practices, but whether those practices had diffused into the broader community — including to people the program never directly reached.
Building diffusion into the design costs relatively little. Leaving it out would have meant leaving one of the funder’s most important questions unanswered.
Going further on methods — and being honest when plans change
Having committed to a comparison design, the question became how to make it as credible as possible. That required being willing to learn.
Being expansive in evaluation design is not only about what questions you ask. It is also about what methods you are willing to try. This evaluation called for matching, which is essentially a way of constructing a fair comparison when you can't randomly assign who gets a program. You find communities or individuals that look as similar as possible on the factors you can measure, then compare outcomes. Neither the design nor the matching approach was something that had been used in this exact form before. That was not a reason to avoid it. Methods can be learned and building that capacity is part of doing the work seriously.
The original plan was to use a principled matching approach to directly optimize balance between matched groups. In practice, however, the structure of the data — imbalances between populations we sought to match, and geographic constraints — made this approach unworkable. We adapted and switched to a nearest-neighbor matching approach more suitable for the data realities of the evaluation. That adjustment is worth naming, not glossing over. The switch was a pragmatic response to what the data actually allowed, not the cleaner solution we had originally envisioned. Whether some precision was lost in the trade is a fair question. What mattered was understanding the underlying logic of the matching approach well enough to evaluate the alternatives and make a reasoned call — not just defaulting to something familiar.
Knowing why a method works matters as much as knowing when and how to run it.
Matching enables comparison across different groups, but, used on its own, leaves the question of attribution unanswered. To address this, we added sensitivity analyses using Rosenbaum bounds. The basic question they address is: even after matching on everything we could observe, what if there is some important difference between the groups that we did not measure — something that explains our results independently of the program? Rosenbaum bounds ask how large that unobserved difference would need to be before it could plausibly account for what we found. If the answer is “very large,” confidence in the finding increases. If the answer is “not very large at all,” that is a reason for caution. It is a useful discipline when randomization isn’t possible or appropriate, and it makes the limits of the design explicit rather than leaving them implicit.
Neither of these choices were strictly required. But if you are seeking to build a rigorous design, it is worth asking whether your analytic tools are doing it justice.
Naming what you cannot know
Expanding the design is only useful if you are equally willing to name its limits.
A central constraint in our evaluation was the rollout of a successor program covering a subset of the population and much of the same content as the initiative under review. By the time fieldwork began, many program participants had already been enrolled in this follow-on program, meaning those assessed as "post-program" were still actively receiving similar programming support. This is the kind of complication that can be glossed over or buried in a methodology footnote. Instead, we built it in. Our surveys included questions on prior and concurrent nutrition program participation, enabling a layering analysis that compared outcomes across participant communities with different patterns of program exposure. And we were direct about the implication: given the rollout of this successor program, the evaluation could not attribute sustained outcomes exclusively to the program under evaluation. Findings on sustainability were reported with that reality explicitly stated, not softened.
Challenges like this also provide opportunities to enrich evaluations, when approached as conditions to accommodate and not as constraints that require compromise. In our evaluation, we chose to move away from treating this multiple programming reality as merely a constraint by designing the evaluation to also examine how participant experiences of multiple programs impacted outcomes. As a result, our findings provided insights into the effects of different levels of program exposure.
The real goal of evaluation design
The Nepal evaluation was not a perfect design. No evaluation of a completed program in a resource-limited setting ever is. What it was is the most rigorous design the context allowed, applied as carefully as we could manage, with findings reported in a way that reflected what the evidence actually supported.
That is the goal. Not the cleanest possible design in theory, but the most credible design in practice — one that is ambitious enough to answer the questions that matter, realistic enough to be carried out, and honest enough that decision-makers can use what it produces.
Design choices made early in an evaluation reverberate through everything that follows. Getting them right means resisting the pull toward the familiar and the convenient, thinking carefully about what a given context actually allows, and — above all — being willing to say clearly what you cannot claim as well as what you can.
From evidence to meaning
Right-sizing evaluation design determines what kind of evidence an evaluation can produce. But evidence does not become useful on its own – it requires interpretation.
In the next article, we look at how to move beyond descriptive statistics toward understanding what the evidence actually means.
Ben Jaques-Leslie is part of the Apricity team. This piece draws on experience gained through that work but reflects his own views and conclusions.