Beyond Descriptive Statistics: Turning Data into Decision-Ready Evidence
Illustrative: In our evaluation in Nepal, aggregate results showed a flat picture. It was only once the data was disaggregated that the real story came into view.
By Sarah Staub and Ben Jaques-Leslie
Most evaluations generate data. Fewer generate understanding. This distinction matters because decision-makers rarely need more numbers. They need evidence to help them understand why outcomes occurred, what factors influenced them, what should happen next, and what lessons can be applied to future programs.
Descriptive statistics are an essential starting point. They tell us how many people participated, how outcomes changed, and whether differences exist between groups or locations. They help us measure who was reached, what was delivered, and what changed. But problems arise when descriptive statistics become the endpoint rather than the beginning of analysis.
Programs do not operate in a vacuum. They unfold in complex environments shaped by geography, infrastructure, social and cultural dynamics, policy contexts, implementation realities, and human behavior. Understanding whether a program worked is important. Understanding why it worked—or why it worked differently in different places—is what makes evaluation useful.
Treating Differences as Signals, Not Noise
The most useful findings often emerge from places where results diverge. The instinct in evaluation is often to look for consistency. When outcomes differ across locations, when stakeholder perspectives do not align, or when indicators point in different directions, those differences can feel like complications: something to control for, smooth out, or explain away.
But differences are where the most useful questions emerge—and where the learning lives. Geographic variation can reveal contextual factors that shape implementation. Divergent stakeholder perspectives can expose assumptions that would otherwise remain invisible. Differences in outcomes can point toward mechanisms of change that a single average will never reveal.
The question becomes less "How do we account for these differences?" and more "What are these differences telling us?" That shift—from treating differences as noise to treating them as evidence—often marks the point where evaluation moves beyond description and toward explanation.
Looking Beyond the Headline Findings
Apricity’s recent evaluation in Nepal provides a good example of this dynamic. One of the questions we addressed was whether outcomes differed across program locations and, if so, what those differences could reveal about how the program operated in different contexts.
At first glance, some findings appeared inconsistent across districts. Looking only at the aggregate results would have produced a flat picture of program outcomes. The more interesting findings emerged when those results were examined in greater detail.
In one district, participants experienced weaker program outcomes than in others. A purely descriptive analysis might have reported the difference and moved on. Instead, that difference became the starting point for further investigation.
As we explored the findings more closely, it became clear that communities were living in very different contexts. In some areas, households had limited access to land, constraining what the program could deliver regardless of implementation quality. In others, participants were enrolled in multiple concurrent programs, making it difficult to isolate the contribution of any single intervention.
In both cases, the differences helped us move beyond observing whether the program was working to understand how and why it impacted beneficiaries differently across areas.
What initially appeared to be a difference in program performance turned out to be, at least in part, a difference in context. That distinction matters. The lesson for future programming was not simply that outcomes were higher or lower in one district than another. It was understanding how local conditions shaped those outcomes and what that meant for future program design. The numbers helped identify where the differences existed. Understanding the context behind those differences helped explain why.
The most valuable insight was not the average result. It was understanding the factors that shaped variation in those results and what those factors suggested for future programming.
When examined carefully, differences help explain why outcomes vary, where adaptation may be needed, and what lessons are most relevant for future programming. The Nepal evaluation reinforced an important lesson we have encountered repeatedly: some of the most important findings only become visible once results are disaggregated.
Moving Beyond Averages
Averages are often less informative than they first appear. They provide useful summaries, reveal broad patterns, and help communicate findings efficiently. But they can also obscure the variation that often contains the most valuable learning. Geographic differences, implementation realities, stakeholder experiences, and contextual factors can all disappear when results are collapsed into a single number.
This does not mean averages are unhelpful. It means they are rarely the end of the story. Building a deeper understanding requires disaggregating findings across geographies or participant groups, combining quantitative and qualitative data, or examining implementation processes alongside outcomes – not just what happened, but how and where.
Often, it simply means asking one additional question: Why do these differences exist? That question frequently produces the most valuable learning—and often the most useful guidance for future decision-making.
From Reporting Results to Supporting Decisions
At its best, evaluation is not simply a reporting exercise; it is a learning process. Organizations invest in evaluation to improve programs, make better decisions, and understand what should happen next. Achieving these goals requires more than documenting results. It demands understanding the factors that produced them.
Descriptive statistics remain essential — but they are the starting point, not the destination. The goal is not to know what happened. It is to understand why it happened differently for different people, in different places, under different conditions.
One important part of that understanding comes from examining how programs are experienced by different stakeholders, particularly the communities they intend to serve. Those perspectives often provide the most direct explanation for why outcomes vary — and why the same program lands differently in different communities. In our next post, we explore how community-centered and participatory approaches help evaluators access those perspectives and build evidence that is grounded and ultimately more useful for decision-making.
Sarah Staub and Ben Jaques-Leslie are part of the Apricity team. This piece draws on experience gained through their work at Apricity but reflects their own views and conclusions.