Ever found yourself having to defend case studies to stakeholders who equate rigour with large samples and statistical generalisation? This post is for exactly those conversations, using a widely cited paper on case study research and the concept of analytic generalisation to show how carefully chosen cases can support learning and credible evaluation.
A scenario
Imagine we’re evaluating a large research fund. The fund has provided grants for projects across many years, sectors, places, and sizes - everything from biomedicine to social policy, from small grants for local pilots to massive investments in multi-site initiatives run by national consortia.
To help evaluate the fund and understand how it works in practice, we’ve selected a sample of grants to use as case studies. Our sample broadly reflects the diversity of grants. It isn’t large enough to provide statistically generalisable findings - but that isn’t our aim. Rather, we plan to dig deep, build understanding by comparing different cases, and draw out lessons about how the fund delivers on its value proposition.
To some stakeholders, an approach that doesn’t produce statistically generalisable findings feels at odds with their understanding of good research. They raise concerns about our methodology, expressing scepticism and challenging us to defend our design.
Five misunderstandings
One of the most-cited papers in the journal Qualitative Inquiry, ”Five Misunderstandings About Case Study Research” (Flyvbjerg, 2006) challenged common criticisms of case studies, systematically pulling apart five prevailing myths - that case studies can't offer theoretical insight, can't be generalised from, only generate hypotheses rather than test them, are prone to researcher bias, and are too hard to summarise or use to inform broader lessons. Through practical examples and philosophical reflection, the paper argued that these beliefs are either oversimplified or just plain wrong. Case studies, when carefully selected and rigorously conducted, are a powerful way to build context-rich knowledge. This post closely follows Flyvbjerg’s “five misunderstandings”, with my gloss on what they mean for evaluation.
Myth 1. Abstract theory is more valuable than context-dependent, practice-based knowledge
The misunderstanding: In practice, people may default to treating high-level theoretical knowledge about ‘the general case’ as more valuable than concrete, practical knowledge from specific cases, especially when they are under pressure to show scalable, comparable results. For example, in our grant fund evaluation, some stakeholders may worry that lessons from individual projects will be too tied to their unique settings to inform broader program strategy.
Flyvbjerg’s argument: Context-dependent, practical knowledge is essential for developing real expertise; theory matters, and it becomes more meaningful when understood in light of specific cases. In complex real-world interventions, lessons from the field are needed to drive learning and improvement, and people responsible for designing, delivering and governing programs move from rule-based “beginners” to genuine experts by engaging deeply with many concrete projects and contexts over time.
Implications for our grant fund evaluation: Case studies will help us learn how things worked in different settings (such as urban academic centres, regional hospitals, and small-town community-led initiatives). They’ll contribute a nuanced understanding that’s hard to get from surveys or large datasets alone. Broad theories also have their place, but the practical lessons from the case studies will be very useful, capturing rich details that are essential for informing practice and policy.
Myth 2. You can’t generalise from a single case
The misunderstanding: A common view is that you can’t responsibly generalise from one (or a few) cases, so case studies have limited value for program-wide decisions. In our grant fund evaluation, some stakeholders are not completely against case studies, but they express doubts about whether in-depth studies of a few projects can say anything useful about the fund as a whole.
Flyvbjerg’s argument: Generalisation is possible, especially when cases are well chosen. For example, a carefully selected typical, or “paradigmatic” case can teach us a lot, by highlighting characteristic features of a system and serving as a practical exemplar of how its mechanisms play out in everyday conditions. A critical or extreme case can help test or challenge broader assumptions. With thoughtful selection, cases can inform wider theory and practice, not by representing a statistical population, but through “the force of example”.
Implications for our grant fund evaluation: Our sample of case studies doesn’t claim to represent all grants in a statistical sense. Instead, it is a route to analytic generalisation - using carefully chosen cases to refine ideas about how the fund works, for whom, and under what conditions (Yin, 2010). By focusing on how individual projects reflect or challenge the fund’s overall value proposition and theory of change, we can spot patterns and principles that matter beyond each case.
To make this work in practice, we will be deliberate about case selection. Rather than a convenience sample, we’ll purposefully choose cases for their information value, including cases that illustrate “typical” conditions, plus a small number of critical or extreme projects (e.g., very small local pilots and very large multi-partner collaborations). This kind of information-oriented selection supports analytic generalisation, because it allows us to see how core mechanisms play out across contrasting conditions, and to identify where the fund’s value proposition holds, where it wobbles, and where it breaks.
Myth 3. Case studies can’t build or test theory - they only generate hypotheses
The misunderstanding: Case studies are often seen as just a starting point, useful for generating ideas but not strong enough to test or develop theories. For example, our stakeholders would be happy if we used case studies as a preliminary exploratory phase, then relegate them to the background once surveys and quantitative methods begin. However, we don’t want to do that because it would miss the opportunity for case-based insights to reinforce or challenge the program’s theory of change.
Flyvbjerg’s argument: Case studies can generate, test, and refine theories. They are powerful tools for both developing new ideas and putting existing theories or assumptions to the test, not only by suggesting new hypotheses but also by subjecting them to tough, context-rich examination where they can be falsified or forced to evolve. This brings us back to analytic generalisation: as Yin (2014) emphasised, when case studies are designed around explicit theoretical propositions and replicated across different contexts, they can advance or revise theory.
Implications for our grant fund evaluation: Case studies are more than just ‘early-stage’ exploration; by examining how the theory of change holds up in specific contexts, they can help refine, validate, or challenge core assumptions. For example, suppose the fund’s theory assumes that providing seed funding plus light-touch mentoring is enough to scale promising pilots. In one case study, we might see that this works well in a well-connected university setting but fails in a small community organisation with limited infrastructure. In another, we might learn that intensive brokering between partners, not the size of the grant, was a decisive mechanism.
Taken together, these cases would not only generate new hypotheses about what drives success, they would also put the fund’s theory of change under pressure in real-world settings. This is genuine theory testing - revising or retaining theoretical propositions based on how well they explain contrasting cases and contexts.
Myth 4. Case studies have a bias toward verification
The misunderstanding: Critics sometimes argue that case studies just confirm what the evaluator already thinks, because you can find whatever you look for in a messy real-world environment. For example, a couple of our evaluation governance group members are sceptical, suggesting that we will cherry-pick stories or interpret findings to support our personal views.
Flyvbjerg’s argument: Actually, in-depth case work often challenges initial assumptions. When researchers get close to the real action, they’re more likely to be confronted by surprises that upend their preconceptions and prompt new lines of thinking. Far from being merely affirming, case studies can and do falsify assumptions.
Implications for our grant fund evaluation: We will go into our case studies expecting our thinking to be challenged. Since case study work is so closely tied to front-line experiences, it’s more likely to expose flaws in the logic or implementation of the program than to simply confirm standard narratives. For us, this is an opportunity to spur learning and honest reflection. To guard against cherry-picking or confirmation bias, we will be transparent about how we select cases and what we are comparing with each case. We’ll look deliberately for rival explanations and disconfirming evidence, and use explicit propositions to link case evidence to findings.
For example, if a case seems to show strong outcomes in a particular sector, we will explicitly look at whether this is because of the fund’s design or because that sector already had highly capable organisations and pre-existing networks. We can also explore whether it was a bit of both, or whether some other factor was behind the patterns observed. Examining these rival explanations helps keep us honest and makes our reasoning easier for others to scrutinise.
Myth 5. Case studies are too detailed to summarise or turn into usable lessons
The misunderstanding: Because case studies are specific and detailed, some think they’re impossible to summarise or turn into clear, usable lessons for wider decision-making. In our evaluation, stakeholders express concern that case study findings will be so context-dependent and narrative-rich that they can’t be organised into useful recommendations or transferrable lessons for the whole fund.
Flyvbjerg’s argument: While it’s true that the richness of a case can’t always be boiled down to one-sentence conclusions, this is a valuable feature. Detailed narratives may resist reductive summary, but they offer deep insights that matter for practice and theory development; the difficulty of neat summarisation reflects the complexity of real social situations, not a flaw in case study methods.
Implications for our grant fund evaluation: We will embrace the detail. Not everything needs to end in bullet points, summary statistics, or simple conclusions. The stories themselves are important, especially when they shed light on how the fund’s aims play out in the day-to-day realities of funded projects. We will summarise what we can, but we’re not worried if some findings resist being squeezed into a brief ‘takeaway’ - because that’s the point. Sometimes, it’s the complexity and specificity that make the case meaningful for future policy or learning.
Discussion
When we draw on case studies in evaluation, many of us run into the demand for “generalisable” findings. The traditional expectation, especially from those more familiar with large-scale surveys or experimental methods, is that you should be able to make inferences from your sample to a wider population. In other words, conclusions would, in theory, hold for all grants or projects, not just the ones you looked at. That’s statistical generalisation.
This is where that familiar question pops up: “But is it representative?”, as if anything short of a large survey automatically fails the quality test.
In our fund evaluation, with its diverse mix of sectors, scales, and settings, statistical generalisation is not the goal. Polit and Beck (2010) described statistical generalisation as just one of three models of generalisation, alongside analytic and case-to-case forms, which can be appropriate for qualitative or case-based work. Here, we are pursuing what Yin (2010) called analytic generalisation: we use case studies to build and test ideas about patterns and mechanisms, then consider where those ideas are likely to hold across different kinds of projects and contexts. Rather than claiming that something is true of all projects, we are asking how the mechanisms and lessons we see in these cases strengthen our understanding of the fund.
In practice, this means stepping back from individual cases to look across them and identify patterns, lessons, and mechanisms that speak to the fund’s theory of change and value proposition. ‘Theory’ here does not have to mean grand, abstract models. It includes mid-range mechanisms, practice-based frameworks, or simple ‘if-then-because’ propositions about how change is expected to happen. We should state up front which ideas or assumptions we’re exploring, so the path from case to conclusion is clear and open to scrutiny.
Typically, analytic generalisation involves comparing the results of case studies to existing theories or frameworks to see how the cases support, refine, or challenge them. We can then argue, with reference to both theory and context, how and where insights may be transferable or relevant to other cases, without pretending they apply everywhere.
By anchoring our findings in the fund’s theory (in this broad sense), we can speak to program-level mechanisms even if our sample is small. We’re aiming for transferable learning, not universal truths: the aim is not to claim that all projects will follow the same path, but to make a reasoned case for when and why certain dynamics observed in our cases might play out elsewhere, in line with Polit and Beck’s (2010) notion of case-to-case transfer.
Analytic generalisation gives us a disciplined way to learn from a small, carefully chosen set of case studies, generalising findings to theoretical propositions, not populations (Yin, 2010; 2014). It provides a bridge from the rich detail of individual projects to broader program learning, supporting strategic thinking and evidence-informed decision-making, even when statistical generalisability isn’t possible or meaningful.
As an aside, this is the same logic I used in my doctoral research. I developed a conceptual model for evaluating value for money, translated it into a series of theoretical propositions, and tested it through two international development case studies, abstracting from the cases back to the propositions through analytic generalisation and replication.
Of course, the case studies may just be one component
If we also want some statistically generalisable findings - i.e., findings that can confidently be extended from our sample to the broader population of projects in the grant fund, we can do that too - for example, we could run a survey across a larger sample of grants.
Using both strategies together (case studies for depth, survey for breadth) is an example of a mixed methods design that gives us a richer and more robust evidence base than we would get from either study alone. While the case studies help us learn how and why things work (or don’t) in real-world contexts, surveys provide a less-detailed picture of a few key areas of enquiry across the fund.
For instance, a short survey could ask all grantees about their perceived level of support from the funder, the extent and quality of collaboration with partners, and whether key milestones were achieved on time and within budget. These items won’t match the nuance of case narratives, but they can show how widespread certain challenges or successes are, and where the in-depth cases are typical or unusual. Each method helps compensate for the other’s blind spots.
Rather than simply run each method in parallel, their findings can interact and sharpen each other. For example, insights from case studies can help prioritise and refine survey questions. Case narratives can provide exemplars that bring survey-based findings to life. Quantitative findings can help us gauge the magnitude of impacts and challenges and determine whether patterns seen in case studies are common or exceptional.
Bottom-line
Case study-based evaluations of complex, diverse programs are not ‘second-best’ apologies. They generate their own insights for real-world learning and analytic generalisation, showing how and why change happens in practice and helping to refine the theories and strategies that guide future decisions.
Further reading
Flyvbjerg, B. (2006). Five misunderstandings about case-study research. Qualitative Inquiry, 12(2), 219–245. https://doi.org/10.1177/1077800405284363
Polit, D. F., & Beck, C. T. (2010). Generalization in quantitative and qualitative research: Myths and strategies. International Journal of Nursing Studies, 47(11), 1451–1458. https://doi.org/10.1016/j.ijnurstu.2010.06.004.
Yin, R. K. (2014). Case study research: Design and methods (5th ed.). Sage.
Yin, R. K. (2010). Analytic generalization. In A. J. Mills, G. Durepos, & E. Wiebe (Eds.), Encyclopedia of case study research (Vol. 1, pp. 21–23). Sage.
Thanks for reading!
If this was useful, a quick tap on the ❤️ helps me know it landed - and it also nudges Substack to show the piece to more people who might benefit from it. Thank you.



Thanks Julian for a useful post. I found Judith Green's article on understanding causality through case studies useful - https://link.springer.com/article/10.1186/s12874-022-01790-8 It gives very explicit arguments against the idea that case studies (and small n qual research in general) can't tell us about cause and effect. Lots of overlap with the points you make eg on testing hypotheses, understanding mechanisms. Working on a cross-disciplinary project with RCT researchers currently, so some of these debates about how we understand causality and the role of qual and small samples in that are fresh in my mind :)
On the same subject please see this post on "MSC ands causal inference" , especially the section on "The value of single cases" https://www.linkedin.com/posts/rickjdavies_msc-and-causal-inference-activity-7395948701296312320-zT44?utm_source=share&utm_medium=member_desktop&rcm=ACoAAAHy7vcB3IC16fzPUA54Tnh3jidtqJAmDaQ