This post explores ambiguities, disagreements and unknowns that inhabit evaluation processes and products. It shows how uncertainty can hide inside stories that look clearer than they are, and offers practical ways to design and report evaluations that surface the uncertainty and support better decisions.
The uncertainty monster is alive and well in evaluation. You might meet it when a commissioner asks for a single “what works” verdict from a complex portfolio, or a board wants a clean-cut benefit-cost ratio to sum up a decade of innovation. Or you might sense its presence in the unease that our judgements rest on incomplete evidence and assumptions, or that the numbers and traffic lights in a dashboard are so simplistic they risk misleading.
Jeroen van der Sluijs (2005) introduced the “uncertainty monster” to science-policy work, as a metaphor for hard-to-tame uncertainties. Judith Curry and Peter Webster (2011) then picked up and extended the concept, applying it to climate science where the stakes and disagreements about uncertainty are high. Their point was that the climate has aspects that are well understood and aspects that remain deeply uncertain, especially for decision-making. If that feels a bit uncomfortable, well, that’s the point. That is the monster.
When I sat with their arguments, it struck me that similar creatures prowl through evaluation - in theories, data, causal inferences, evaluative judgements, and more. The discomfort they describe is a recurring theme in how people handle ambiguity, disagreement, and ignorance in policy and program evaluation.
This post borrows the monster metaphor and some of the structure of the argument from Curry & Webster’s 2011 “uncertainty monster” paper, and Curry’s 2023 book, Climate Uncertainty and Risk. I’m focusing here on the arguments about uncertainty and decision-making, not Curry’s positions in climate science and policy.
Meet the uncertainty monster
Most evaluators have met some version of it. A data analysis may throw up some surprising results, leaving us unsure whether findings cast doubt on the theory of change or are just an anomaly. A benefit-cost ratio could swing wildly when one assumption is tweaked. Different evidence sources might tell conflicting but equally plausible stories. Stakeholders may fundamentally disagree about what good value looks like, meaning there’s no single, uncontested standard against which the outcomes can be judged. Sometimes the monster is baked into the scope from day one, when evaluators are asked to look at implementation and outcomes but explicitly told not to question whether the program was well-conceived or well-designed in the first place, effectively ruling consequential uncertainties out of bounds before we even start.
In each of these cases, uncertainty is more complex than just statistical noise. It’s a signal about contested purposes, imperfect models, missing answers, or forbidden questions. But the work still has to land on a valid, actionable set of findings.
The monster gives a name to the discomfort that arises in decision-making where knowledge and ignorance collide, or when debates straddle boundaries often treated as separate - for example, facts and values, objectivity and subjectivity, prediction and speculation, science and policy (van der Sluijs, 2005). In Curry’s climate work, the monster lurks in model inadequacy and structural errors, parameters whose values are only loosely pinned down by data or theory, scenario uncertainty, and recognised ignorance about key processes, as well as in the gap between what those uncertainties can honestly support and the precise probabilities and ‘likely’ ranges decision‑makers would ideally want.
In Curry’s subsequent book, Climate Uncertainty and Risk (2023), she used the metaphor as part of a more general diagnosis of how societies grapple with wicked problems under deep uncertainty. She argued that political and institutional pressures tend to demand simplified narratives (sometimes preordained) and confident numbers, even though the underlying systems, being dynamic, non-linear, and context-dependent, resist being squeezed into narrow probability distributions or definitive statements. The result is the monster: the neglected remainder of what we can’t be certain about, can’t yet know, or can’t adequately quantify, but must still act upon.
How we try to tame it
Curry and Webster (2011) described several pathological strategies for “taming” the monster, such as:
“Monster hiding”: uncertainty is downplayed or buried in technical annexes to avoid complicating the policy message. In evaluation, for example, a summary report might headline a single rubric rating or benefit-cost ratio, while the sensitivity analysis and caveats are tucked away in a detailed report that few decision-makers will read.
“Monster simplification”: ignorance and structural uncertainty are turned into precise probabilities or tight confidence intervals, even when the evidence base is thin. For example, a complex social program may be scored with estimated likelihoods of impact in different domains based on one widely cited study, even though the underlying theory of change is contested and the data are sparse.
“Monster exorcism”: calls for more research are used as a way to postpone uncomfortable conversations about limits of knowledge or about values and trade-offs. In evaluation terms, recommendations may lean heavily on the need for more data or a larger trial, instead of helping commissioners make well-informed decisions now, under acknowledged uncertainty.
These examples sit alongside other patterns described in the literature, such as “monster denial” (Knotters et al., 2024, completely ignoring uncertainty) and “monster adaptation” (van der Sluijs, 2005, of which simplification is one form, forcing messy uncertainties into something that can be quantified and modelled). In each case, the theme is a reluctance to let uncertainty appear in the main story.
Another strategy in Curry and Webster’s paper is “monster detection”. This spans several strands, some helpful, some not. For example, one strand is the boundary-pushing researcher who probes accepted claims and expands the edges of what is known. Another is the watchdog who scrutinises methods and evidence to protect rigour and transparency. Both of these roles are valuable in science and policy.
A third, more corrosive strand is the “merchant of doubt” who cherry-picks uncertainties to stall decisions or protect vested interests. That kind of uncertainty work is as unhelpful in evaluation as it is in climate research. Our aim should be to surface and understand uncertainty so that action can be more robust and accountable.
In her book, Curry also highlighted the political economy of these moves: different actors have incentives either to suppress or weaponise uncertainty, to strengthen their particular arguments for or against a course of action.
For Curry, the way forward is “monster assimilation”: learning to live with uncertainty, making it explicit, integrating it into risk management and decision processes instead of pretending it can be eliminated.
The monster in evaluation
Evaluators work with complex, adaptive systems, uncertain evidence1, contested values, and decisions with unknowable consequences. Yet our findings are sometimes expected to display an almost heroic confidence because commissioners, boards, and political systems are under pressure to make clear decisions based on imperfect information. The uncertainty monster tends to get tamed in familiar ways, such as:
Logic models as monster cages: Theories of change and results frameworks, though useful, can act as beautifully packaged diagrams that hide unresolved causal ambiguity, unknown pathways, feedback loops that blur the distinction between causes and consequences, unstated assumptions, and more. They risk domesticating complexity into something that looks linear and predictable.
Point estimates as monster masks: Cost-benefit analyses (CBA) and social return on investment (SROI) studies sometimes present benefit–cost ratios (BCRs) as if we could know, with precision and certainty, that every dollar we invest in a program creates exactly $3.72 worth of social value. Structural uncertainties, like political risk, institutional capacity, knock-on effects and behavioural responses, may be acknowledged, but the takeaway number still reads like a fact. For example, a social investment business case might headline a single BCR, even though political risks play out in ways that defy probabilistic sensitivity analysis.
Summary reports as monster tranquilisers: Evaluations may respond to (understandable) pressures to provide a single line to the board, e.g., a manageable narrative about “what works”. Dissonant findings, big contextual caveats, or deep disagreements about mechanisms may be smoothed over in the interests of coherence and communicability. In this way, the framing of evidence can be a political act as much as a scientific one.
Curry and Webster, drawing on Agassi (1974), used the term “bootstrapped plausibility” for a form of circular reasoning in which a claim is accepted as plausible‑looking, and that apparent plausibility then lends credibility back to some of its more doubtful premises. In evaluation, theories of change, benefit‑cost ratios, and summary narratives (among other things) can work this way. For example, a neat causal chain in a diagram can make its starting assumptions about behaviour change or system response look more warranted than the evidence really allows; a BCR can make a stack of assumed parameters and proxies feel more solid than they are; a carefully negotiated line for the board can make underlying disagreements seem more resolved than they are. The overall story becomes so convincing that we may discount our own biases, “believe our own press”, and stop scrutinising the fragile assumptions underneath.
In each case, the monster hasn’t disappeared. It is still there, forgotten but not gone. The risk is that all of us, from evaluators to funders and policymakers, end up making apparently decisive choices based on overly tidy stories and over-confident metrics that fail to surface the uncertainty hiding underneath.
Lessons from Curry’s risk framing
Some of the most interesting parts of Climate Uncertainty and Risk for evaluators are about how to think about risk when uncertainty is deep. Curry distinguishes between situations where probabilities can reasonably be estimated and those where the problem is dominated by structural ambiguity, contested models and evolving systems. In the latter, she leans toward approaches that emphasise adaptability and local risk management over centrally imposed, one-size-fits-all solutions.
Transposed into evaluations in complex settings, this suggests several shifts:
From optimising for the ‘most likely’ scenario to asking how a decision performs across very different futures
From assuming away low-probability shocks to asking what would break the strategy and how we would respond
From treating disagreement as noise to using it to map what we don’t know, which assumptions are contested, and where values differ.
Curry’s work on “possibilistic” scenario analysis is about exploring plausible, high-impact futures without pretending to know their exact probabilities. Instead of asking “what is the most likely outcome?”, the question becomes “what are the serious futures we cannot rule out, and how would we cope if they arrived?”.
For evaluators, a lesson is to ask how a decision can keep options open and create room to learn and adapt, rather than locking everything in on day one. To borrow some investment language, this connects to ideas like real options (building in choices later), adaptive pathways (sequencing decisions so they can change over time), and explicitly valuing flexibility and learning in appraisal and evaluation.
Doing evaluation with the monster
My practical take: If the monster isn’t going away, the question for evaluators is how to work with it and make it useful. That means building designs and reporting habits that normalise uncertainty and help decision-makers act with their eyes open. The strategies that follow are small shifts that can be woven into existing evaluation practices, shifting attention from taming or hiding uncertainty to working with it explicitly. Many evaluations already do these things.
1. Disaggregate uncertainty types
Instead of a generic caveat paragraph, use a simple table in the methods or limitations section of an evaluation report to distinguish at least three layers of uncertainty in designs and reports, mirroring Curry’s typology of statistical, scenario, and deep uncertainty:
Statistical uncertainty: e.g., sampling error, measurement noise, wide confidence intervals, volatile results in subgroups. The appropriate response is better design and measurement (larger or more representative samples, improved instruments, appropriate models) and transparent presentation of intervals in addition to point estimates.
Scenario uncertainty: alternative plausible ways political, economic, or implementation events could play out. In evaluation, this could include uncertainty about future funding, policy shifts, workforce capacity, or partner behaviour that could change how (or whether) impacts materialise, even if current trial results look strong. The response is to stress-test conclusions with simple scenarios (e.g., different rollout speeds, adoption rates, shocks) and to frame recommendations as contingent on particular conditions. I’ve written before about scenario analysis and break-even analysis in CBA to do exactly this.
Recognised ignorance (deep uncertainty): areas where causal structures, system boundaries, long-run dynamics, etc, are poorly understood. For evaluators, this covers phenomena such as unknown knock-on effects in complex systems, emergent behaviours that were not in the theory of change, and contexts where there is little prior evidence or where stakeholders fundamentally disagree. The response is epistemic humility: naming these limits of knowledge explicitly, avoiding spurious quantification, and recommending adaptive management, learning cycles, or further exploratory work rather than making unjustifiable claims about long-run value.
2. Match your claims to what you know
Curry’s critique of “monster simplification” – using precise probabilities where evidence is thin – has obvious application in impact and cost-benefit estimates. For some questions, directional judgments (“likely positive but highly uncertain”) or wide ranges are more honest than a spurious decimal and deserve to be up-front in key findings. This demands a shift in how commissioners read reports, actively rewarding candour over false precision.
3. Expose value choices
Curry’s book is explicit that risk assessment isn’t value-neutral; it embeds judgements about whose risks count, which time horizons matter, and what trade-offs are acceptable. Evaluations do the same, via things like explicit criteria and standards, and transparent positionality. Economic evaluations expose value choices through decisions about which impacts are counted as material, measurable, attributable and monetisable, how they’re monetised, how long they’re expected to last or matter, and how steeply their present value diminishes with time. Making these choices transparent and open to deliberation is a way of bringing part of the monster into the light.
4. Analyse possibilistic scenarios
A practical way to bring the monster into view is to work with Curry’s “possibilistic” scenarios (plausible, high-impact futures whose probabilities are not well understood). Instead of asking only “what is the expected return?”, we can ask “how do our conclusions hold up if the world turns out to be tougher or kinder than our central case?”
To illustrate, imagine a government funder reviewing a portfolio of youth mental health programs. A standard CBA might estimate an overall benefit-cost ratio (BCR) of $2.50 for each $1 invested, based on historical service utilisation and costs, and modelled long-term savings to the health and justice systems. A monster-aware approach would ask different questions, such as: how does the BCR hold up under a series of plausible, hard-to-quantify shocks - for example, a change of government that tightens eligibility rules, a workforce shortage that lengthens wait times, or a social media trend that suddenly increases help-seeking among young people? The evaluation would explicitly construct scenarios around shocks like these, and examine how portfolio performance changes if they occur, separately or together.
Crucially, this framing can shift our understanding of what good looks like: the ‘best’ intervention in this portfolio might not be the one with the highest BCR, but the one that remains acceptable, or creates options to adapt, across a wide range of futures.
5. Use extended peer communities to interpret uncertainty
Curry’s call for “extended peer review”, drawing in a wider range of perspectives to interrogate assumptions and interpretations, is readily translated into evaluation practice. Bringing implementers, communities, and diverse disciplinary voices into co-interpretation workshops or sense-making sessions can surface alternative framings, contested values, and overlooked risks that a narrow technical team might miss. This way of working is integral to my view of evaluation, as readers of this series will recognise.
6. Design for robustness and learning
Robustness, in a deep‑uncertainty sense, is about program design and evaluation that hold up across a range of possible futures, rather than being optimised for a single best guess. As an obvious example, it is why we diversify our investments even if we feel a particular racehorse is a sure winner.
In the climate field, Curry argues for risk frameworks that prioritise resilience and options for adaptation under uncertainty. In evaluation, it suggests value may be placed on intervention designs that create options, capabilities, and adaptive capacity (for example, flexible delivery models and modular components) rather than just immediate outcomes. It suggests giving weight to investments that leave systems better able to learn and adjust, even if ultimate impacts and benefit–cost ratios are harder to pin down.
Designing these principles in aligns with longstanding evaluation approaches such as Michael Patton’s Developmental Evaluation, which treats complexity, emergence and uncertainty as normal conditions for innovation and emphasises real‑time learning and adaptation. It also resonates with approaches like Andrew Hawkins’ Propositional Evaluation, which treats programs as propositions for action and uses explicit argument structures to develop valid, well‑grounded propositions and to identify and manage risks of program failure over time.
Not every evaluation needs the full monster-handling toolkit.
The point is to do something appropriate and proportionate for the context. For example, even a short report section that distinguishes different kinds of uncertainty, or sketches a couple of plausible scenarios, is a meaningful step beyond a single heroic number.
Living with the monster
Judith Curry works in evolving and contested spaces in the climate field, her perspective is one among many, and the scientific community debates some of her arguments about evidence and policy implications.2 That contestation matters and should be acknowledged, because it is part of the very uncertainty dynamics she describes and is, after all, one of the ways science advances. But the point I’m exploring here is a different one: debates about how uncertainty is characterised, analysed, communicated, and used in decision-making are real and consequential in many of the contexts evaluators work in.
One does not need to form opinions on Curry’s climate thesis to borrow insights from her treatment of the monster metaphor. Where her work is particularly useful for evaluation is in insisting that uncertainty, disagreement, and ignorance are central features of the terrain on which decisions are made, and that how we deal with the monster can affect what gets decided. That’s my focus here.
For our profession, the invitation is to shift from treating uncertainty as an inconvenience to downplay, or as something that can be tamed with confidence intervals and probability weights, to instead recognising it as a more complex characteristic of things we evaluate. The uncertainty monster won’t go away, but it can become a guide, provoking better questions, richer conversations, and evaluation designs that are honest about uncertainty. Like a lot of monsters, this one might not be so scary once we get to know it.
None of this implies that ‘nothing is known’ or that all claims are equal; it is about matching the strength and style of our claims to the type and depth of uncertainty we actually face. It’s about evaluative thinking.
And, none of it means evaluations should shy away from providing clear answers to important questions. Commissioners and communities are entitled to well-reasoned judgements, not endless fence-sitting. The point is that these judgements are stronger when they’re honest about the uncertainty that surrounds them, rather than shoving the monster under the bed and hoping it keeps quiet.
Thanks for reading!
These posts are offered as contributions to an ongoing professional conversation. They’re not the last word; thoughtful questions or comments are very welcome. What other ways can we live productively with uncertainty monsters in evaluation?
References
Curry, J. A., & Webster, P. J. (2011). Climate science and the uncertainty monster. Bulletin of the American Meteorological Society, 92(12), 1667–1682.3
Curry, J. A. (2023). Climate uncertainty and risk: Rethinking our response. Anthem Press.
Knotters, M., Bokhove, O., Lamb, R., & Poortvliet, P. M. (2024). How to cope with uncertainty monsters in flood risk management? Cambridge Prisms: Water, 2, e6.
van der Sluijs, J. P. (2005). Uncertainty as a monster in the science–policy interface: Four coping strategies. Water Science and Technology, 52(6), 87–92.
I’m a little obsessed with humanity’s tendency to equate evidence with “facts”. In evaluation, a lot of the evidence we have to consider is uncertain. For example, real-world evidence includes probabilities, tendencies, theories, hypotheses, propositions, interpretations of incomplete or ambiguous data, modelled scenarios based on constellations of assumptions, expert opinions, educated guesses, and endeavours to synthesise and communicate complex issues in digestible terms - all of which involve human judgement. Even “hard” facts can change: they remain open to revision as new evidence emerges (because science). No matter what methods are used, facts are not neutral or values-free - they’re influenced by values and beliefs. For example, people make decisions about what counts as credible evidence, what data to collect, how to collect it, how to analyse it, and how to interpret it. All these methodological decisions are guided by values (whether declared or not). The reporting of facts, too, reflects values and biases; what is considered a fact is influenced by what is deemed important or relevant by the people presenting it. Any time we look at a fact, we could think about where it came from and what biases may lurk beneath its surface. Examples include measurement bias, sampling bias, selection bias, recall bias, confirmation bias, and publication bias. And that’s when everybody’s on their best behaviour. Unfortunately, purveyors of evidence are not immune to racism, sexism, pecuniary interests, career ambition, politics, social conformity, groupthink, noble lies, dogma, hubris and bullying. These human frailties can affect the veracity of any facts. Recognising that facts are produced and interpreted through human systems does not mean ‘anything goes’, but it does mean we attend carefully to method, bias, accountability, and working effectively with the uncertainty monster. This footnote summarises selected points from my post on Objectivity and Subjectivity in Evaluation.
Judith Curry is an atmospheric scientist and former professor who has worked on tropical cyclones, climate dynamics, and climate risk, and who has also been a public commentator on climate policy debates. She now works as a consultant on climate risk and adaptation, advising organisations whose assets, operations, and communities are exposed to real risks in the face of uncertain climate futures. Curry argues that climate uncertainty is deeply rooted in the complex, nonlinear, and dynamic nature of the climate system, and that these uncertainties limit what traditional models can say with confidence about long-term or regional futures. She is critical of approaches that present probabilistic projections as if they captured the full range of plausible scenarios, and instead advocates dynamic, adaptive risk frameworks and scenario methods that work more explicitly with deep uncertainty and robustness, rather than prediction alone. In a field where strong consensus narratives can make it uncomfortable to talk about what is contestable or not yet known, her writing has been both criticised and appreciated for putting uncertainty, disagreement, and institutional handling of risk squarely on the table.
Bear in mind that Curry and Webster’s 2011 paper responded to the IPCC Fourth Assessment Report; later assessments (AR5 and AR6) have updated methods and conclusions. Curry’s climate science and policy analysis sit well outside my lane; my interest here is in the treatment of uncertainty and what still translates into evaluation practice.



Fascinating. Thank you for articulating this so clearly. The idea of explicitly surfacing the 'uncertainty monster' rather than trying to hide it is critcal, especially when we're dealing with complex systems and incomplete data. It's vital for transparent and ethical decision-making.
Nonlinearity is the rule! This is an unruly business.