Imagine an AI‑enhanced evaluation lab where you and your stakeholders rapidly simulate 1,000 evaluations using the same prompt, rubric, and evidence base, to reveal patterns, disagreements, and known unknowns. In this blog, I outline how multiple one-minute evaluations could visualise the range and distribution of reasonable ratings, communicate uncertainty, and ultimately strengthen human judgement.
We evaluators sometimes talk as if evaluation aims to uncover a single definitive truth. Give us a program, policy, or intervention, and we’ll tell you if “it works” and if “it’s good”. Many of us believe it’s more complex than that, of course, but still we’re often expected to navigate our way from clear questions to decisive answers.1 But what if there’s more than one possible valid answer?
Previously, I’ve shared an account of a one-minute evaluation I conducted in Perplexity Pro, using a single prompt. It was pretty good in some ways, but not good enough to inform a consequential public policy decision. In particular, the lack of stakeholder input and proper evaluator oversight were serious limitations.
Beyond that, there’s also an inherent problem with large language models (LLMs): their indeterminacy means you can keep using the same prompt over and over again and get different answers. If we potentially get something different each time, how do we know what to trust? After all, “hallucination” is widely understood to be a built-in mathematical property of these models. It hardly inspires confidence, does it?2
Yet, this problem isn’t unique to AI. Both humans and AI systems, when asked the same evaluative question multiple times, can give you different answers - because evaluation itself is fundamentally indeterminate.3
This raises an intriguing possibility.
What if running a quick (say, one-minute) evaluation a thousand times, under controlled settings (fixed prompt, evidence, and rubric), could tell us things we couldn’t get from a single assessment? What if variability in evaluative ratings4 isn’t noise to eliminate, but signal to be understood?
What follows is a case for treating repeated, rapid AI‑assisted evaluations as a way to map the landscape of reasonable judgements - not as an abdication of human evaluative responsibility, nor to exclude real deliberation with affected people, but as a way to help inform human judgements by providing extra information to deliberate on.
The mirage of one right answer
While some evaluation theorists and practitioners argue that careful attention to methodology can narrow indeterminacy to acceptable levels, constructivist and realist traditions challenge the notion that there even exists a single “appropriate” evaluation design or conclusion for any given program.
I was reminded of this recently, when checking out a blog from the Institute of Development Studies (Apgar, Cislaghi, and Wein, 2025), examining how holistic community-led development initiatives defy simple, standardised assessment. The Full Spectrum Coalition argues that for participatory and community-led programs grounded in diverse community visions of wellbeing and emergent change, evaluation must embrace plurality, flexibility and collaboration, recognising the idea of “one best design” as a mirage rather than a realistic goal. This resonates with what I’m describing here as task indeterminacy in evaluation; the recognition that one design and one conclusion may be insufficient in complex contexts.
Fundamentally, when we evaluate a program, we’re usually not measuring something fixed and objective like the boiling point of water. There may be aspects that we can measure or estimate statistically, but above that, we’re interpreting complex social phenomena through particular lenses, influenced by our own and others’ values, by context, and the specific moment of inquiry. Different evaluators may reasonably emphasise different aspects, and the same evaluator at different times may reach different conclusions. From a constructivist or realist perspective, this variation reflects the nature of evaluative work. From a positivist perspective, it might be seen as a limitation to be minimised through methodological refinement. Either way, some degree of indeterminacy remains baked in. More on this in the footnotes.
Indeterminacy in human and AI evaluation
Let’s be precise about what we mean by indeterminacy. In evaluation contexts, I’m referring to the reality that there’s often more than one warrantable conclusion that can be reached about the merit, worth or significance of a policy, program, or intervention.5
When we use AI in evaluation, indeterminacy has two distinct layers:
Ambiguity in the intervention and criteria, where several value judgements could be defensible in light of evidentiary, axiological, cognitive, contextual, or algorithmic differences, or because of mathematical properties of the problem itself; and
The inherent randomness in the AI tool, where the same task elicits different outputs across runs.
This post is mainly about the first kind, and it explores harnessing the second as a deliberate probe into the indeterminate landscape.
For human evaluators, indeterminacy means that two equally competent evaluators, working from the same evidence base, may reasonably disagree on whether a program represents (say) good or adequate value. Our interpretive frameworks, value commitments, cognitive biases and more affect what we notice and how we weigh it.
AI tools6 exhibit analogous behaviour, for different reasons. Most modern AI systems that generate text are built to be a bit unpredictable rather than perfectly repeatable every time. There’s no single “best” next word, so they pick from several likely options. In practice, built‑in unpredictability in LLM outputs is intentionally used to support creativity, adaptability, and more nuanced answers to complex questions. Sophisticated users can turn this unpredictability up or down, using settings like “temperature” and other sampling controls to make models behave in a more or less deterministic way. However, typical consumer setups keep some unpredictability in their outputs.
Language use in human communication is highly variable and underdetermined, and LLMs are generally designed to reflect that (as well as balancing other objectives like being helpful and harmless). Running the same prompt several times with default settings is likely to give you different, but still plausible, responses. This reflects both the fact that there are many reasonable ways to continue a sentence (even when the underlying meaning doesn’t change much) and the limits of what the model has captured from the training data. This range of responses may turn out to have a definable frequency distribution.
However, some of the apparent randomness may reflect an unreliable side of LLMs. When these systems confidently invent details that are not true, those errors are not random; they arise from the way models are trained to sound convincing without reliably checking facts, a behaviour often referred to as “hallucination.” The practical challenge is distinguishing informative variation from error and fabrication. Some hallucinations will occur more sporadically if we run the AI model multiple times, while reliably grounded content will tend to recur more consistently. But systematic hallucinations can also repeat. Overall, repetition can reduce the impact of hallucinations, but it isn’t a complete safeguard.
So: language is non-deterministic; evaluation is non-deterministic; people are non-deterministic; and LLMs are non-deterministic. These indeterminacies have implications for how we think about evaluation. Rather than trying to eliminate indeterminacy, perhaps we can harness it as a feature by deliberately sampling across it.
The Monte Carlo analogy
If you’ve worked in project management or economic modelling, you’re probably familiar with Monte Carlo simulation. Rather than making a single prediction about things like project completion dates or investment returns, Monte Carlo methods run a large number of scenarios, each with different input values that are drawn from the probability distributions of the input variables.
For example, we may not have a crystal ball to tell us if there will be budget overruns in a particular project, but we may know that past projects have gone over budget because of common drivers whose means and standard deviations are known. We can use these parameters as model inputs. The output of the Monte Carlo simulation is a distribution of possible project costs, revealing not just the most likely spend, but the range of possibilities and the probability of extreme events.
This principle of exploring a distribution of outcomes rather than a single forecast offers an analogy for how we might think about repeated AI-assisted evaluations.
Running an evaluation prompt through AI a thousand times is effectively conducting a Monte Carlo simulation of the AI tool’s evaluative ratings. Each iteration represents a random draw from the distribution of possible interpretations that the AI model can make from the evidence, criteria, and standards. By aggregating these results, we can estimate the central tendency (what judgement appears most often), the spread (how much disagreement exists), and the outliers (unusual but possible conclusions).
Monte Carlo simulations usually draw from explicitly specified probability distributions, whereas repeated LLM runs would be sampling from an implicit, unknown distribution defined by model weights, decoding settings, and prompt context. So in effect, the results would be more “black box” than a true Monte Carlo simulation - but could still be illuminating.
The “multiple runs” principle applies equally to human evaluation, though it’s impractical to assess something thousands of times. We can approximate this by engaging diverse evaluators or stakeholder groups - e.g., running a survey, deliberative democratic, or other process to gather evaluative judgements, as this study did.7 Each individual contributes a distinct perspective, conceptually analogous to drawing a sample from the wider population of possible human judgements about the program. In practice, a few dozen runs may suffice to reveal the main patterns, as we can expect diminishing returns from each additional run.
AI opens the possibility of simulating multiple runs, more quickly and at greater scale. However, doing a good job of this requires an appropriate process. For example, though you can ask an AI tool to design and conduct a whole evaluation with just one prompt, I really don’t recommend it. An iterative process between human evaluators, stakeholders and AI is ethically and conceptually more defensible. Also, breaking evaluation steps into discrete tasks is a more effective way to keep AI tools on track. This blog I came across recently explains it beautifully.
So, within an evaluation, we could isolate a key task and run it hundreds or thousands of times. In this post, the main example I’m using to illustrate the idea is prompting an AI tool to simulate evaluative ratings by applying a consistent rubric to a consistent set of evidence. However, possibilities abound at every stage of an evaluation - for example, prioritising and defining criteria from a value proposition, transcribing interviews, producing thematic summaries of transcripts, rapidly mapping perceived causal links into an emerging theory of change, and a lot more.
Why repeat? Is there value in aggregating and comparing multiple evaluations?
What might we gain from this approach? Running an evaluation task thousands of times could offer several advantages that we can’t get from a single-shot evaluation:
Revealing variability and consensus
A data visualisation of 1,000 evaluation runs could help us see which conclusions are robust across iterations and which are highly sensitive to small variations in framing or interpretation. For example, if 95% of runs converge on a rating, that points to meaningful consensus. If ratings are more evenly split, that is also meaningful because it highlights areas of ambiguity.
Detecting outliers
Repeated sampling could surface potential conclusions we might not have considered in a single evaluation. That unexpected interpretation that appears in 2% of runs might identify a crucial risk or opportunity that conventional analysis could miss. In complex social programs, these edge cases sometimes matter enormously.
Producing intervals or bands
By analogy with confidence intervals in Monte Carlo simulations, repeated evaluation runs can give us interval‑like summaries around our judgements. Instead of saying simply “this program is effective”, we might say “across 1,000 simulated evaluations, most ratings cluster between moderately and highly effective, with highly effective being the most frequent”.
These bands are rating distributions, not formal confidence intervals, but could well provide valuable information, helping decision‑makers visualise where judgements concentrate, how wide the plausible range is, and how much weight to place on outlier conclusions.
Reducing single-judgement bias
Any individual evaluation (whether human or AI) can be skewed by idiosyncratic factors. Aggregating many independent judgements can reduce the influence of unsystematic biases, moving us toward more representative conclusions. However, it won’t remove systematic bias; e.g., if a rubric is biased toward a particular set of cultural values, that bias will persist no matter how many times you run it.
Critically, this whole approach depends on the quality of the prompt, rubric, and evidence. Garbage in, garbage out. If poorly designed, repetition would simply multiply poor evaluations.
Practical applications and reflections
What does this mean for evaluation practice? Having a distribution of potential ratings available as an input to human judgement opens up several possibilities.
For program and policy evaluation, this approach offers a way to acknowledge and quantify the uncertainty inherent in our judgements. An evaluator could: (1) co-develop a clear rubric and gather evidence; (2) run the AI evaluation 1,000 times using the same prompt; (3) visualise the frequency distribution of ratings; (4) interpret the way ratings cluster or spread; (5) generate questions for further exploration; and (6) reach better-informed judgements, determining what the range and distribution of possible ratings means overall. This might better serve decision-makers who must act despite uncertainty.
For policy analysis, repeated rapid assessments might efficiently map the landscape of possible interpretations before investing in detailed analysis. For example, before commissioning a full evaluation, a policymaker might run 100 rapid AI assessments, looking for areas of strong convergence (needing less scrutiny) versus areas of greater heterogeneity (which warrant deeper investigation).
For evaluator training, a trainer might: (1) ask trainees to evaluate the same program twice, a week apart, and compare their ratings; (2) facilitate a workshop where trainees independently use the same rubric on the same evidence and compare results; and/or (3) get trainees to run their own multiple-run AI simulations. Experiencing variability in one’s own judgements, or seeing how different people/runs interpret the same evidence, could turn out to be profoundly educational. Discussion of why ratings diverge could help build appreciation for multiple perspectives and foster epistemological humility.
For AI-enhanced evaluation, this approach moves beyond treating AI as a virtual assistant to speed up analysis and writing tasks, and it steers away from mistaking AI for an oracle that magically delivers single right answers. It opens up a possibility for extending how we evaluate into new methodological territory. By simulating distributions of possible conclusions, 1,000 one-minute evaluations could enhance our understanding of indeterminacy by visualising the interpretive space surrounding a policy or program.
These applications share a common thread: they treat indeterminacy as information. To be clear, this proposal isn’t about removing human judgement and stakeholder participation from evaluation, but about exploring multiple interpretations to overcome one-shot bias and enrich understanding.
Bottom line
I began with a provocation: could running a one-minute evaluation a thousand times add new insights compared to conducting a single comprehensive assessment? The answer, as with most evaluation questions, is “it depends” - but I find the possibilities intriguing.
Both human evaluators and AI tools operate in spaces of indeterminacy, where multiple reasonable interpretations coexist. Perhaps we can better embrace the indeterminacy of evaluation through repeated rapid assessment as described here.
This article is written from a perspective of seeking to understand complexity and ambiguity, particularly in exploratory contexts where patterns of disagreement may matter more than a single verdict. On the other hand, in accountability and high-stakes decision contexts, more definitive conclusions may be essential. Either way, evaluators would be helped by understanding the degree of indeterminacy involved before finalising judgements.
This indeterminacy has different implications for different audiences. Some decision-makers and implementers may welcome insights into the range of ways their programs are valuable or valued. On the other hand, those who crave simple answers may find it frustrating to be given extra information about uncertainty. However, extra information doesn’t hold us back from delivering a clear verdict; rather it can enhance the clarity.
Some evaluators may reject this approach on principle, seeing it as an inappropriate involvement of AI in a human responsibility. I respect that concern and think the decision about whether to use AI in evaluation should remain with evaluators and stakeholders, not be imposed.
Rather than make final determinations for us, AI can simulate what rating distributions might look like. This encourages us to think differently about evaluation - rather than a search for a single truth, as mapping the landscape of reasonable interpretations. It means designing evaluations to surface pluralism. This mindset could foster a greater appreciation for the complexity we’re trying to understand and the intellectual humility to recognise that well-characterised uncertainty may serve decision-makers better than false certainty.
I’m flying a balloon here - let me know what you think
Is there potential in this idea? Could AI indeterminacy be harnessed to simulate human evaluation indeterminacy? Is it conceptually valid? Could it provide intel worth having?

Cautiously, I think so – provided we’re clear about what is and is not being claimed. The underlying ambiguity in complex interventions and criteria means that more than one evaluative judgement can be defensible, even with the same evidence. Stochastic variation in AI tools, under disciplined conditions (e.g., a fixed prompt, rubric, evidence base, and AI temperature settings), could give us a practical way to sample across that space of plausible judgements, much as simulation methods in risk analysis explore a range of outcomes rather than a single forecast.
I have boldly assumed that repeated AI runs and repeated human evaluations might sample from overlapping distributions of reasonable ratings. This working hypothesis would need empirical testing. The sources of indeterminacy differ: human indeterminacy arises from different values, contexts, and cognitive biases; AI indeterminacy arises from random sampling in a mathematical model. Studies could be designed to investigate how well modelled variation fits real-world variation. However, the distribution of ratings should be read as a heuristic picture of how a particular model under a particular setup tends to interpret the evidence, rather than as probabilities about the world.
The resulting distributions wouldn’t offer formal statistical analysis, but could still be a useful lens on how robust or fragile different evaluative conclusions may be, with human evaluators being responsible in the end for deciding what it all means.
The outputs are nothing more than an aid to deliberation, and the value of the analysis ultimately depends on human oversight to weed out hallucinations, sense‑check whether the central tendencies look broadly compatible with informed professional judgement, interrogate any surprising outliers, and recognise that any systematic bias in the rubric, evidence base, or model will be reflected in the resulting distributions.
A one-minute evaluation conducted 1,000 times adds up to a lot of minutes. It also requires significant computational resources and human hours for interpretation, so there’s a value for investment question here too: are the extra insights we gain worth the extra costs? This is itself an evaluative question, and the answer will depend on what decision-makers and stakeholders value and can afford.
I look forward to your thoughts.
Holiday season offer: 50% off, forever!
Evaluation and Value for Investment is a weekly blog for people who care about using evidence and explicit values to make better decisions. It focuses on practical, thought-provoking pieces that aim to be genuinely useful, interesting, and occasionally fun for evaluators, commissioners, and decision-makers.
It’s free to sign up, and most new articles stay freely available for the first two months after posting. Paid subscribers get full access to the complete archive, so they can revisit past pieces, frameworks, and tools whenever they need them.
For a limited time, you can purchase a paid subscription at 50% off the usual price. You can use this to upgrade your own access, or gift a subscription to that special evaluator in your life.
👉 Follow this link to see pricing and complete your subscription or gift purchase. If you already read the free posts and find them valuable, this is a low-cost way to support the work while unlocking the full back catalogue.
If you’re an evaluator working in a low- or middle-income country, a student, or you are between roles and the subscription cost is a stretch right now, please feel free to contact me directly. It will be my pleasure to open up complimentary access so that cost is not a barrier to using these resources.
Thanks for reading!
And thanks to Perplexity Pro for iterative review, conducted through a multi-criterial workflow I established in Labs using similar reviewing prompts to those I shared last week. All opinions, errors, and omissions herein remain my sole responsibility. I write on Substack to share my ideas, expertise, and views, retaining full creative and editorial control. Every topic, insight, and opinion presented in these articles reflects my own judgement and voice, with obsessive attention to every word.
Also see
Evaluator-free evaluation
AI can complete in one minute, a stepped evaluation design and implementation process that might take a team of humans 40-50 days. It’s good, but not good enough - needing deeper human oversight and stakeholder participation to be trustworthy.
AI-enhanced evaluation
Organising human and AI evaluation inputs around the principle of comparative advantage could help us do more, better, quicker, cheaper - making evaluation more accessible and boosting demand for our services.
AI and the Gell-Mann amnesia effect
It’s easy to spot LLM mistakes on familiar topics, but we’re more likely to trust its confident tone when working in new areas. Slick but shallow outputs, misquotes, misclassifications, and hallucinations can contaminate evaluation findings if not caught. This post offers practical strategies for evaluators to stay ahead of AI’s unreliable side.
New UK Evaluation Society AI Guidelines
These guidelines are built around four key principles: 1) transparency, accountability and competence; 2) human control and proportionate AI use; 3) active risk management and harm prevention; 4) quality assurance and verification. Includes explicit “do not proceed” scenarios and a checklist for evaluators. Check it out.
Evaluation is more than just measurement and causal inference. At its core, evaluation is about weighing information about a program/policy (e.g., its quality, success, and value) - including what matters to people - and making (or working with stakeholders to make) considered, transparent value judgements - such as judgements about how well things are going, what actions to take next, and so on. While causal inference helps us determine whether an intervention made a difference, it does not tell us whether that difference is valuable, clinically significant, appropriate, or sufficient. That step requires an evaluative judgement.
Indeterminacy refers to situations where an outcome, meaning, or process is fundamentally uncertain, ambiguous, or open to multiple valid interpretations. This could arise from limitations in knowledge, inherent properties of systems (as in quantum mechanics), or the complexity and openness of certain tasks. In the context of large language models (LLMs), indeterminacy manifests in two main ways: 1) non-deterministic output, where the same prompt can lead to different results on multiple runs due to random sampling, floating-point arithmetic, or the inherent stochasticity of their architecture; and 2) task indeterminacy, where prompts may allow multiple reasonable or “correct” completions, reflecting real ambiguity or vagueness in language and meaning. So indeterminacy in LLMs refers not only to technical non-determinism (e.g. variability in output even with identical inputs) but also to the open-endedness of language tasks.
An evaluative question can often be answered in multiple ways, with no single “right” answer. Two evaluators may reach different but equally reasonable judgements, because (for example) what pieces of evidence feel most important, how criteria are weighed, and how context is read are contestable. Even with agreed criteria and careful procedures, there can be several divergent responses that are all defensible, so the variation in judgements reflects a feature of the task rather than a flaw in the evaluator. This indeterminacy is not only practical and interpretive but also mathematical, more like under-constrained problems that allow many acceptable solutions until more and more constraints are added, than uniquely-solvable problems with a single right answer, which is why disagreement and spread are often exactly what a well-run evaluative process should be expected to produce.
When I refer to conclusions reached by AI about whether performance and value meet agreed definitions of (say) excellent, good, or adequate, I am being careful not to call the AI’s conclusions “judgements”. I prefer to view judgement as a fundamentally human trait. I’m using “ratings” instead, till I think of a better word.
Throughout this post, I use the term “AI” for simplicity. But these systems aren’t “intelligent” in the way we apply this term to humans. The AI tools most of us are using are more accurately referred to as Large Language Models (LLM), and for the most part, that’s what I’m talking about in this post. However, AI also includes computer vision, speech and audio processing, robotics and more. Any of these could have useful applications in evaluation, but here I’m writing mainly with LLMs in mind. LLMs work by analysing huge volumes of text and predicting the most likely next word or phrase based on patterns in their training data. They don’t possess understanding, consciousness, or intent. While some newer AI models incorporate forms of “reasoning”, it’s still fundamentally different from human reasoning. They follow statistical and pattern-based processes, and their reasoning is limited to structured tasks and information present in their training data. They can’t comprehend, reflect, intuit, empathise, or apply judgement as people do. This can be a strength and a weakness, depending how you choose to use the technology.
Congratulations to Heidi Peterson on the publication of this hot-off-the-press article in Evaluation, which makes important contributions to how we think about democratic deliberation in VfM and how value is created in complex systems; and to Carolina Mejia Toro and colleagues on the recent publication in BMC Public Health of this participatory VfI evaluation of Ka Ora, Ka Ako, New Zealand’s original (pre budget cut) free, healthy school lunches program, which showcases a multi-stakeholder approach to eliciting and aggregating multiple evaluative judgements.






Let's set up a 'good enough' case study to demonstrate this. Create synthetic values inquiry data about hypothetical evaluation participant value frameworks, apportion their respective operationalized criteria and standards of merit (or set up 333 runs for three paradigms with respective criteria and questions). Develop synthetic performance data corresponding to dimensions of merit collected in theory by different methods (that would hold varying degrees of credibility of evidence types). Provide different modes of evaluative reasoning and synthesis options to be selected based on varying participant values.
In short, identifying the variable inputs for simulations would be the main task. This almost sounds like a task for an orchestra of AI agents.
Interesting! In comparing it to Monte Carlo simulations, it should be noted that those simulations allow for two types of uncertainty; the inherent randomness in the environment, reflected in the use of quasi-random numbers in the simulations; and uncertainty in the input parameters, which can be explored by varying those parameters in a controlled way. (aka sensitivity testing). The problem with the 1000 AI queries is that it is not obvious whare the variation comes from. Is it possible to 'look under the hood' and get at least an impression of what drives the variation in AI results? If we do not feasibly know, then the controllability available in the Monte Carlo situation is just not there.