In evaluation, one of the hardest steps is getting from an eclectic mix of evidence and perspectives to a defensible overall conclusion that is more robust than a finger-in-the-air opinion. This post explores that “synthesis problem” by comparing two practical approaches, multi-criteria decision analysis and rubrics. Both are legitimate ways of guiding evaluators through the synthesis step, so long as the process makes values visible, centres the right voices, navigates value tensions, and supports value judgements with clear rationale.
Methodological absolutism has been on my mind lately. A recent LinkedIn exchange reminded me how easily any of us can fall into implying there is one “right” way to do evaluation, even when we know better. Rubrics are often my preferred tool to support the synthesis step in evaluative reasoning, but they are not the only valid way.
This post unpacks how multi-criteria decision analysis and rubrics sit inside the same evaluative logic, and when each may be more or less helpful.
Quick recap: there’s a general logic involved in evaluating, whether explicit or not
Underneath any evaluative process is a general logic of evaluation: judging how good something is based on evidence, criteria (what aspects of it matter), and standards (what good looks like). Whether you’re choosing a mango at the market or reaching consequential conclusions about social programmes, you’re relying on some form of this logic - whether you know it or not (this point is the focus of a previous post, freely available here).
In policy and programme evaluation, where people’s lives and rights are affected, it is a good practice to make the logic explicit, defining criteria and standards that make judgements transparent and grounded in the values of people and groups who have a stake in the intervention. That is one area where I would argue a bit of “absolutism” is warranted: Explicit evaluative reasoning is, in my view, essential.1
There are many ways to implement this logic. I’ve written before about a range of alternatives, their situational strengths, and advantages of combining them (“mixed reasoning”). Here, I will deep-dive into two contrasting approaches to handling the synthesis step: multi-criteria decision analysis (MCDA), and rubrics.
Two layers: aggregation and sense-making
The process of bringing together criteria, standards, and evidence to make evaluative judgements is called the synthesis step. When I think about this process, I see two distinct layers at play:
Aggregation: how the elements are combined
Sense-making: the cognitive and social processes through which people interpret evidence in light of values, negotiate tensions and ambiguities, and reach judgements.
Both MCDA and rubrics are tools for organising how multiple and often diverse criteria, standards, and pieces of evidence are brought together - the structure they bring supports aggregation. Sense-making, by contrast, is embedded in the design work (choosing and defining criteria and standards) and in the judgement work (interpreting what the synthesis means). The sense-making step is led by human brains, not technocratic tools - though the tools support it, and the choice of tool can affect how and what judgements are made.
MCDA in plain language
In policy and economics circles, multi‑criteria decision analysis (MCDA) usually means some form of what Michael Scriven called Numerical Weight and Sum (NWS), rooted in multi‑attribute utility theory (MAUT)2 and decision analysis.3 MCDA as a field is broader than this, including approaches that work with qualitative judgements as well as numerical models, but here I focus on a simple, NWS-style multiply-then-add pattern that is common in policy work. In its everyday form, it typically follows a process like this:
Define options to compare
Specify a finite set of criteria (for example, equity, efficiency, sustainability)
Assign numerical importance weights to each criterion (e.g., equity gets 40% of the overall score, efficiency 30%, etc.)
Score each option on each criterion
Multiply scores by weights and sum to get a composite score
Rank options according to their composite scores.
For example: imagine health ministry policy analysts need to compare five mental health initiatives competing for the same pot of funding. They agree a small set of criteria such as reach, impact, equity, and implementation risk, then work with stakeholders to derive or assign importance weights to each criterion (for example, through structured preference-elicitation or more informal deliberation). Each initiative is scored on every criterion using available evidence and expert judgement. The scores are converted to a 0-100 scale, multiplied by the weights, and summed to give a composite score for each option. Options are ranked according to their composite score. The resulting ranking helps decision-makers see which package of services looks strongest overall given the stated value trade-offs, but they still discuss the numbers, test how sensitive the ranking is to different weights, and may override the top-ranked option if additional nuances come into play.
Although designed primarily for ex-ante (future-facing) decision-making when choosing between multiple options, the same machinery can be applied ex-post (evaluating what has already happened) to a single programme by scoring it relative to some “ideal” score instead of against competitors. In essence, the “standards” part of this process is an expectation about how well the actual score should stack up against the ideal score.
The appeal is obvious: NWS is systematic, transparent in its arithmetic, and allows sensitivity analysis to test how much conclusions depend on the chosen weights and scores. However, for all the numbers, NWS doesn’t quantify the “right answer” for us - that’s the judgement step. NWS assists in aggregating criteria, standards, and evidence, paving the way for clear deliberation and judgements.
In fact, judgement is needed at several places in this process:
Choosing criteria
Setting weights
Deciding what scores the evidence justifies
Judging the nature and extent of uncertainty and risk, and how that should be incorporated in the assessment
Interpreting whether the resulting composite score represents acceptable, good, or excellent value in the context; the score on its own doesn’t say, for example, whether we should invest, scale, adapt or exit, though it can help us decide.
The broader MCDA field includes more sophisticated approaches than the simple NWS approach I have described above. These include methods where stakeholders’ preferences are elicited in a structured way (using tools such as MAUT‑based protocols or discrete choice experiments4), turned into individual value models, and then aggregated to show how different stakeholder groups would rank options under their own value structures. Essentially, this is akin to each stakeholder providing their own evaluation, and then mathematically aggregating the individual evaluations to help determine what it means overall. In this way, MCDA can provide a robust and traceable representation of where stakeholders’ values sit in the aggregate.5
However, in the real world, weights and scores are not always elicited empirically. Sometimes they are pulled out of the air. When this happens, it gives NWS a veneer of “objectivity” that hides arbitrary assumptions. This issue is not with NWS or MCDA as such, but with weakly justified weights and scores - a form of “bad MCDA” (a label that could equally apply to rubrics if somebody sat at their desk and ChatGPT’d them out of the sloposphere without rigorous stakeholder involvement). Without an empirical basis for its numbers, NWS can give a false sense of precision and technocratic solidity that is not warranted by the underlying reasoning.
Scriven (1991) also argued - damning with exquisitely faint praise - that although the common additive application of NWS is “sometimes approximately correct, and nearly always clarifying” (p. 380), it can mislead. One concern is that when many relatively minor criteria are included, their combined weighted scores can drown out a small number of critical criteria, so that decisions end up favouring secondary or tertiary objectives at the expense of the primary purpose of the intervention. He also highlighted problems arising from treating all values as continuous variables, when in practice some criteria function as minimum requirements or “bars” that should exclude options regardless of how well they perform elsewhere. A further issue is the assumption that utility increases in a straight line across the full range of each variable, when in reality the value of an extra unit often changes depending on where you are on the scale.
The upshot of these limitations is that NWS does not always work as well as intended. Jane Davidson (2005, p. 177) put it in a way that perfectly aligns with my early experiences of using the approach in public policy work:
The reality is that although NWS seems simple and intuitive, it can often leave the evaluation team looking at a conclusion that does not seem quite right. The temptation at that point is often to fiddle with the numbers to see whether the right answer can be coaxed out of the data. An alternative is to work with a synthesis strategy that incorporates the key elements of how the human brain naturally weights considerations, making them explicit so that they can be applied to larger numbers of dimensions.
This is the intuition that underpins qualitative and rubric-based synthesis.
Qualitative Weight and Sum and rubrics
Michael Scriven’s alternative is “Qualitative Weight and Sum” (QWS). Instead of assigning precise numerical weights to each criterion, QWS:
Uses a small set of ordinal importance categories (for example, essential, high, medium, low), reflecting Scriven’s view that “beyond this very modest level, validity in allocating utility points is hard to justify” (1991, p. 294).
Synthesises evidence qualitatively within categories, without aggregating across fundamentally different importance levels
Treats some criteria as bars or minimum requirements that cannot be compensated by strengths elsewhere.
In practice, Scriven argued that qualitative weighting and synthesis should be the default in most evaluation contexts, with numerical weighting reserved for relatively narrow conditions where its assumptions can genuinely be warranted.
Evaluative rubrics provide a practical framework for implementing the general logic of evaluation through qualitative weighting and synthesis; they operationalise that logic through qualitative descriptions and ordinal levels of performance for each criterion. A rubric typically specifies:
Criteria of merit, worth, or significance
Standards (a few performance levels - for example, excellent, good, adequate, poor)
Descriptors of what performance looks like at each level for each criterion (Davidson, 2005).
Rubrics can accommodate both qualitative and quantitative evidence, and they express weighting ordinally - e.g., good is better than adequate. A little counterintuitively, adequate is more important than good, because the adequate level stipulates must-haves, whereas good describes additional features of performance. In keeping with this, descriptors linked to adequate performance can be treated as non‑negotiable, while good and excellent describe attributes that build on the basics.
Importantly, rubrics are especially well‑suited to participatory development and use: the process of co‑designing criteria and standards surfaces stakeholders’ diverse values, supports a shared understanding of what “good” means in context, and sets up a transparent basis for later synthesis, deliberation (including constructive disagreement), and judgement.
How different are rubrics and NWS really?
At a conceptual level, both NWS and rubric‑based synthesis are expressions of the same underlying evaluative logic: they apply explicit criteria and standards to evidence, and the resulting synthesis helps evaluators reach judgements. The differences are mainly:
Structure: NWS uses cardinal numbers and arithmetic, rubrics use ordinal categories and qualitative description.
Treatment of importance: NWS assumes commensurable weights that can be multiplied and summed; rubrics treat some criteria as non‑compensable (i.e., bottom lines where trade-offs would be inappropriate) and only aggregate by contextual choice within qualitative groupings.
Typical style: Both approaches combine technocratic structure with human deliberation; NWS tends to lean technocratic while rubrics lean deliberative but both can be used in either mode.
From a terminological perspective, “MCDA” can be understood broadly as any approach that uses multiple criteria to support decisions, in which case both NWS and rubrics fall under that umbrella. In practice, however, economists and many policy analysts use MCDA synonymously with NWS approaches. This suggests a pragmatic approach to language: use the terms that resonate with your audience, but understand how each works and be explicit about whether you are talking about numerical or qualitative weighting and synthesis.
When I’m introducing rubrics to economists, I find it helpful to introduce them as “a qualitative form of MCDA”. This description fits rubrics into a family of approaches that is already widely accepted and used in economic analysis.
Value for Investment isn’t exclusively rubric-based
The “synthesis problem” (or aggregation problem) is the challenge of pulling together multiple criteria, standards, and streams of evidence into coherent evaluative conclusions. It becomes more acute as:
The number of criteria grows
The diversity of stakeholder perspectives increases
The evidence base becomes more mixed (e.g., the breadth of quantitative and qualitative methods used)
The evidence base becomes less certain (please enjoy my soapbox moment in this footnote6).
Value for Investment (VfI) was developed to answer evaluative questions about whether a policy, programme, or other investment is good use of resources. Answers to that question must be grounded in explicit evaluative reasoning, incorporating criteria that represent what relevant people value, standards of “good” resource use, reliable evidence, synthesising it all and arriving at warranted judgements.
NWS and rubrics are both viable synthesis methods in evaluation. They can help structure the synthesis process and facilitate well-formed deliberations and judgements. As long as the aggregation and sense-making processes are explicit, well-designed and well-conducted, either NWS or rubrics can meet the minimum specification for explicit evaluative reasoning.
In VfI, evaluative reasoning provides the means to combine economic and evaluative insights, qualitative and quantitative evidence, diverse stakeholder values, technocratic and deliberative inputs, and different kinds of evidence with differing levels of certainty, in a way that remains transparent and challengeable. NWS and rubrics can both support the synthesis step within this larger reasoning process; either can be a valid choice but neither is “the evaluation” in its entirety.
The tool isn’t the judgement
It is tempting to look for a tool that “does the judgement” for us. But that’s not how this works. People make judgements, not tools:
Tools like NWS or rubrics help structure and document the synthesis step
Judgement happens in design (choosing criteria and standards), in interpretation (making sense of ratings and scores), and in deliberation (balancing considerations and reaching conclusions).
Whether using NWS, rubrics, or another approach, the key is that values are collectively determined, explicit, and open to scrutiny. NWS supports this when weights and scores are empirically elicited and discussed. Rubrics support it by making criteria and standards co-defined and visible, and by inviting stakeholders into a deliberative process.
Using numbers does not remove subjectivity; it relocates it into the choice of criteria, weights, and scales, just as rubrics locate it in how criteria and performance levels are defined. Either way, the art and the science is to make values explicit and co-create a shared subjectivity that is used transparently to support defensible value judgements.
Crucially, the output of a synthesis method (whether a number, a rating, or a set of narrative findings) is not the judgement. It is a structured input to human judgement, which also takes into account uncertainty, debate, contextual nuance, professional and lived experience, and the social and political realities around the investment.
Let me illustrate this with a quick example. We could work with stakeholders to surface what we might call first-order values – for example, when it is agreed that an evaluation should focus on accessibility, acceptability, and equity of outcomes. It is often helpful to go further and define second-order values – for example, we might debate what accessibility means in our context, and land on an agreed definition of good accessibility, such as “a clear majority of service users, including those with specific accessibility needs, report that they find the services to be accessible”. Those definitions might make it into the evaluation framework.
A third-order set of values comes into play when the evidence comes in, and we have to deliberate on whether, say, 80% of service users overall and 51% of the deaf community finding services to be accessible should be judged “good” or not. Criteria and standards are not intended to set out values algorithmically so that a machine could make judgements for us. They are a scaffold and a focusing tool to guide those deliberations.
When to lean toward NWS
NWS is most useful when certain contextual and quality conditions hold:
A few key criteria, which can be meaningfully quantified and measured on commensurable scales7
There is a strong empirical basis for deriving numerical importance weights, such as stakeholder preference studies, utility elicitation exercises, or robust deliberative processes that explicitly explore trade‑offs
There are multiple options to compare, and decision‑makers need a transparent, reproducible way to rank them (for example, in procurement, or portfolio prioritisation)
The evaluation context allows for and values sensitivity analysis to understand how fragile results are to plausible changes in weights and scores
Transparent acknowledgement that the formula is just an input to deliberation, not decision rules or a replacement for judgement.
In such contexts, NWS‑style MCDA can be a powerful way to structure options analysis, highlight the implications of value judgements embedded in weights, and surface where additional data or deliberation would make the most difference.
When to lean toward rubrics
Rubrics become particularly attractive when:
The evaluation needs to be transparent and communicable to non‑technical audiences, including communities and frontline practitioners
There are minimum standards or ethical “bars” that should not be traded off against other gains (for example, basic safety, rights protections, minimum equity thresholds)
Evidence comes from mixed methods, and it would be misleading or opaque to force everything into a composite numerical score
Criteria are deeply intangible (e.g., equity, cultural safety, dignity)
It is appropriate for judgements to be reached through collective stakeholder deliberation - e.g., criteria are contested, stakeholders have diverse worldviews, there is strong interest in co-defining what ‘good’ looks like.
Rubrics excel at supporting democratic deliberative sense‑making by providing a shared language for discussing performance, priorities, and disagreements. They are well‑aligned with VfI work that aims to blend economic criteria (for example, cost‑effectiveness) with other dimensions such as equity, sustainability, and cultural significance, within a single evaluative framework.
When either could work (and how to choose)
There are many situations where either NWS or rubrics could work tolerably well. In these “either/or” cases, questions that can help guide the choice include:
What kind of conversation do you want to open up? If the aim is to support participatory, values‑rich deliberation, rubrics can create room for free-flowing dialogue. If the immediate need is to produce a ranked list under time pressure, or a more structured empirical approach to values-elicitation is desired, NWS may be more acceptable to decision‑makers.
How much precision can you legitimately claim? Where the evidence and value base are inherently approximate, uncertain, or indeterminate, ordinal categories and narrative explanations may be a better fit than weighted numerical scores.
How important is comparability or aggregation across a large portfolio? If you need a consistent, portfolio‑wide way to aggregate ratings across many units, NWS can help, provided the underlying weights are justifiable (even defaulting to an unweighted aggregate score is a normative choice). Rubrics can also support this through shared frameworks and meta‑synthesis, albeit in a more qualitative form.
Where will deliberation actually happen? If the key deliberation is at the design stage (co‑creating criteria and standards) and at the sense-making stage (weighing diverse criteria, standards, and pieces of evidence, and developing value), rubrics are a natural fit. If deliberation is more likely to happen in between design and sense-making (eliciting value judgements from stakeholders as part of an evidence gathering process), or more likely to focus on quantifiable trade‑offs among options (for example, budget allocations), NWS can scaffold these kinds of conversations.
In some evaluations, a hybrid approach may be considered - for example, using a set of rubrics to make ratings against criteria such as relevance, coherence, sustainability, etc, and then using NWS to aggregate the ratings and help make an overall judgement against the criteria collectively. Or, using rubrics as the overarching framework and embedding NWS‑style analyses as one stream of evidence for particular criteria (such as efficiency or cost‑effectiveness), while keeping the final synthesis qualitative and deliberative.
Technocratic and deliberative uses of the same tools
Neither NWS nor rubrics are inherently technocratic or inherently deliberative. Both can be used:
Technocratically, as if the formula or rubric is a decision rule; or
Deliberatively, as a structured aid to collective reasoning.
That said, NWS tends to be associated relatively more with technocratic decision‑making, particularly when guidance is written for “decision‑makers” and focuses on the mechanics of weighting and scoring. Rubrics tend to be associated more with deliberative, participatory processes, especially when used to collaboratively design criteria and standards with communities and other stakeholders. In both approaches, the deeper question is how we are using them - as rigid decision rules or as aids to explicit, defensible sense-making.
My applications of VfI have tended to push toward the deliberative end of the spectrum, foregrounding considerations such as whose values are represented, how credible the evidence is, and how judgements are warranted in ways that are acceptable to diverse stakeholders. In this space, rubrics offer some natural advantages, but NWS can still be implemented in participatory ways - for example, by eliciting weights from multiple groups and presenting separate “accounts” for each.
Bringing it back to methodological absolutism
Returning to methodological absolutism, the key messages are:
The logic of evaluation is always there (you can make it transparent or opaque, but you can’t make it go away) and evaluations are more credible when we make it explicit
Criteria and standards, grounded in stakeholder values, are essential in serious evaluation (and therefore in VfI)
The choice of synthesis tool is contextual.
Michael Scriven himself was wary of methodological absolutism about specific tools or designs, but he was relatively uncompromising about the general logic of evaluation: using explicit criteria and standards to reach warranted judgements from evidence. My own view adds, following Thomas A. Schwandt and co-authors over the years, including Robert Stake and Emily Gates, that this logic is enacted through social and political processes, so any method also has to be judged by whose values and voices it honours.
NWS and rubrics are both valid strategies for addressing the synthesis problem, each with situational strengths and limitations. Both can be implemented well, with carefully elicited values and transparent reasoning, or poorly, with arbitrary weights or descriptors that create a misleading veneer of “objectivity”. The important questions are not so much about which tool is “right” but what kind of sense‑making is needed in a context, and how we can design the synthesis step so that it supports sound, explicit, inclusive, and shared judgements about value.
The final responsibility in evaluation lies not in the tool but in the people and conversations around it: whose values are visible, whose voices are heard, how value tensions are navigated, and how well the final judgements can be defended as logical, transparent, and fair.
How do you make your evaluative reasoning explicit?
Thanks for reading!
And a big thank you to Dr Jay Whitehead for helpful peer review. All errors, omissions, and opinions in this piece are my sole responsibility.
If this was useful, a quick tap on the ❤️ helps me know it landed - and it also nudges Substack to show the piece to more people who might benefit from it. Thank you.
In practice, some evaluations do not conform with the identity of systematically making value judgements. For example, some studies labelled as evaluation are framed more narrowly around “did it work?” type questions - such as estimating whether and to what extent welfare benefit sanctions incentivised return to work. On the face of it, this might not look like an obvious case of judging merit, worth, or significance, but about estimating an effect. But even here, evaluative questions may be lurking in the background, like, “is this a meaningful result?” or “did it work well enough to be worth doing, given its costs, side-effects, and alternatives?” There is value in making those judgements explicit rather than leaving them implied or unexamined.
Multi‑attribute utility theory (MAUT) is just a formal way of doing something humans already do: making trade‑offs between things that matter. Economists use it to build a “preference model” of a decision‑maker, based on questions like, “Would you rather have a bit more of A or a bit more of B?” or “How much extra of A would compensate for a loss in B?”. From those trade‑off judgements, they construct a mathematical function (a utility function) that assigns a number to each combination of attributes, so options can be compared and ranked. In practice, MAUT gives structure to conversations about value and compromise: it turns many small, intuitive judgements about “what matters more, and by how much?” into a coherent scoring system, without claiming that the numbers themselves are objective facts.
Decision analysis is basically the general logic of evaluation applied to choosing between options under uncertainty. It follows the same steps: clarify what matters (objectives and criteria), articulate what “good” would look like on those dimensions (preferences or standards), describe how each option is likely to perform (evidence about consequences and risks), and then reason explicitly to a conclusion about which option is better, given those stated values. The tools that make decision analysis look technical – decision trees, probabilities, utility or value functions – are just ways of structuring this evaluative reasoning so that trade‑offs, assumptions, and uncertainties are laid out clearly and can be discussed, rather than leaving them as hidden, implicit judgements.
Multi‑attribute utility theory (MAUT)‑based protocols and discrete choice experiments (DCEs) are structured ways of turning people’s gut feelings about trade‑offs into usable numbers. In a MAUT protocol, stakeholders are taken through a series of carefully designed “what if” questions (for example, “How much more of attribute A would make up for losing some of B?”) so their underlying value trade‑offs can be mapped into a utility or value function that scores options on multiple attributes. DCEs do something similar by asking people to choose repeatedly between pairs or sets of hypothetical options that differ on key attributes; from those repeated choices, you can infer how strongly people value each attribute relative to the others. In both cases, the goal is not to manufacture objectivity, but to make stakeholders’ preferences and trade‑offs explicit, consistent and available as an input to evaluative reasoning or MCDA, instead of leaving them as unexamined impressions.
A similar idea can be implemented with rubrics, by asking different stakeholders or groups to apply a rubric and then comparing or aggregating their evaluations. This connects with the broader possibility of running many evaluations to explore the interpretive space and how robust conclusions may be under different reasonable value judgements.
This footnote is a repeat, but it’s worth saying again for new readers. I’m endlessly fascinated with humanity’s tendency to equate evidence with “facts”. In evaluation, a lot of the evidence we have to consider is uncertain. For example, real-world evidence includes probabilities, tendencies, theories, hypotheses, propositions, interpretations of incomplete or ambiguous data, modelled scenarios based on constellations of assumptions, expert opinions, educated guesses, lived experiences, and endeavours to synthesise and communicate complex issues in digestible terms - all of which involve human judgement. Even “hard” facts can change: they remain open to revision as new evidence emerges (because science). No matter what methods are used, facts are not neutral or value-free - they’re influenced by values and beliefs. For example, people make decisions about what counts as credible evidence, what data to collect, how to collect it, how to analyse it, and how to interpret it. All these methodological decisions are guided by values (whether declared or not). The reporting of facts, too, reflects values and biases; what is considered a fact is influenced by what is deemed important or relevant by the people presenting it. Any time we look at a fact, we could think about where it came from and what biases may lurk beneath its surface. Examples include measurement bias, sampling bias, selection bias, recall bias, confirmation bias, and publication bias. And that’s when everybody’s on their best behaviour. Unfortunately, purveyors of evidence are not immune to racism, sexism, pecuniary interests, career ambition, politics, social conformity, noble lies, dogma, hubris and bullying. These human frailties can affect the veracity of any facts, which is why evaluators attend carefully to method, bias, accountability, and evaluative thinking.
To unpack the requirement for NWS that “criteria can be meaningfully quantified and measured on commensurable scales”: First, it is essential that you can sensibly put numbers (or at least well‑defined rating levels) on performance for each criterion, so that “more” and “less” really do reflect better or worse in a way stakeholders recognise. Second, once you have done this, the different criteria must be on value scales that can be compared and traded off inside one model – so that a change on Criterion A and a change on Criterion B are both expressed in a common “currency of preference”, allowing you to weight and add them without pretending they are literally the same thing.



Hi Julian...here is my discussion of rubrics vs more weighted criteria based approaches. https://mandenews.blogspot.com/2020/04/rubrics-yes-but.html
The Basic Necessities Survey BNS, mentioned in above, generates weighs democratically
Great summary Julian. Keen to understand how this approach may be used to support a systems strategic performance framework? Where some standard system metrics are routinely collected and could be grouped together in a way to reflect performance of an aspect of performance in a larger system but where we might also know (or assume) there are other criteria (known and unknown) for which there is no current metrics data readily available. Because science evolves as we discover more about it like you say. Perhaps worth a conversation when you are next in town? Hope you and yours are well.