Douglas Adams had some keen insights on bureaucratic absurdity that resonate with me as an evaluator. A favourite set of books from my youth, The Hitchhiker's Guide to the Galaxy offers all sorts of lessons for anyone working in evaluation and value for money (VfM) assessment. Here are ten reflections. A mix of light-hearted and more reflective takes, each points to a real evaluation challenge.
If you haven’t read the books, don’t panic: each point below uses a short scene or idea from the series as a jumping-off point for a real-world evaluation issue.1
1. Beware of the Leopard: transparency matters
In the novel’s opening scene, Arthur Dent lies down in front of a bulldozer to block the demolition of his home. A council officer tries to persuade Arthur to move, arguing that the plans for a new road through the property have been on display for months, to which Arthur memorably responds that the plans were displayed:
…in the bottom of a locked filing cabinet stuck in a disused lavatory with a sign on the door saying Beware of the Leopard.
#BewareOfTheLeopard remains a popular hashtag because it so aptly skewers bureaucratic opacity.
Evaluations and VfM assessments suffer from endemic Leopard problems. For example, evaluation reports may be written but unpublished, published on a website without telling anyone, written in inaccessible technical language, or may require the skills of a forensic methodologist to avoid being misled.
Meaningful transparency is more than just sharing accessible reports; Leopard problems can occur at any stage of an evaluation. For example, commissioners may stipulate methods without checking whether they meet stakeholders’ needs, consultations may be held in ways that systematically exclude marginalised voices, evaluators might interpret data exclusively on the basis of their own understanding, biases and assumptions without engaging affected people.
Let’s all beware of the leopard problems that can pop up in evaluation.
2. The universal translator: evaluators are Babel Fish
The small yellow Babel Fish was very useful. When placed in your ear, it would feed on brainwave energy and instantly translate any language into one you understand - a bit like ChatGPT today, only cuter.
Evaluation and VfM sit at the junction of multiple languages and cultures, including theoretical and methodological jargon, operations-speak, policy, politics, lived experience, local languages and cultures. Without effective translation, evaluation becomes an exercise in exclusion or speaking past one another. For example, when economists talk of allocative efficiency or net present value, can communities see what it means for their lives? When service users speak of dignity, self-determination, and cultural fit, do funders understand how to operationalise these principles and bring them to life?
Many evaluation terms do not travel well across languages or cultures, at least not by literal translation alone. Words like “evaluation,” “evidence,” “rigour,” “participation,” “impact,” and “value for money” are bound up with particular histories, power relations, ways of knowing. When moved into other linguistic and cultural settings, they often need reinterpretation rather than assuming direct equivalence. For example:
Evaluation might sometimes move closer to collective reflection or ‘checking our path’
Evidence may expand to include story, land, spirituality, and lived experience
Rigour could shift from methodological purity to being faithful and accountable to relationships and deliberation
Participation may expand from answering questions to making shared decisions
Impact may have less to do with effect sizes and more to do with surfacing changes that communities recognise as movement toward wellbeing
VfM might be bigger than efficiency, encompassing wider questions about justice, reciprocity, stewardship, sovereignty, responsibilities to community and ancestors.
These conceptual shifts make translation much more than a technical language task - it’s a political and ethical one, requiring evaluation teams to include competent cultural interpreters and avoid inappropriately exporting conceptual toolkits that don’t fit. For a concrete example, see this reflection.
Being a Babel Fish is an essential part of everyday evaluation practice, moving back and forth between diverse worlds of understanding, connecting dots between:
Criteria that get to the heart of what good value looks like to affected communities, in terms that also resonate with decision-makers (e.g., justifying relationships and trust as core aspects of efficiency)
Statistical and economic findings that relate abstract concepts like counterfactuals, compensating variations, generalisability and uncertainty with what actually matters to stakeholders
Complex mixed-methods evidence, with complementary needs for technical rigour, inclusive rigour, ethics, and practical utility to produce actionable insights.
The Babel Fish role is acutely important when Indigenous evaluators carry the responsibilities of bridging local languages, values, ways of knowing and being, with the expectations and machinery of Western governments or international donors, working “with a foot in both worlds” to facilitate mutual understanding and positive impact.
Evaluation requires co-learning - reciprocal knowledge exchange, and stepping back to question our assumptions (triple-loop learning - e.g., “are we asking the right questions?”). This can’t happen effectively without continuous Babel Fish translation that not only respects but also deeply understands different languages, values-systems and ways of knowing.
3. Don't be a telephone sanitiser: make VfM count!
The cautionary tale of the telephone sanitisers from the planet Golgafrincham is a metaphor for too much of what passes for VfM assessment today, as I detailed in a post last December dedicated to this theme.
The Golgafrinchans, convinced that people in certain occupations like telephone sanitisers, advertising jingle writers and TikTokers added no value to society, shipped them off to colonise another planet. The irony was that those left behind all died from a virulent disease contracted from a dirty telephone.
The moral, of course, is to be resolutely useful. The moment VfM assessment becomes disconnected from real purpose, it turns into busy work that serves nobody. Too often, VfM evaluations are treated as exercises in compliance or defensive wagon-circling rather than crucial opportunities for learning, improvement, and good decision-making. Reports that sit on shelves, tick boxes, measure the wrong things, or consume resources without creating insight, help nobody.
The most useful VfM evaluation provides critical information that decision-makers actually need. What decisions does this evaluation inform? Who uses the findings, and how? What will change as a result? If these questions can’t be answered clearly, the evaluation team risks being as dispensable as a telephone sanitiser.
4. Vogons and bureaucratic decision-making
Making VfM evaluations useful also depends on decision-makers holding up their end of the bargain. Vogons epitomised the perils of taking procedural machinery to its most absurd extreme. Not truly malevolent, Vogons were relentlessly officious, joyless administrators who ran the galaxy's most impenetrable administrative labyrinth.
We’ve all had to deal with Vogons from time to time - right? For example, in a recent house-building project it took a full year of industry professionals passing bits of paper around before we could start digging the foundations. The Dilbert Principle says that the time required for a decision to be made is two weeks multiplied by the number of in-trays it has to pass through, so presumably the various forms and plans and consultant reports had to traverse something like 26 Vogon in-trays at our local authority.
Even the best evaluations can’t overcome a system that elevates process over purpose. If every decision is governed by filing the right forms and waiting in queues, then actual improvement is stifled, innovation is crushed, and VfM is defined by compliance at the expense of impact and value. No matter how good we get at VfM assessment, our work achieves little unless decision-makers resist bureaucratic, risk-averse Vogon tendencies and instead focus on learning and meaningful impact.
5. Deep Thought's “42”: the danger of answers without questions; numbers without context
Millions of years before the main events covered in The Hitchhiker’s Guide story, the book explains how a race of hyperintelligent pan-dimensional beings built a supercomputer called Deep Thought, commissioning it to compute the Answer to “the Ultimate Question of Life, the Universe, and Everything”. After 7.5 million years of contemplation and calculation, Deep Thought (voiced by Helen Mirren in the 2005 movie) said it had the answer, “but you’re not gonna like it… 42”.
Once the confusion and disappointment had worn off, it dawned on them that they had never actually clarified what the Ultimate Question was. It’s hard to provide clear answers without clear questions.
So Deep Thought proposed to design an even greater computer, whose intricate program would, over millions of years, uncover the precise question for which “42” was the answer. That supercomputer was the planet Earth.
“42”, to me, is a perennial indicator problem in a nutshell. Indicators can tell us something important, but only if they address a question someone was actually asking, and only if we know how to interpret them. I can calculate for you that the average cost per program participant was $42 - but without context, how can you judge whether that’s high, low, or just right? The other program down the road might deliver its stuff for $38, but how good is its stuff? $84 might be better than $42 if it creates more than double the value. Even a benefit-cost ratio can mislead if it doesn’t measure and monetise things that matter.
If we don’t define clear questions, and clarify what good value looks like in a context, value risks being reduced to whatever is easiest to count or measure instead of what matters. No number is an answer to a VfM question. It’s just a piece of evidence. At its best, VfM assessment becomes a process of meaning-making. It invites reflection, learning and improvement, acting as an intervention to help stakeholders make sense of performance, context, perspectives, and trade-offs.
6. Philosopher Unions: focusing inward at the expense of broader connections and relevance
If Vogons illustrate bureaucratic excess, Philosopher Unions show the pitfalls of guild-like insularity. The philosopher unions protested anything that would diminish their authority or resolve big questions too efficiently. Their objection to Deep Thought was not concern for truth, but fear that an answer would render them obsolete, undermining their professional raison d’être.
Could there possibly be a cautionary parallel for the evaluation field here? Like the philosopher unions, evaluators sometimes take a dim view of those perceived as outsiders - be they economists, AI tools and experts, data scientists, impact measurement specialists, or other fields whose methods might impinge on evaluation’s established ways and sense of jurisdiction.
Setting standards and boundaries, defending rigorous, ethical and socially just evaluation are important. But the field can become so wrapped up in internal argument and boundary-policing that it forgets to present a united, compelling case to the world for why evaluation matters. In today’s world, we need that united voice more than ever.
We evaluators rightly celebrate our pluralism and debate. With over 100 recognised approaches (see Michael Quinn Patton’s 13-minute video) grounded in different purposes, theories, methods, and values, the everlasting discourse about what constitutes good evaluation (to which this blog aims to contribute) is a strength of our profession.
Pluralism is a strength when it drives intellectual rigour and innovation, but it’s a weakness if it fragments the profession’s external voice or leaves critical audiences puzzled about evaluation’s value. Evaluation is “the largest profession no one has heard of” (as John Gargani memorably put it) partly because so much of our attention is absorbed by debates within the guild, and too little on clearly communicating our value or advocating for our critical societal role.
7. Infinite Improbability, imperfect measurement, and multiple perspectives
Adams' Infinite Improbability Drive harnessed the laws of quantum probability to allow a spaceship to travel to every conceivable point in the universe almost simultaneously, making travel across vast interstellar distances instantaneous - and unfortunately (but hilariously), highly unpredictable.
The Improbability Drive offers a crucial lesson about measurement limitations. Evaluation often proceeds as if precise quantification is not only possible but necessary for credible analysis. Yet many important aspects of public value, like dignity, community cohesion, or democratic legitimacy, resist easy measurement and monetisation.
Even if something is hard to measure, there can be immense value in describing it, perhaps even estimating its value and conducting statistical or sensitivity analysis to understand the implications of uncertainty. We may not be able to divine “one right answer” but, as statistician John Tukey famously wrote: “Far better an approximate answer to the right question, which is often vague, than an exact answer to the wrong question, which can always be made precise”.
The infinite improbability principle also applies to deeper complexities. Improbable though it may seem, evaluation is itself indeterminate, in the sense that there may be more than one “right” (warrantable) answer to an evaluative question. Moreover, even the most purportedly “objective” indicators can mask irreducible differences in what matters to people. Different people value the same things differently. What is in my interests might directly contravene yours; what is respectful in one community may look quite different to another. There is rarely absolute consensus on what matters, and it can shift across cultures, times and places.
Programs also face multiple possible futures, each with different value propositions. Today's resilience investment might later prove invaluable during an extreme weather event, moderately valuable peace-of-mind during more stable times, or irrelevant if better solutions emerge. Scenario planning and robust decision-making approaches help navigate this temporal uncertainty by evaluating how programs perform across different plausible futures rather than optimising for a single predicted outcome.
The improbability principle suggests that rather than seeking the impossible, such as perfect certainty about hard-to-measure and contested values across unknowable futures, effective evaluations embrace structured approaches to uncertainty. Examples include stakeholder-centric processes that make different values explicit, scenario analysis that explores alternative futures, adaptive frameworks that can evolve as both values and circumstances change, and getting AI tools to run the same evaluation 1,000 times to see the range and distribution of possible conclusions. Like the crew on the Heart of Gold (the stolen starship powered by the Infinite Improbability Drive), evaluators can navigate infinite improbability with tools and approaches designed for uncertainty rather than precision.
8. The Total Perspective Vortex: context matters
The Total Perspective Vortex was a device designed to show any individual their true place in the infinite cosmos, with utterly devastating psychological consequences for those who experienced it. The machine revealed an overwhelming view of the entire universe and then pinpointed the observer’s location - a tiny dot labelled “you are here” - making palpable the insignificance of their existence compared to the scale of creation.
For VfM evaluation, this highlights the importance of a whole-system perspective and recognising that things can look very different depending where we choose to draw boundaries. A robust VfM assessment demands more than checking a project’s costs and benefits in isolation; it calls for examining how well an intervention fits within a bigger picture (maybe not the whole universe, but at least some relevant adjacent parts) including overlapping initiatives, strategies, and policies. This can include, for example, assessing how coherently a program builds synergies with related efforts, complements broader goals, and avoids duplication or unintended conflict.
This systems perspective acknowledges that value constantly emerges from, and depends on, its context: the same intervention might deliver vastly different results depending on surrounding investments, community and individual characteristics, and the unpredictable ‘ocean currents’ that carry them all. Therefore, evaluators ask not only “what works?” but “for whom and in what circumstances?”, “how does this program interact with the wider landscape of change?” and “what does it even mean for something to ‘work’ in the context of a system transformation?”
9. The Restaurant at the End of the Universe: the temporal dimension in evaluation
Adams' second book features a restaurant that exists not at a place, but at the end of time, allowing diners to watch the universe's final moment while enjoying a meal. What a mind! How did he come up with this stuff?
Thinking in temporal perspectives is important in evaluation and VfM: we have to consider different time horizons and their implications for value assessment.
Value is rarely static; it emerges and shifts over time, sometimes only becoming clear years after decisions are made. As programs progress, what is considered "valuable" or "successful" naturally shifts, demanding that evaluators update criteria and adjust the weight given to different evidence according to a project's stage. Meaningful evaluation considers not only results but also the quality of the change process, anticipating future impacts and recognising that today's actions can create ripples that extend well beyond the present. Contemporary frameworks increasingly embrace these dynamics, reinforcing that value creation involves both careful stewardship of resources and forward-looking, adaptive reasoning.
The temporal dimension suggests that VfM frameworks ought to treat criteria like the 5Es (economy, efficiency, effectiveness, cost-effectiveness, and equity) as an evolving balancing act - an exercise in optimisation rather than maximisation. Early in the life of a program, we might focus predominantly on economy and efficiency, but the most valuable long-term impacts probably won’t come from “spending less” on inputs. Working efficiently is a dynamic endeavour, adapting strategies, acting on emerging opportunities, exiting ineffective investments to cut losses. If you want to maximise equity gains, you might have to compromise on efficiency. And so on.
There is always a tension between the 5Es in value for money, and there are some concerns when we talk about “has it reached value for money?”. We want everything; we want efficiency, we want effectiveness, we want equity, but you cannot have them all at the same level. We need to be very clear; what is the value that is most needed? Is it equity? Is reaching the poorest of the poorest the main goal? You want to be able to do very well in the other Es, but you need to have one or two that are above the others.
(Anisa Berdellima, Director of Evidence and Impact at MSI Reproductive Choices, testifying to the International Development Committee on FCDO’s approach to value for money, 17 June 2025).
10. The Paranoid Android's wisdom: accepting imperfection
Marvin the Paranoid Android (voiced in the movie by the inimitable Alan Rickman), perpetually depressed by the gap between possibility and reality, offers a final lesson for evaluators. Perfect evaluation is impossible, but that’s no excuse for inaction. Humility about limitations must be paired with commitment to usefulness. The goal isn’t perfection, but informing good decisions.
Evaluation can serve democracy by informing good decisions that create value for people and by supporting transparency, accountability, learning, and adaptation. In our world of infinite improbability, rigorous and accessible VfM evaluation is an act of hope - betting that despite our inconsequential existence in the context of the entire universe, better information can lead to better decisions and more just allocation of resources in the spaces that matter to us. Better VfM assessment includes transparent reasoning, inclusive sense-making, rigorous measurement, and ensuring essential information is accessible, not hidden behind a “Beware of the Leopard” sign.
Don't Panic. But do demand better VfM evaluation.
Thanks for being part of this community! This is the last instalment for 2025. See you next year.
Special offer: 50% discount forever
Evaluation and Value for Investment is a weekly blog for people who care about using evidence and explicit values to make better decisions. It focuses on practical, thought-provoking pieces that aim to be genuinely useful, interesting, and occasionally fun for evaluators, commissioners, and decision-makers.
It’s free to sign up, and most new articles stay freely available for the first two months after posting. Paid subscribers get full access to the complete archive, so they can revisit past pieces, frameworks, and tools whenever they need them.
For a limited time, you can purchase a paid subscription at 50% off the usual price. You can use this to upgrade your own access, or gift a subscription to that special evaluator in your life.
👉 Follow this link to see pricing and complete your subscription or gift purchase. If you already read the free posts and find them valuable, this is a low-cost way to support the work while unlocking the full back catalogue.
If you’re an evaluator working in a low- or middle-income country, a student, or you are between roles and the subscription cost is a stretch right now, please feel free to contact me directly. It will be my pleasure to open up complimentary access so that cost is not a barrier to using these resources.
Douglas Adams’ The Hitchhiker’s Guide to the Galaxy is a seminal work of comic science fiction that began as a BBC radio series in 1978 and was published as a novel the following year. The story follows Arthur Dent, an Englishman whose home is scheduled to be demolished to make way for a new road project. Putting Arthur’s predicament in perspective, the very same day Earth is demolished by a Vogon construction fleet to make way for an interstellar bypass. Escaping at the very last minute with his eccentric friend Ford Prefect, Arthur is swept up by the spaceship Heart of Gold, crewed by the two-headed Zaphod Beeblebrox, brainy Trillian, and chronically depressed robot Marvin. Blending philosophical musings and satirical wit, the book lampoons bureaucracy, technology, and the search for meaning with surreal set pieces and existential humour, which may be why it continues to resonate so strongly with evaluators and other bureaucracy-watchers. Its cult status was further cemented with adaptations as a BBC TV series (1981), computer game, and a quarter-century later, a Hollywood film featuring Martin Freeman, Sam Rockwell, Alan Rickman, Zooey Deschanel, and brilliant cameos from Bill Nighy, Helen Mirren, Stephen Fry, and John Malkovich.




