When it comes to evaluating public policies and social programs, I think it’s important to consider a wide range of methods for understanding causal relationships between actions and impacts, and for exploring factors that influence outcomes.
But, when I’m prescribed a pharmaceutical drug, I expect to see high-quality randomised controlled trial (RCT) evidence, preferably from several studies, and ideally a meta-analysis or systematic review of multiple trials (yes, I’m one of those patients who does my own literature review before accepting or declining my doctor’s recommendation).
Is this a double-standard? Can I reconcile this apparent inconsistency? And if so, how? In this post I’ll explore what sits behind my thinking, as curious as you may be whether I can make it stack up…
Why I seek RCT evidence for medical interventions
I don’t brook “gold standards” in evaluation - but that label is often applied to RCTs as an approach to causal inference, especially for pharmaceuticals and many other medical/surgical interventions. I’m mindful that evidence hierarchies are themselves value-laden, reflecting the priorities, assumptions, and debates of scientific communities. RCTs aren’t perfect, but the following features make them desirable in medical research:
Medical interventions are often amenable to RCT research: Not every intervention is a good candidate for an RCT. But a lot of health care interventions are. For example, drugs can be standardised in formulation, dosing, and administration, so they’re inherently suited to the consistent conditions required in RCTs. In a clinical trial setting, we can have a reasonable expectation of precise measurements of effects and confident attribution of outcomes to the intervention. The structure of clinical trials allows for pre-specified outcomes and analysis plans, well-powered studies, robust statistical analysis, reproducibility and transparency.
Minimising (some kinds of) bias and confounding: RCTs use random assignment and ideally blinding to balance both known and unknown factors between groups.1 This minimises the chances that results are due to something other than the intervention itself. For medical interventions, where small differences in outcomes can have major implications for safety and effectiveness, I value the edge this provides.
High stakes and risks: Drugs and surgery can have powerful and sometimes harmful effects on individual patients and populations. The risks are high, so the causal inference bar is set accordingly. I tend to err on the side of not taking a medication unless the benefits and risks are well understood. You might apply a different standard, of course.
Clinical guidelines: RCTs inform clinical practice guidelines, treatment protocols, and health care decisions at scale. These guidelines are influential. Doctors may feel significant professional and legal pressure to follow guidelines, as deviations can raise questions of liability, especially if adverse outcomes occur. RCT evidence helps doctors trust the validity of the guidelines.
Regulatory and ethical standards: RCTs are generally (with a few exceptions) required by regulatory agencies such as the Food and Drug Administration (FDA) in the U.S., and the Therapeutic Goods Administration (TGA) in Australia, as a core part of the process to approve new drugs, ensuring that medicines are both safe and effective before they reach the public. This standard was brought in to help protect patients from ineffective or harmful treatments and to ensure that marketing claims are backed by solid evidence.2
Public trust: The requirement for RCTs, along with trial registration and reporting standards, promotes transparency in drug development. This helps maintain public trust in medicine and the healthcare system by ensuring that claims about safety and efficacy are based on open, verifiable evidence - though, as with all institutions, public trust also depends on the integrity with which evidence is generated and communicated.
These aren’t reasons to favour RCTs at the exclusion of other methods. But it’s enough to make me think I want them in the mix. For the collective weight of the reasons above, I’m more likely to trust medical interventions backed by reliable and replicated RCT evidence that suggests their clinical benefits outweigh their risks.
Why I embrace diverse causal inference strategies in policy and program evaluation
Value and causation are different things
At its core, evaluation is about weighing information about program/policy performance (e.g., its quality, success, and value) - including what matters to people - and making (or working with stakeholders to make) considered, transparent value judgements - such as judgements about how well things are going, what actions to take next, and so on.
Judging value is conceptually and methodologically distinct from inferring causation. You can’t judge value with an RCT.3 While causal inference (through RCTs or other methods) helps us determine if an intervention made a difference, it does not tell us whether that difference is valuable, clinically significant, appropriate, or sufficient. That step requires an evaluative judgement. Causal inference is, however, an important supporting component of many evaluations because, before we judge the value of an impact, we must ascertain if there is an impact - and if so, how large, and to what degree of precision and certainty.
RCTs alone aren’t enough in policy and program evaluation
I am not against the use of RCTs to support causal inferences in evaluation. Looking back at my reasons for wanting to see RCT evidence in medical research, some of the considerations apply here too - e.g., high stakes, potential for policies and programs to have significant positive or negative impacts at scale, need for public trust. However, evaluating social programs and public policies demands a broader toolkit than RCTs alone. Here are some reasons why:
When we’re all in, there’s no control group: Some interventions affect whole communities, social structures or systems. In these cases, randomisation is often impractical. When evaluating global initiatives or large-scale policy reforms, we don’t have a measurable counterfactual (e.g., a parallel Earth or a parallel New Zealand, without the intervention).
Multiple inter-related outcomes and synergistic effects: Social interventions rarely operate in isolation or target a single, easily measurable outcome. Instead, they tackle a web of interconnected issues across diverse populations. The relationships between causes and effects are often “many-to-many” - a single intervention can involve multiple causes influencing multiple outcomes. The combined effects of multiple interventions may be greater (or different) than the sum of their individual impacts. If we attempt to isolate the specific contribution of any one part, our estimate may be off.
Emergence: Social programs operate within complex systems where outcomes often can’t be fully predicted in advance. Interventions may lead to emergent effects - new patterns, behaviours, or outcomes that arise from the dynamic interactions of people and contexts. When the impacts of a program can evolve in unexpected ways over time, predetermined outcome measures could miss something important.
Stakeholder diversity: Evaluations need to account for the perspectives and needs of multiple stakeholder groups - such as policymakers, service providers, recipients, and subgroups - each of whom may value things differently and require different kinds of evidence. Multiple, and inclusive methods that help make sense of what’s happening, for whom, and why, can help to paint a richer picture.
Implementation variation: There are often significant differences between how a program is designed and how it is actually delivered in practice, due to the realities of resource constraints, adaptations to respond to local conditions or individual circumstances, and unanticipated challenges. Fidelity is less like adhering to a rigid recipe and more like being mindful of a set of guiding principles. Kinda the opposite of a rigid treatment protocol.
Adaptive practice: In community settings, programs may operate alongside (or in conjunction with) related initiatives and within shifting policy landscapes. Rather than holding things constant to measure them consistently over a duration, effective programs need to be able to adapt and evolve in response to ongoing learning and new priorities. This flexibility allows for more responsive, contextually relevant interventions that can better meet the needs of communities - and it requires commensurately responsive, contextually relevant causal inference approaches.
Mechanisms and variation: While RCTs quantify effect sizes (including variation in effects across subgroups), they must be complemented by additional methods to unpack underlying causal mechanisms or explain why interventions work well for some people but less well (or not at all) for others. Understanding nuances such as reasons behind differential effectiveness, the role of contextual factors, and the sometimes diverse pathways through which impact occurs requires methods that capture personal experiences, contextual influences, and other information.
Experimental adaptations are possible, within limits
All sorts of innovative RCT designs have been developed over the years to address some of the real-world challenges of experimental research in complex settings. These designs can make RCTs more practical and flexible - for example, by phasing in different interventions over time, building in protocols for iterative adaptation, or finding creative ways to accommodate ethical or logistical constraints. However, each of these adaptations involves trade-offs; we have to accept modifications to service implementation, such as when, how, or to whom interventions are delivered, to make the study feasible and statistically robust.
These modified designs are ingenious and genuinely useful. I am not opposed to their use. What I resist is the mindset of privileging experimental designs as if they’re a universal “gold standard” at the expense of other approaches tailored to complex settings. Alternative approaches may simply be better suited to some problems; there’s little point designing cleverer hammers when we could just reach for a screwdriver. Over-reliance on RCTs would narrow our understanding and undervalue the insights we can get from other kinds of evidence.
Towards methodological pluralism
The upshot of all this complexity is that theorists and practitioners have developed multiple strategies for identifying, quantifying and qualifying causality and contribution, and seeking to understand the nuances of policies and programs in real-world contexts. Examples in the footnote.4 Effective evaluation in complex settings demands flexibility, drawing on complementary strengths of diverse causal inference strategies.
There are shortcomings to RCTs in clinical research, and broader approaches are needed there too
I’ll continue to search for RCT evidence before I decide whether to pop a new pill. However, I also know RCTs in real-world clinical research are far from perfect. Among their shortcomings:
The transferability problem: RCTs address threats to internal validity, ensuring results are reliable for ‘this’ study population in the conditions of ‘this’ trial. However, external validity (the extent to which findings are applicable to broader, real-world populations and settings) is a separate challenge. For example, a treatment found to be effective in a carefully selected trial population may not work the same way for older patients with multiple conditions in regular practice. Results of RCTs don’t predict what outcomes will be under different circumstances - e.g., differences in healthcare environments, patient eligibility, intervention protocols, and real-world adherence. These factors, which are ever-present, can significantly affect whether and how trial results translate into everyday practice.
They mask complexity and may not be all that helpful for individual patients: I’ve written previously about the need to complement RCTs with research that better reflects the messy realities of everyday healthcare. A reductive focus on average effects in idealised populations creates evidence that, although valuable to inform population-level resource allocation decisions, often fails to help individual patients navigating multiple, overlapping challenges.
They exclude populations: Despite some recent improvements in response to advocacy, RCTs often systematically under-represent women, minority populations, and people with complex health conditions. This limits our understanding of how treatments work across diverse groups. Potentially, it masks differences in efficacy, safety, and side effects that may be specific to certain populations. Consequently, medications could be less effective or even harmful to those not adequately represented in trials, and clinical guidelines may not be appropriately tailored to the needs of the broader patient population. The risks of excluding diverse groups from drug trials include perpetuating health disparities, eroding trust in research and healthcare, and holding back equitable access to new treatments.
They’re not always done: A 2025 investigation reported that 73% of U.S. Food and Drug Administration (FDA)-approved drugs from 2013–2022 lacked robust evidence of efficacy, with 39 drugs meeting none of the agency’s core scientific standards. Cancer drugs fared particularly poorly; 81% were approved based on preliminary evidence (like tumour shrinkage) rather than survival benefits, and only 2.4% met all FDA criteria.
They’re not always done well: While many RCTs adhere to high standards, serious lapses in design and reporting do occur. Poor randomisation can lead to biased group allocation. Underpowered samples may fail to detect clinical differences. Surrogate endpoints - relying on markers rather than actual health outcomes - can be misleading. There’s a well-known reproducibility problem with RCTs, adding fuel to the claim that “most published research findings are false”, as John Ioannidis famously wrote. It can sometimes take a lot of expert knowledge and skill to pick apart RCTs and tell what is trustworthy and what is not - one of the reasons I follow Peter Attia.
There are incentives to do poor RCTs: Financial and competitive pressures can sometimes compromise scientific rigour. Common strategies to design and conduct RCTs that maximise the likelihood of positive results include comparing a new drug to an inferior one rather than a placebo or best available therapy, selecting relatively healthy participants to inflate efficacy and minimise side effects, using inadequate randomisation or blinding, relying on surrogate endpoints instead of meaningful clinical outcomes, or shortening follow-up periods to avoid detecting late-emerging risks. Composite outcomes (combining disparate metrics, such as hospitalisations and lab values) can overstate benefits. Companies may also selectively publish favourable studies while withholding negative results, or exert control over data analysis and reporting to present their products in a positive light. These practices systematically bias trial findings and undermine the reliability of the evidence base used for regulatory approval and clinical decision-making.
The answer to some of the above may be to demand more and better RCTs. However, part of the answer is also to conduct more than just RCTs, to bring in broader insights. I wrote about broader approaches in clinical research in a previous article.
Conclusion
I don’t think it’s entirely inconsistent of me to seek RCTs for medicines while embracing a broader range of methods for program evaluation. My standards for causal inference reflect the contexts and practicalities of different domains. In medicine, where interventions are amenable to experimental study designs, I place a premium on RCT evidence, both for population and individual-level decision-making. In social program evaluation, where complexity features prominently and contexts vary widely, I value a broader, more flexible approach. This isn’t a double standard - it’s a recognition that different questions and settings demand different tools.
I don’t mean to set up a false dichotomy though. RCTs and other methods should be seen not as competitors, but collaborators in understanding complex realities. Experimental and non-experimental causal inference strategies all have their place, in health and other social contexts. We can have both. We should have both. I recognise the value of experimental study designs in program and policy evaluation, where appropriate. And I am all for the use of mixed methods approaches to causal inference in medical research to complement the insights we get from RCTs.
Uncritical faith in RCT evidence (or anything else) as a standalone “gold standard” is misguided, even in medicine. Many drugs are approved on the basis of weak or incomplete RCT evidence, and RCTs often fail to represent the diversity and complexity of real-world patients. In my view, progress lies in integrating judicious use of RCTs with real-world data, patient perspectives, adaptive frameworks, and the rich toolkit of research and evaluation methods, recognising that all evidence has limitations and requires transparent, critical scrutiny.
In Australia currently, Andrew Leigh (Assistant Minister for Productivity, Competition, Charities and Treasury in the Federal Government) stands out as a strong advocate for greater use of RCTs in policy analysis to provide robust evidence for policy decisions. Minister Leigh acknowledges that RCTs are not always feasible or ethical and that some policy questions require alternative approaches.
My views differ from the Minister’s in terms of emphasis. He may well be right that there’s potential for more RCTs, but I think we need a broad menu of options and mixed methods approaches to causal inference - and not just as a fallback when RCTs aren’t possible.
I place greater weight on methodological pluralism: selecting and combining approaches contextually to address the complexities of real-world policy, spanning technocratic and deliberative, interpretive approaches. Sound causal reasoning (e.g., “did it work?”) is essential, no matter the methods. Sound evaluative reasoning is too (e.g., “was it worthwhile and fair?”) - and RCTs don’t do that. Both medicine and program evaluation will benefit from a more systemic, inclusive, and context-sensitive approach to evidence. Complexity is an inherent reality of many public policies and social programs, and it’s on us to design methods that can keep up.
Acknowledgements
I’d like to thank Kate McKegg for peer review. This post represents my opinions alone, and not the organisations or individuals I work with. Moreover, these are my opinions today. They’re updatable. Go ahead and change my mind. We are all teachers and learners here!
👆 Recognise this? Be first to correctly identify it in the comments section of this post and win a free 12-month subscription, giving you access to my full archive.
RCTs can be designed to address several kinds of bias that threaten the internal validity (accuracy of causal inferences within the study) of clinical research findings. These include selection bias (skewed sample selection), performance bias (unequal treatment between groups), detection bias (differential outcome assessment), and attrition bias (skewed results from unequal dropout). These biases are addressed through randomisation (assigning participants to treatment and control groups randomly) and blinding (concealing treatment and control group assignments from participants and researchers). However, RCTs don’t eliminate all biases. Some residual biases (e.g., from imperfect implementation, or unmeasured confounding) can still occur. Threats to external validity (applicability beyond the study to real-world populations) include contextual variability in factors like healthcare settings, eligibility criteria, intervention protocols, patient adherence and other things that aren’t replicated outside trials.
The regulatory and ethical standards governing RCTs for new medicines were significantly influenced by the Thalidomide tragedy of the late 1950s and early 1960s. The widespread congenital anomalies caused by Thalidomide, which had not been adequately tested before hitting the market, exposed critical gaps in drug evaluation processes. In response, the U.S. passed the 1962 Kefauver-Harris Amendments, mandating that drug manufacturers provide robust evidence from well-controlled clinical trials demonstrating both safety and efficacy before approval, and requiring informed consent from trial participants. Similar reforms occurred internationally, with new systems for adverse event reporting and stricter requirements for preclinical and clinical testing, especially regarding the risk of congenital conditions. These changes established the foundation for modern ethical and regulatory frameworks in clinical research, emphasising participant protection, transparency, and rigorous scientific methods.
My point “you can’t judge value with an RCT” is true but has exceptions. In health economics, studies using utility measures like quality-adjusted life years (QALYs) blend causal and value judgements, blurring the distinction. Conceptually, the distinction is still there but the value and the causation are all woven into the measure. Even so, I still maintain that RCTs can’t on their own, capture the full spectrum of contextual and stakeholder value judgements required in evaluation.
Here’s a guide from Patricia Rogers covering key concepts in causal inference. And here are some examples of various causal inference strategies. This list isn’t systematic or comprehensive. These examples aren’t mutually exclusive - some cover overlapping concepts, and different strategies can be used in combination to build up a credible and nuanced picture of posited causal relationships and learning about what works, for whom, when, etc. To illustrate the breadth of causal inference options available, examples include: theory of change; realist evaluation; causal pathways analysis; qualitative impact protocol; contribution analysis; process tracing; contribution tracing with Bayesian updating; case-based and comparative methods; checklists (e.g., Jane Davidson’s causal inference strategies in her book, Evaluation Methodology Basics: Bradford Hill’s criteria sometimes used in epidemiology; rubrics (e.g., see Aston & Apgar); quasi-experimental designs such as matching and propensity score analysis, difference-in-differences and time series analysis, estimated counterfactual, equivalent groups designs and regression discontinuity; and RCTs. If I forgot your favourite, let me know. Fun game: draw any two evaluators at random and they will find something on this list to disagree about.



The most lettered critique of RCTs I have read. Thanks
Battersea Power Station?