This post is timed to coincide with a panel presentation I was in today, with Patricia Rogers and Scott Bayley at the Australian Evaluation Society Conference in Canberra.
Our topic: Evidentiary criteria for causal inference. In recent decades, persistent debate within the evaluation field about the "gold standard" for impact evaluation methods has come at the cost of progress, with many practitioners agreeing that context-specific choices based on intervention type, evaluation questions, and available resources are essential. Despite this, the debate remains stuck. Scott suggested it may be more productive to refocus on the evidentiary criteria required for establishing credible and defensible causal claims - and hence this panel was born.
Here I will focus on just one example from the presentation, going into a bit more depth to illustrate one set of evidentiary criteria. We’ll write more about the broader topic too - stay tuned.
Bradford Hill’s criteria
In 1965, Sir Austin Bradford Hill, an epidemiologist, presented an essay to the Royal Society of Medicine. In it, he presented what he called, not criteria, but “nine different viewpoints from all of which we should study association before we cry causation”.
He argued that these viewpoints were not hard-and-fast rules of evidence. They could not provide indisputable evidence for or against a cause-and-effect hypothesis. But they could help us to weigh the available evidence for or against various possible interpretations of cause and effect.
Now, these criteria were not designed for policy and program evaluation. They were specifically developed for appraising potential causal relationships in observational data in epidemiology - for example, to determine whether or not a particular chemical might be considered carcinogenic on the basis of exposure-disease associations in a free living population, with no control group.
However, what the field of evaluation may be able to take away is the idea that a similar, but distinct menu of criteria might apply when we have to consider the balance of messy, real-world evidence to make judgements about causality or contribution.
Reasoning, not methods
Our panel presentation suggested that, instead of fixating on method superiority, evaluators could advance the field by focusing on criteria for interrogating the validity of causal claims.
The Bradford Hill criteria model what a similar approach looks like in another field, weighing multiple lines of evidence. This approach encourages epidemiologists to transparently reason and argue why causal claims may or may not be warranted, integrating both quantitative and qualitative data, and reflecting on whether the totality of evidence supports a defensible conclusion.
Example: can antihistamines cause dementia?
Observational studies have linked first-generation antihistamines, like diphenhydramine (e.g., Benadryl), to risk of dementia in older people, especially when used often and for a long time. The mechanism is thought to be impairment of acetylcholine signalling, which is vital for memory. Newer allergy medicines like loratadine and fexofenadine are much less risky.
These findings come from long-term observation, not experiments. Researchers have monitored groups over time to see if people taking anticholinergic drugs developed dementia more frequently than those who do not. Data typically comes from pharmacy dispensing records, health surveys, and clinical follow-up. Analyses rely on statistical models to control for confounders (such as age, comorbidities, and other medication use). To date, there have been no RCTs directly testing whether antihistamines cause dementia, presumably due to ethical concerns given the potential risk.
Applying the Bradford Hill criteria
Applying Hill’s criteria to the observational evidence linking anticholinergic antihistamines to dementia risk suggests moderate support for causality, with some limitations mostly due to the nature of observational studies.
Strength: association hints at causality
Association is not causation. However, a strong association (e.g., correlation, between-group difference, or within-group change), in combination with other considerations, can lend support to a cause-effect interpretation. A weaker association would contribute weaker support. No association would contribute toward a judgement of non-causality.
Studies consistently show a statistically significant association between cumulative anticholinergic use and dementia risk, though hazard ratios are usually modest (1.5–2.1, meaning that people exposed to anticholinergics were 50% to 110% more likely to develop dementia than those who weren’t exposed). This implies a small but valid increased risk.
Consistency: findings repeat across studies/populations
Consistency means that a cause-effect association is observed repeatedly across different studies, populations, and circumstances (i.e., reproducability). When multiple sources or types of evidence show the same association, this strengthens the inference of causality by demonstrating that the finding isn’t unique to a single context or study. Conversely, inconsistent results would tend toward a judgement of non-causality.
Multiple studies across diverse populations, settings, and drug classes suggest that chronic use of anticholinergics increases dementia risk.
Specificity: association tied to a specific cause-effect combination
A causal inference is strengthened if the effect is observed only in association with the suspected cause, and not seen in the absence of the suspected cause. If the effect is seen more often in association with the suspected cause than without it, this might lend weaker support to a causal inference. Absence of specificity doesn’t rule out a possible causal relationship.
The association is observed most strongly for drugs with high anticholinergic activity, but not exclusively for antihistamines (other drug classes show similar patterns). Specificity is present but not absolute.
Temporality: cause comes before effect
If the expected temporal relationship is observed, this lends support to our cause-effect interpretation. Absence of temporality may suggest lack of a causal relationship - but be careful: in complex systems like social programs, feedback loops might mean causality is bidirectional and multi-factorial. Moreover, some effects happen in anticipation of causes, e.g., anxiety before a social interaction.
Exposure to anticholinergic medication precedes dementia diagnosis in cohort studies, supporting the criterion of temporal sequence.
Dose-response gradient: effect rises with greater exposure
A strong and consistent dose-response gradient would contribute strong support to a cause-effect interpretation. A reasonably consistent gradient might lend weaker support. Absence of a dose-response relationship doesn’t rule out causality. Bear in mind that these gradients may not be linear. For example, some interventions may be beneficial in small doses but harmful in higher doses and/or depend on the initial state (e.g., dietary iron).
Some studies show a dose-response relationship - higher cumulative exposure to anticholinergics is associated with higher risk of cognitive decline.
Plausibility: reasonable mechanism established or theorised
A scientifically credible mechanism by which the proposed cause could produce the effect supports a causal interpretation. Absence of a plausible mechanism doesn’t rule out a possible causal relationship. As Bradford Hill noted, “What is biologically plausible depends on the biological knowledge of the day”.
The biological mechanism, blocking acetylcholine essential for memory/learning, is well-understood and plausible. Acetylcholine helps nerve cells in the brain to transmit information, supports the formation of new memories, and boosts learning by enhancing communication between key regions such as the hippocampus and cortex. Drugs that interfere with acetylcholine can make it harder for the brain to store and retrieve information, leading to memory problems and, possibly, a higher risk of dementia over time.
Coherence: no major conflict with existing knowledge
Coherence considers whether the proposed causal relationship fits with the broader body of existing knowledge about the subject. While coherence supports a causal claim, lack of coherence with current knowledge does not rule out causality if the empirical evidence is persuasive. Scientific understanding can evolve in response to new findings.
Findings are coherent with established knowledge about acetylcholine’s role in cognition and dementia pathology.
Experiment: testing cause through any intervention
Bradford Hill’s criterion of “experiment” means more than just RCT - e.g., it also includes natural experiments, real-world learning through intervention and practical trial, studies showing changes in disease risk after removing an exposure, animal or lab research that manipulates variables to reveal mechanisms or outcomes, quasi-experimental and community intervention designs where exposures aren’t randomly assigned but are otherwise manipulated and outcomes are measured.
In the case of anticholinergics and dementia, evidence from any form of experimental human trials is lacking; most support comes from observational and animal studies.
Analogy: causal link seen in similar situations
Analogy contributes weak support for a causal inference. For example, if we know how a process works in one context, we can use this knowledge as an explanatory model to propose how it might work in a new context. Absence of analogy doesn’t rule out a possible causal relationship.
Other medications with similar anticholinergic actions (such as some antidepressants, bladder medications, and antipsychotics) have also been consistently associated in multiple studies to cognitive decline, supporting this criterion.
Conclusion
Bradford Hill criteria of strength, consistency, specificity, temporality, gradient, plausibility, coherence, and analogy, all lend support to a causal interpretation. No experimental evidence was found. Confounding remains possible, as many users are older and may have underlying risk factors. The effect sizes are modest, and causality is not “proven” but on the basis of the observational evidence, a precautionary approach is warranted.
Contemporary critiques of Bradford Hill criteria
As Sir Austin himself cautioned, the criteria are not a rigid checklist: not all need be satisfied to infer causality, nor does satisfying them guarantee a causal relationship. They don’t replace expert judgement.
The Bradford Hill criteria aren’t for every situation. They’re at their most decisive in situations with a dominant cause of a single effect/risk (e.g., smoking and lung cancer; asbestos and mesothelioma). They aid clarity in multifactorial relationships such as the link between first-generation antihistamines and dementia discussed above (dementia is also associated with other medications as well as age, genetics, education, head injury, cardiovascular risk factors, and lifestyle). The criteria struggle with greater complexity - e.g., multiple interacting causes and systems, feedback loops, context-dependent effects, active interventions (like policies and programs) that can be varied, or where different criteria give conflicting signals.
Over-reliance on the criteria would risk misinterpretation or over-confidence if important nuances and contradictory findings are ignored. For all of these reasons, the Bradford Hill criteria are best seen as inductive guides to inform reasoning, not as definitive standards. Scientific context and common sense remain essential.
Despite these cautions, the criteria are in widespread use. They are commonly applied in epidemiology, public health, and biomedical research as a flexible framework for assessing causality, even as their limitations and appropriate application continue to be debated. Researchers frequently use the criteria to structure discussions about causal associations and guide evidence assessment, especially in policy, guideline development, and risk communication.
How might our evidentiary criteria differ for policy and program evaluation?
This is not a proposal to adopt Bradford Hill wholesale. A stronger focus on evidentiary frameworks - not Bradford Hill’s criteria per se but something (or things) analogous, tailor made for evaluation - could facilitate constructive dialogue on what makes impact evaluation findings credible, shift attention to the logics behind causal inference, and acknowledge the messy realities of evaluation contexts. This focus may offer more productive debates about rigorous, context-appropriate evidence for determining intervention effects.
We plan to share more from this panel in future writing, but for now, here’s a brief outline of my co-panellists’ presentations.
Scott Bayley argued that impact evaluation is fundamentally about examining causal relationships. He noted that causality is a philosophical concept with many schools of thought. Scott outlined critical multiplism (his preferred approach) and identified four key evidentiary standards for concluding that an intervention (X) caused an outcome (Y):
Demonstrating an association between participating in X and achieving Y
Establishing temporal order (X occurs before Y)
Ruling out rival/alternative explanations for Y
Identifying a plausible mechanism linking X and Y.
Patricia Rogers built on this by grouping different evidentiary standards into those testing congruence with a causal relationship and those testing alternative explanations. In particular, she:
Emphasised that good program theory (e.g., theory of change) is necessary to move beyond simple association, and introduced the idea of causal packages (groups of factors or conditions that together produce an outcome).
Argued that causal questions must be clearly framed - for example, questions about whether an intervention contributed to specific long-term changes differ fundamentally from those asking whether the intervention caused those changes.
Drew some concepts from process tracing, distinguishing evidence that is consistent with a causal claim (‘straw in the wind’), evidence which strongly rules it out (‘hoop’ - e.g., timing), and evidence that strongly supports it (‘smoking gun’).
Referenced Marina Apgar and Tom Aston’s rubrics for assessing evidence and the types of statements being made (e.g., contrasting retrospective recollection by an unreliable witness with contemporaneous official recording).
Stay tuned - more to follow!
There was a rich discussion from the audience too. We intend to write more. Subscribe and stay tuned!
Bonus content
Acetaminophen (aka paracetamol) and autism: after recent Trump administration announcements linking acetaminophen use in pregnancy to autism risk, I ran it through the Bradford Hill Criteria here. The verdict: no credible link. (added 23/09/25).
Here’s something I prepared (much) earlier on Bradford Hill criteria, for the 2015 ANZEA Conference. Like the Ship of Theseus, I am a different person now (nearly all of the cells I was made of ten years ago will have been replaced multiple times; some of my brain cells may be the same but hopefully they continued learning and evolving in the intervening decade). I still mostly agree with my former self, but in retrospect I was overly optimistic in transferring these criteria directly to evaluation. Our field needs its own evidentiary criteria for causal inference, though they may retain a resemblance to some of Hill’s.
In this study, we used Bradford Hill criteria to assess statistical trends in longitudinal observational data collected by an alcohol and drug rehabilitation program. The program administered multiple validated psychometric tools at several time points to track clients’ recovery. Our analysis found that the observed trends met criteria for strength, consistency, temporality, and coherence. Additionally, a study of the broader literature on similar programs found that the overall body of evidence aligned with Hill’s criteria.
Previous posts exploring other aspects of causal inference include the following:






