Bradford Hill viewpoints: what they can and cannot settle
The Bradford Hill criteria explained: what Hill's nine viewpoints were for, why he warned against checklist use, and how they are applied today.
By Manouchehr Hessabi, MD, MPH
The Bradford Hill criteria are nine viewpoints that Austin Bradford Hill proposed in 1965 for judging whether an observed association between an exposure and a health outcome is best explained by causation. They are not a test a finding passes or fails. Hill offered them as aids to judgment, and he said so directly in the same address that introduced them.
That distinction has not survived well. The viewpoints are among the most cited ideas in epidemiology and among the most misapplied, routinely presented as a nine-item checklist where meeting more items means a stronger causal claim. This article returns to what Hill actually wrote, walks the viewpoints in plain language, and looks at how the questions changed once molecular data entered the field.
What are the Bradford Hill criteria?
They come from a presidential address Hill delivered in 1965 to the Section of Occupational Medicine of the Royal Society of Medicine, titled "The environment and disease: association or causation?" The address was published in the Proceedings of the Royal Society of Medicine and republished by the Journal of the Royal Society of Medicine in 2015, fifty years on.
The setting matters. Hill was speaking to occupational physicians during the period in which the relationship between smoking and lung cancer was being argued out, a debate in which he had been a central figure. His question was practical: given an association that is unlikely to be due to chance, what should make an observer move from "these things occur together" to "this one produces the other"?
His answer was nine considerations, usually named as follows:
- Strength. How large is the association?
- Consistency. Has it been observed repeatedly, by different investigators, in different populations and circumstances?
- Specificity. Is the exposure tied to a particular outcome, in a particular group?
- Temporality. Did the exposure precede the outcome?
- Biological gradient. Does more exposure correspond to more of the outcome?
- Plausibility. Is a causal relationship biologically credible given what is known?
- Coherence. Does the interpretation conflict with what is otherwise known about the disease?
- Experiment. Does intervening on the exposure change the outcome?
- Analogy. Have similar exposures produced similar effects elsewhere?
Why did Hill warn against treating them as a checklist?
Because he said, in the address itself, that they could not function as one. Hill wrote that none of his nine viewpoints could "bring indisputable evidence for or against the cause-and-effect hypothesis" and that "none can be required as a sine qua non." What they could do, he argued, was help an observer answer the underlying question: is there any other way of explaining the set of facts, any other answer equally likely or more likely than cause and effect?
That framing is closer to structured skepticism than to scoring. The viewpoints are prompts for asking what else could account for the pattern, not boxes that accumulate toward a verdict.
Two of the nine sit differently from the rest, and it is worth separating them out.
Temporality is the one genuine requirement. A cause must precede its effect. This is not a matter of weight or degree, and it is different in kind from the other eight: an association in which the outcome preceded the exposure is not a weaker causal claim, it is not a causal claim at all. Epidemiologists have long treated this as the non-negotiable element, and the 2015 reappraisal in Emerging Themes in Epidemiology by Fedak and colleagues notes the general agreement on that point.
Specificity was the weakest of the nine for most modern questions. Hill's era imagined a cleaner correspondence between one exposure and one disease than the evidence generally supports. A single environmental exposure can plausibly affect several organ systems, and a single outcome usually has many contributing causes. Demanding a one-to-one relationship would discard a great deal of real causation. Fedak and colleagues argue that molecular data has begun to give specificity a different and more defensible meaning, discussed below.
What does each viewpoint actually ask?
Read as questions rather than criteria, the nine become more usable.
Strength asks whether the association is large enough that it would be difficult to produce with plausible confounding alone. A strong association is harder to explain away, but a weak one is not thereby non-causal. Many established causal relationships are modest in size.
Consistency asks whether the finding reappears across different investigators, populations, designs, and settings. This is one of the more informative viewpoints, because a result that repeats under varied conditions is less likely to be an artifact of any single study's peculiarities.
Biological gradient, often called dose response, asks whether more exposure corresponds to more of the outcome. Where present, a gradient is persuasive. Its absence is not disqualifying: some biological responses have thresholds below which nothing observable occurs, and some are non-monotonic, so a flat or irregular curve does not by itself refute causation.
Plausibility and coherence ask whether a causal reading fits what is already known, biologically and epidemiologically. Hill himself noted the limitation here, which is that plausibility is bounded by the biological knowledge of the day. An association can be real and causal while appearing implausible simply because the mechanism has not yet been described.
Experiment asks whether intervening changes the outcome. This is the strongest evidence available when it can be obtained ethically, which in environmental health is often not the case. Natural experiments, such as a regulatory change that reduces population exposure, sometimes stand in for it.
Analogy is the loosest of the nine, asking whether comparable exposures have produced comparable effects. It widens the imagination rather than settling anything.
How has molecular data changed how the viewpoints are used?
Fedak and colleagues argue that the arrival of biomarkers, mechanistic toxicology, and molecular measurement has changed what several of the viewpoints can contribute. Plausibility and coherence in particular now have far more to work with than Hill had available. Where he could appeal only to general biological credibility, a modern argument can often point to a measured intermediate step between exposure and outcome.
They also reconsider specificity. Rather than requiring one exposure to map to one disease, molecular evidence can demonstrate specificity at the mechanistic level, identifying a particular biological pathway through which an exposure acts. That is a more defensible version of the same idea.
Two cautions follow from their analysis. The first is that better mechanistic evidence does not lower the bar anywhere else. Knowing a plausible pathway does nothing about confounding, selection, or exposure misclassification, which remain problems of study design and analysis. It strengthens one part of the argument and leaves the rest intact.
The second is the checklist problem stated plainly. As the authors put it, Hill's nine aspects of association "were never intended to be viewed as rigid criteria or as a checklist for causation, yet have been popularized as such over the past 50 years." They warn that mechanical application invites a form of inductivism, in which the available data are read to fit the criteria rather than used to test the hypothesis. That warning is Hill's own, restated with fifty years of misuse behind it.
What the viewpoints cannot do
Several limits are worth stating explicitly.
They cannot convert a confounded association into a causal one. If a third factor drives both the exposure and the outcome, no amount of consistency or analogy repairs that. Confounding is addressed through design and analysis, not by satisfying other viewpoints.
They do not produce a score. Counting how many of the nine are met is not a method, and a study that satisfies six is not therefore stronger than one satisfying four. The viewpoints differ in weight, in kind, and in relevance depending on the question.
They are not a decision rule for an individual. A population-level causal claim, that an exposure raises risk across a group, is a different statement from a claim about what caused a particular person's illness. The viewpoints were designed for the former.
This article is educational and is not a substitute for personal medical advice. Questions about an individual exposure or health concern belong with a qualified clinician.
A better question to ask of a headline
For a general reader encountering a claim that some exposure causes some outcome, the useful move is not to run down the nine viewpoints. It is to ask Hill's underlying question directly: what would have to be true for this association to be something other than cause? Then ask whether the study design could have ruled that out.
If the answer is that the two groups differed in ways nobody measured, or that the outcome could have preceded the exposure, or that only one study has ever reported it, the association may still be real and still be worth investigating. It is simply not yet a demonstrated cause. That distinction is what Hill's address was written to preserve, and it is the standing subject in research on autism and the environment, where exposures are common, outcomes are multifactorial, and the temptation to read association as causation is strongest.
Readers interested in the underlying methodological work may wish to consult the peer-reviewed literature in this area.