Dr. Manouchehr Hessabi
← All writing
8 min readresearch methods · epidemiology · study design

Selection bias: when who gets studied changes the answer

Selection bias explained with examples: how the way people enter a study can create or reverse an association, and what careful researchers do about it.

By Manouchehr Hessabi, MD, MPH

Readers who have absorbed one lesson about observational research have usually absorbed this one: correlation is not causation, because some third factor may explain the link. That third factor is a confounder, and the standard reply is statistical adjustment.

There is a second reason correlation is not causation, and it is less familiar and harder to fix. A study can only speak about the people who are in it. If the way people came to be in it depends on both the thing being studied and the outcome being measured, the study can display an association that does not exist in the wider population, or hide one that does. Adjustment does not repair this, because the damage was done before the analysis began.

This is educational and not a substitute for personal medical advice.

What is selection bias, in the structural sense?

Four terms first. An exposure is whatever might influence health: a metal in drinking water, a medication, a body-mass category. An outcome is the health state being studied. The analytic sample is the set of people whose data actually enter the analysis. The target population is the group the researchers want their conclusion to apply to.

The definition epidemiologists now use came from a 2004 paper by Hernán, Hernández-Díaz, and Robins in Epidemiology, which argued that a range of seemingly unrelated biases share one causal structure. Their conclusion was that selection bias arises from conditioning on a common effect of two variables, one of which is the exposure or a cause of it, and the other the outcome or a cause of it.

Conditioning means restricting attention to people with a particular characteristic, which is exactly what happens when a study enrolls only people who were tested, admitted, imaged, or still enrolled at the end of follow-up. A variable that is a common effect of two others is called a collider, because two causal arrows collide in it.

The paper draws the contrast that makes the concept usable. Confounding comes from a common cause of exposure and outcome, a variable with arrows pointing out toward both. Selection bias comes from a common effect, a variable with arrows pointing in from both. The same authors note that this structure underlies both inappropriate selection of controls in case-control studies and informative censoring in cohort studies, which had traditionally been taught as separate problems.

One refinement is worth a footnote. A 2024 commentary in the American Journal of Epidemiology by Lu, Gonsalves, and Westreich argued that restricting the sample to one level of a collider is selection bias in the strict sense, while putting a collider into a regression model is better described as overadjustment bias, because adjustment does not involve selecting a sample. The distinction is terminological rather than practical. Both distort the estimate.

How can the tested sample invent a risk factor?

The clearest recent demonstration came out of the early COVID-19 literature. Many studies of who was at risk were built on people who had been tested for infection or admitted to hospital, because those were the people whose data existed.

Griffith and colleagues, writing in Nature Communications in 2020, pointed out the difficulty. Being sampled at all was itself an outcome influenced by many things, so restricting analysis to that group could induce associations between variables that affect the likelihood of being sampled. They then showed it empirically. In UK Biobank, the participants who had been tested for COVID-19 were, compared with the wider cohort, highly selected for a range of genetic, behavioural, cardiovascular, demographic, and anthropometric traits.

That is the mechanism in one sentence. If both a suspected risk factor and the severity of illness influenced whether a person was tested, then within the tested group those two things will appear related to each other whether or not they are related in the population.

Their recommendation is the part most often skipped. Collider bias should be explored in existing studies, but the authors are explicit that the optimal way to mitigate the problem is to use appropriate sampling strategies at the study design stage. This is a design problem with a design solution.

Can the same data give opposite answers depending on the control group?

It can, and a 2026 analysis in the American Journal of Epidemiology by Ueland and colleagues showed it inside a single dataset.

Using de-identified electronic health record data, they examined metabolic exposures and rotator cuff tears. Because imaging is often required to diagnose a tear, some researchers argue that controls should also be required to have imaging, so that cases and controls are held to the same diagnostic standard. That sounds like methodological care. It is also conditioning on a descendant of a collider, because imaging is ordered in response to symptoms, and symptoms sit downstream of both the exposure and the outcome.

The consequence was not subtle. Type 1 diabetes was positively associated with tears using controls without imaging, at an adjusted odds ratio of 1.78 (95% CI 1.64 to 1.92), but inversely associated using controls with imaging, at 0.75 (95% CI 0.57 to 0.97). Same data, same exposure, same outcome. One rule about who counts as a control, and the association changes sign.

The scope of that result has to be stated plainly. It is a demonstration of a methodological mechanism inside one health-record dataset. It is not a clinical finding about diabetes and shoulders, and neither number should be read as an estimate of anyone's risk.

For readers who want the underlying logic of how cases and controls are assembled in the first place, that is treated separately in the explainer on case-control and cohort designs.

Can selection alone manufacture a disparity?

This is where the concept moves beyond clinical epidemiology.

A study published in Social Science and Medicine in July 2026 by Bashir and Al-Kassab-Córdova examined whether measured differences between social groups might reflect how people entered the sample rather than any underlying social mechanism. Using directed acyclic graphs, which are formal diagrams of assumed causal relationships, together with simulations, they showed that conditioning on a selection variable influenced jointly by social characteristics and the outcome induces collider bias.

The sharpest result is what happened when they built a simulated world containing no real association at all. Even there, restricting analysis to a selected sample produced apparent inequality, and the authors report that such bias can generate, amplify, attenuate, or reverse apparent differences. They also note that the distortion is not confined to the interaction terms researchers usually examine, since it propagates to general descriptions of inequality as well.

The practical reading is disciplined rather than dismissive. A difference measured inside a clinic population, a volunteer survey, or a disease registry is a statement about that sample until the selection process has been accounted for. It is not evidence that the difference is unreal. It is a reason the sample cannot settle the question by itself.

How do researchers measure what the missing people could have done?

The honest answer is that they cannot observe the missing people, so they reason about them explicitly instead. The method is quantitative bias analysis: state several plausible assumptions about the unobserved group, recompute the estimate under each, and report the range.

Bokern and colleagues applied it in a 2026 paper in BMC Medical Research Methodology. During the pandemic, some patients with severe COVID-19 were never admitted to hospital, so a study restricted to hospitalized patients was missing part of its own target population. Working within a cohort of people with chronic obstructive pulmonary disease, they compared two inhaler regimens and reported that the odds ratio after inverse probability of treatment weighting was 1.01 (95% CI 0.59 to 1.72).

They then recomputed it under four scenarios that varied the assumed death rates among the patients who were never hospitalized. The corrected odds ratios ranged from 0.81 (95% CI 0.52 to 1.23) to 1.28 (95% CI 0.83 to 2.00). Their stated conclusion was that the scenarios were in line with the null hypothesis, that the confidence intervals were wide, and that death rates in the non-hospitalised would have needed to be substantially different in the treatment groups to change the study conclusions.

That is what a useful bias analysis looks like. It does not rescue the estimate or prove it. It states how much unobserved reality would have to differ before the reported answer stops holding, and lets the reader judge whether that much difference is plausible.

What should a careful reader look for?

Two questions, and they can be asked of any observational paper before a single result is read.

  • How did these people get into this study? Were they tested, admitted, imaged, referred, still enrolled at the end, or willing to volunteer? Each of those is a filter, and a filter influenced by both the exposure and the outcome is a collider.
  • Did the authors say what the people who did not get in could have done to the estimate? A sensitivity analysis or quantitative bias analysis is a sign that the question was taken seriously.

If neither is addressed, the association reported is a fact about the sample, and its extension to anyone else is an assumption rather than a finding. The verb still matters as well: observational designs of this kind establish that things are associated with one another, and the move to causation requires ruling out both confounding, treated separately in the explainer on correlation, causation, and confounding, and the structural problem described here.

It is worth noticing that a definition published in 2004 is still being actively refined more than twenty years later, in commentaries about terminology, in health-record demonstrations, and in simulation work on social inequality. That is not a sign of confusion. It is what a genuinely difficult problem looks like while it is being worked out, and it is a reason to read the methods section of a study with the same attention usually reserved for its conclusion.

Readers who prefer the primary literature to summaries of it may find the peer-reviewed publications listed on this site a reasonable starting point.

About the author. Dr. Manouchehr Hessabi is a physician-epidemiologist and Senior Research Scientist at the BERD core of UTHealth Houston's Center for Clinical and Translational Sciences. See his peer-reviewed publications or research programs.