Dr. Manouchehr Hessabi
← All writing
7 min readresearch methods · biomarkers · clinical trials

Validation and qualification: what a biomarker has to earn

Biomarker validation and qualification answer different questions. What context of use means, and why a reliable measurement can still be the wrong endpoint.

By Manouchehr Hessabi, MD, MPH

A biomarker that measures something accurately has not yet earned the right to stand in for anything. Those are two separate achievements, and confusing them is one of the more expensive mistakes in translational research.

The formal vocabulary for keeping them apart comes from the FDA-NIH Biomarker Working Group, which published the BEST Resource in 2016 as a shared glossary. Three of its terms do most of the work: analytical validation, clinical validation, and qualification. A fourth, context of use, quietly governs all three.

What is a biomarker, exactly?

BEST defines a biomarker as "a defined characteristic that is measured as an indicator of normal biological processes, pathogenic processes, or biological responses to an exposure or intervention, including therapeutic interventions."

Two things in that definition are easy to skim past. The characteristic must be defined, meaning the thing being measured is specified precisely rather than gestured at. And it is an indicator, which is a claim about what the measurement stands for, not merely what it is.

A biomarker is also not an endpoint. BEST keeps a separate category for a surrogate endpoint: "an endpoint that is used in clinical trials as a substitute for a direct measure of how a patient feels, functions, or survives." Every surrogate endpoint is a biomarker. Very few biomarkers ever become surrogate endpoints, and the distance between those two states is the subject of this article.

BEST also sorts biomarkers by the job they do: susceptibility or risk, diagnostic, monitoring, prognostic, predictive, response, safety, and multicomponent. A biomarker that performs well in one of those roles has demonstrated nothing about the others.

Validation and qualification answer different questions

Analytical validation asks whether the measurement works. BEST defines it as "a process to establish that the performance characteristics of a test, tool, or instrument are acceptable in terms of its sensitivity, specificity, accuracy, precision, and other relevant performance characteristics." This is a laboratory question. Does the assay detect the analyte it claims to detect, at the concentrations that matter, reproducibly, across operators and instruments and days?

Clinical validation asks whether the measurement means anything. BEST defines it as "a process to establish that the test, tool, or instrument acceptably identifies, measures, or predicts the concept of interest." Here the question moves out of the laboratory. Does this quantity actually track the biological or clinical state it is supposed to represent, in the population where it will be used?

An assay can pass the first and fail the second completely. A protein can be measured with excellent precision and still bear no useful relationship to the disease it was hoped to reflect. The reverse is also possible in practice: a measurement that plainly associates with an outcome in a research setting, but with an assay too unreliable to act on.

Qualification is a third thing, and it is regulatory rather than scientific. BEST defines it as "a conclusion, based on a formal regulatory process, that within the stated context of use, a medical product development tool can be relied upon to have a specific interpretation and application in medical product development and regulatory review."

The phrase carrying the weight there is within the stated context of use. Qualification is never a general endorsement of a biomarker. It is a verdict about one specified use.

What context of use actually means

BEST defines context of use, usually abbreviated COU, as "a statement that fully and clearly describes the way the medical product development tool is to be used and the regulated product development and review-related purpose of the use."

In practice a COU statement names the disease and population, the stage of development, the specific decision the biomarker will inform, and the interpretation attached to a given result. It is deliberately narrow. A biomarker qualified to enrich a trial population for a particular condition has not thereby been qualified to select therapy for individual patients, or to serve as the endpoint a drug is approved on.

The BEST Resource is explicit that fitness does not travel. A tool "that has been determined to be fit-for-purpose for a non-regulatory use, or a given regulatory purpose, is not necessarily sufficient for another regulatory purpose, such as those needed for marketing approval or qualification."

This is the single most useful idea in the whole framework, and it generalizes well past regulatory science. Every validated measurement carries an invisible footnote describing the conditions under which it was validated. Reading a biomarker result without reading that footnote is where most overinterpretation begins.

Why a reliable measurement can still be the wrong one

The clearest illustration of what goes wrong when a biomarker is trusted outside its evidence comes from cardiology.

Ventricular ectopy, meaning extra irregular heartbeats, is easy to measure and is associated with sudden cardiac death after a heart attack. Drugs that suppressed those extra beats were therefore expected to reduce deaths. The reasoning was that the suppressible thing and the fatal thing were tightly linked, so acting on the first should improve the second.

The Cardiac Arrhythmia Suppression Trial tested that hypothesis directly. As Echt and colleagues reported in the New England Journal of Medicine in 1991, 1,498 patients whose ectopy had been successfully suppressed were randomly assigned to continue active drug or receive placebo. After a mean follow-up of 10 months, 89 patients had died. Fifty-nine deaths were attributed to arrhythmia, 43 in the drug groups against 16 on placebo. Twenty-two were nonarrhythmic cardiac deaths, 17 against 5. The encainide and flecainide arms were stopped because of excess mortality.

The biomarker had behaved exactly as advertised. The drugs suppressed the extra beats, the measurement was accurate, and the association between ectopy and death was real. What failed was the assumption that the association ran through a pathway the drug would improve rather than worsen. The authors noted plainly that the mechanism behind the excess mortality remained unknown.

This is the difference between a biomarker being correlated with an outcome and being a valid substitute for it. An intervention can move the marker in the desired direction while moving the patient in the opposite one.

How qualification actually works

The FDA's Biomarker Qualification Program provides a formal route for establishing a biomarker's value for a stated context of use in drug development and regulatory review. Under the process specified in statute, a submission proceeds through three stages: a Letter of Intent, a Qualification Plan, and a Full Qualification Package.

The structure reflects the difficulty of the task. The Letter of Intent proposes the biomarker and the COU. The Qualification Plan sets out what evidence would be sufficient to support that specific use. Only then does the full evidentiary package get assembled and reviewed. Qualified biomarkers, along with submissions still in progress, are published in the agency's Drug Development Tool qualification database, so the record of what has been qualified, and for what, is public.

The payoff is that once a biomarker is qualified for a COU, other developers can rely on it for that same use without re-establishing the case each time. The constraint is that the qualification travels no further than the COU it was granted for.

Questions worth asking of any biomarker claim

When a study or a press release reports that a biomarker predicts, detects, or tracks something, four questions separate a strong claim from a fragile one.

  • Which validation is being claimed? Analytical performance and clinical meaning are frequently reported together in a way that lets the stronger evidence cover for the weaker.
  • In which population was it established? A marker validated in a specialty referral population may behave very differently in general practice, where the underlying prevalence is lower.
  • What decision is it being used to make? Enriching a trial, monitoring a known condition, and approving a drug are three different burdens of proof.
  • Is the marker being treated as a substitute for the outcome, or as a signal alongside it? The first requires evidence that intervening on the marker changes the outcome, which is a much higher bar than association.

None of this is an argument against biomarkers. Modern trials would be slower, larger, and less informative without them. It is an argument for reading the context of use as carefully as the result, because a biomarker is only ever valid for the question it was tested against.

Further reading on study design and measurement in clinical research is available through the peer-reviewed publications listed on this site.

About the author. Dr. Manouchehr Hessabi is a physician-epidemiologist and Senior Research Scientist at the BERD core of UTHealth Houston's Center for Clinical and Translational Sciences. See his peer-reviewed publications or research programs.