All insights

Heterogeneity in indirect comparisons: your design matters

You have built your indirect comparison. Your trial on one side, the best comparator evidence you could find on the other, adjusted as well as the data allows.

The assessor’s first objection is that the two groups of patients are not alike enough for the comparison to mean anything.

That objection is heterogeneity. Together with risk of bias, it forms the most common criticism made of indirect comparisons in Europe. And by the time it arrives, its causes are already fixed in the design.

Heterogeneity sits inside the most common objection

When an HTA body assesses a medicine, it publishes an assessment report. The report sets out what the committee accepted, what it doubted, and why. Those reports are public, and researchers have begun reading them in bulk.

The findings are consistent.

A 2024 review of five European agencies found heterogeneity and risk of bias raised in 48% of reports.[1] A 2026 review of submissions to the French Transparency Committee put the combined figure at 59%.[2] For external control arms it was 47% in England, 56% in France and 58% in Germany.[3]

Missing or unclear data runs it close, at 43% in the same 2024 review.[1] For external control arms in Germany that criticism is raised more often still, at 64%.[3]

Better methods do not make it go away

The same 2024 review looked at whether the criticism depended on the technique used. It did not.

Heterogeneity and risk of bias were raised against 54% of network meta-analyses. They were raised against 67% of matching-adjusted indirect comparisons, a method built specifically to correct for population differences.[1]

A more sophisticated method did not answer the objection. In that review it drew the criticism more often, not less.

Agencies expect the groundwork first

Guidance from HAS, NICE and Germany’s IQWiG asks for systematic identification of the evidence, and for potential confounders, effect modifiers and prognostic variables to be identified in advance.[2]

Only 37% of the French comparisons had been predefined in a protocol.[2]

And in a review of NICE technology appraisals, not one submission using an anchored comparison had justified its assumptions about effect modification in advance.[4]

None of that is a complaint about the statistics. Each of those concerns starts with a decision taken before any statistics were run.

Some of it is two studies answering different questions

Cochrane splits heterogeneity three ways: differences in patients and treatments, differences in how studies were designed and run, and the statistical variation that follows.[5]

Cochrane also draws a distinction that gets missed. Where heterogeneity comes from differences in design and conduct, it “suggests that the studies are not all estimating the same quantity, but does not necessarily suggest that the true intervention effect varies.”[5]

Two studies can differ because the drug genuinely works differently in different patients. Or they can differ because they were never designed to answer the same clinical question.

Population adjustment was built for the first situation. It cannot turn two different clinical questions into the same one.

Two trials can look identical and still answer different questions

The precise description of what a study is designed to estimate is called its estimand. It is the clinical question written down in a form specific enough that two teams would build the same trial from it.

Specifying an estimand means stating five things. Which patients. Which treatments are being compared. Which outcome is measured. How events after treatment starts are handled. And how the result is summarised across the population.[6]

That fourth attribute is the one that catches people out. Events after treatment starts are called intercurrent events. A patient switches therapy. A patient stops early. A patient receives something else afterwards.

Two trials can share a drug, a population and an endpoint, and still answer different questions. One trial counted deaths whatever treatment followed. The other stopped counting at the switch.

No baseline characteristics table will show you that difference. And most studies do not clearly define the research question, so readers have to infer it from the trial design, the handling of the data and the analysis.[7]

If two studies are answering different questions, matching their populations makes the numbers more comparable and leaves them just as wrong.

Baseline tables do not establish comparability

Baseline characteristics matter too, and here the checks are more familiar. Faced with two populations, most submissions do the same three things. Put the baseline characteristics side by side. Report standardised differences. After weighting, report how many patients are effectively left.

All three are worth doing. None of the three establishes comparability.

Those checks cover the characteristics you happened to measure, on the scale you happened to pick. About the rest they are silent. The FDA makes the point from the design side: important prognostic factors may not be known at all, and what is not known cannot be used to build an external control arm that matches your patients.[8]

Overlap is the harder limit. Adjustment reweights your patients to resemble the comparator’s. Where the two populations barely overlap there is little to reweight towards, and NICE says these approaches are unlikely to perform well when the target population is poorly represented in the sample you are analysing.[9] The FDA makes the same point by degree: controlling for differences becomes more challenging the more dissimilar the populations, and differences in calendar time are hard to address by analysis alone.[8]

Hazard ratios move when the patients differ, even if the drug does not

One problem survives perfect balance.

Odds ratios and hazard ratios are non-collapsible. The population-level number is not an average of the individual patients’ numbers. So a factor that predicts outcome, without changing anyone’s response to treatment, still moves the result you report.[10]

Survival is worse. Where prognostic factors affect the hazard, the population-level hazard ratio must change across follow-up even when the effect in each individual patient is constant. The marginal hazard averages over whoever is still at risk, and that mix keeps changing.[10]

Two studies of the same drug, with the same effect in every patient, will report different hazard ratios if their case mix, their follow-up or their censoring differs.[10]

Report absolute effects alongside relative ones, and say which population your estimate belongs to.

Three things to do about it, in order

Heterogeneity is not one problem, so there is no single fix. There is a sequence, and the order is what makes it work.

  1. Define the estimand before you choose the data.
  2. Specify the target trial before you touch the data.
  3. Adjust last, once you know what adjustment can and cannot reach.

One. Define the estimand before you choose the data

Write down the five attributes. Population, treatments compared, outcome, intercurrent event strategies, summary measure.

Write them for your own study, and for every comparator you intend to use.

Differences that surface at this stage are the ones you can still design around. ICH E9(R1) is explicit that estimand definition comes before the choice of data and method of analysis, not after it.[6]

Two. Specify the target trial before you touch the data

Write the protocol of the randomised trial you would have run if you could. Eligibility, treatment strategies, how patients are assigned, when follow-up starts and ends, outcomes, the comparison you want, the analysis. Then state how each part maps onto the data you actually have.[11]

The component that does the most work is when follow-up starts. Follow-up has to begin at the moment eligibility is met and treatment is assigned.[11]

Get that wrong and you manufacture a difference that has nothing to do with the drug or the patients. One dataset on tocilizumab in covid-19 gave a mortality hazard ratio of 0.71 when the target trial was emulated properly. A simpler analysis of the same data gave 0.91. Two things had changed: patients were compared whenever they started treatment rather than at a common time zero, and clinical characteristics at time zero were not carefully adjusted for.[12] The postmenopausal hormone therapy literature contradicted its own randomised trial for years until the two analyses were aligned in the same way.[12]

The name for this class of problem is design-related bias. It arises from decisions the analyst makes, not from the patients and not from the treatment.[13] Specifying the target trial properly prevents it. Specifying the target trial does not remove confounding, and nobody claims it does.[12]

The specification earns its keep a second time when you go looking for comparator evidence. It gives you eligibility criteria for the search. A study that cannot be mapped onto your target trial is not a comparator, however similar its patients look, and you will have decided that before the results are in front of you.

Three. Adjust last, and know what it cannot reach

What you must adjust for depends on what the design already gives you.

In an anchored comparison, each trial has its own randomised control arm. Randomisation has balanced the prognostic factors within each trial, so population adjustment concentrates on the factors that modify the treatment effect, on the scale you are working on.[14]

In an unanchored comparison there is no such protection. You have to account for every factor that predicts the outcome as well. NICE Decision Support Unit Technical Support Document 18 is blunt about what that assumes: it is “largely deemed unreasonable (if it were, there would be no reason to undertake randomised controlled trials).”[14]

Two things adjustment cannot do.

Adjustment cannot reach differences in dosing, co-treatment, permitted switching, outcome definitions or study conduct. Those differences are bound up with the treatment itself, and no weighting scheme separates them.

Adjustment also always answers for a particular population. Matching-adjusted indirect comparison reweights your patients to resemble your competitor’s trial. The answer tells you how your drug would have performed in their study population, which is not the population your payer is deciding for.

If your comparator has to be assembled from several sources rather than lifted from one trial, you inherit the pooling problems too. A random-effects confidence interval tells you where the average sits, not how much the studies disagree.[5] Across 479 significant meta-analyses with any heterogeneity, 72% had a prediction interval that included no effect.[15] Report the prediction interval as well.

Take home

  • Heterogeneity and risk of bias together form the most frequent criticism of indirect comparisons in European HTA, and more sophisticated methods do not escape it.

  • Some of the disagreement between studies is two studies answering different clinical questions. Matching the populations does not fix a mismatch of questions.

  • Define the estimand first, and pay particular attention to intercurrent events. Two trials with the same drug, patients and endpoint can still be answering different questions.

  • Specify the target trial before you touch the data. Align eligibility, treatment assignment and the start of follow-up. In one published dataset, a proper emulation gave a hazard ratio of 0.71 where a simpler analysis with a misaligned time zero and weaker adjustment gave 0.91.

  • Use that same specification as the eligibility criteria when you search for comparator evidence.

  • Adjust last, and say which population your answer belongs to. An unanchored comparison assumes every prognostic factor and effect modifier has been measured, which is an assumption nobody would accept as a reason to skip randomisation.

  • Report what you could not adjust for, and how large a residual difference would have to be to change your conclusion.

Heterogeneity is not a statistic you calculate at the end. It is a set of choices you make at the beginning.

References

  1. Macabeo B, Rotrou T, Millier A, et al. The acceptance of indirect treatment comparison methods in oncology by health technology assessment agencies in England, France, Germany, Italy, and Spain. PharmacoEconomics Open. 2024;8(1):5-18. https://doi.org/10.1007/s41669-023-00455-6

  2. Monnereau M, Baschet L, Jarne A, et al. Methodological advances and challenges in indirect treatment comparisons: a review of international guidelines and Haute Autorité de Santé Transparency Committee case studies. Value in Health. 2026;29(6):1025-1033. https://doi.org/10.1016/j.jval.2025.12.013

  3. Monnereau M, Delord JP, Michiels S, et al. Acceptance of external control arm by HTA agencies: a retrospective review assessed in France, England, and Germany from 2021 to 2023. Value in Health. 2024;27(12):S2. Conference abstract, ISPOR Europe 2024.

  4. Phillippo DM, Dias S, Elsada A, et al. Population adjustment methods for indirect comparisons: a review of National Institute for Health and Care Excellence technology appraisals. International Journal of Technology Assessment in Health Care. 2019;35(3):221-228. https://doi.org/10.1017/S0266462319000333

  5. Deeks JJ, Higgins JPT, Altman DG, et al. Chapter 10: analysing data and undertaking meta-analyses. In: Cochrane Handbook for Systematic Reviews of Interventions, version 6.5. Cochrane, 2024. https://www.training.cochrane.org/handbook

  6. International Council for Harmonisation. Addendum on Estimands and Sensitivity Analysis in Clinical Trials to the Guideline on Statistical Principles for Clinical Trials, ICH E9(R1). 2019. https://database.ich.org/sites/default/files/E9-R1_Step4_Guideline_2019_1203.pdf

  7. Kahan BC, Hindley J, Edwards M, et al. The estimands framework: a primer on the ICH E9(R1) addendum. BMJ. 2024;384:e076316. https://doi.org/10.1136/bmj-2023-076316

  8. US Food and Drug Administration (CDER, CBER, Oncology Center of Excellence). Considerations for the Design and Conduct of Externally Controlled Trials for Drug and Biological Products. Draft Guidance for Industry. February 2023. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/considerations-design-and-conduct-externally-controlled-trials-drug-and-biological-products

  9. National Institute for Health and Care Excellence. NICE Real-World Evidence Framework. Corporate document ECD9. June 2022. https://www.nice.org.uk/corporate/ecd9

  10. Phillippo DM, Remiro-Azócar A, Heath A, et al. Effect modification and non-collapsibility together may lead to conflicting treatment decisions: a review of marginal and conditional estimands and recommendations for decision-making. Research Synthesis Methods. 2025;16:323-349. https://doi.org/10.1017/rsm.2025.2

  11. Matthews AA, Danaei G, Islam N, Kurth T. Target trial emulation: applying principles of randomised trials to observational studies. BMJ. 2022;378:e071108. https://doi.org/10.1136/bmj-2022-071108

  12. Hernán MA, Wang W, Leaf DE. Target trial emulation: a framework for causal inference from observational data. JAMA. 2022;328(24):2446-2447. https://doi.org/10.1001/jama.2022.21383

  13. Cashin AG, Hansford HJ, Hernán MA, et al. Transparent reporting of observational studies emulating a target trial: the TARGET statement. JAMA. 2025;334(12):1084-1093. https://doi.org/10.1001/jama.2025.13350

  14. Phillippo DM, Ades AE, Dias S, et al. NICE DSU Technical Support Document 18: Methods for Population-Adjusted Indirect Comparisons in Submissions to NICE. Decision Support Unit, December 2016. https://www.sheffield.ac.uk/nice-dsu

  15. IntHout J, Ioannidis JPA, Rovers MM, Goeman JJ. Plea for routinely presenting prediction intervals in meta-analysis. BMJ Open. 2016;6:e010247. https://doi.org/10.1136/bmjopen-2015-010247

If your pivotal trial leaves HTA and payer questions unanswered, we should talk.