All insights

When there's no head-to-head: a field guide to indirect treatment comparisons

Two drugs. One question: which one works better?

The cleanest answer is a head-to-head trial. Randomise patients to A or B, follow them, compare. Job done.

But that trial often doesn’t exist. Maybe your drug was tested on its own in a single-arm trial. Maybe it was compared with placebo, not with what patients actually receive. Maybe the competitor launched after your trial began.

So you compare across studies instead. That is an indirect treatment comparison, or ITC.

Done well, an ITC can carry a health technology assessment (HTA) dossier. Done badly, it misleads quietly and confidently. This is a field guide to the why, the what, the when, and the how. That last one, the how, is the part we have spent months living inside.

Why: when the trial you need doesn’t exist

The short version: you reach for an ITC when there is no head-to-head trial to reach for.

Three situations come up again and again.

Your evidence is a single-arm trial. Common in oncology and rare disease, where randomising patients against placebo is neither practical nor ethical. There is no comparator arm at all, only your treatment and its outcomes.

You have a randomised trial, but against the wrong comparator. Your trial used placebo. Or a drug nobody uses any more. Or a standard of care that differs from the market you are trying to enter.

The competitor moved. A new treatment became standard after your pivotal trial locked, so the two were never tested against each other.

In all three, the comparison a payer wants is not the comparison your trial ran. An ITC borrows evidence from other studies to build the missing bridge.[2,8]

What: the toolbox, from blunt to brilliant

ITCs split into two families, and that split matters more than any single method.

Anchored comparisons keep a common comparator. Both trials share an arm, usually placebo or a shared control. You compare A with B by travelling A → shared arm → B, and randomisation is preserved along the route.[1]

Unanchored comparisons have no shared arm. A single-arm study, or two trials with nothing in common. You compare absolute outcomes head-on, and randomisation is broken.[1,2]

That one question, whether there is a common anchor, drives almost everything about how believable the result is. We come back to it under “how.”

Now the methods themselves, roughly from simplest to most ambitious.

Anchored, aggregate data only:

  • Bucher: the classic. Two trials, one shared comparator, subtract one relative effect from the other. Simple and transparent.[4]
  • Network meta-analysis (NMA): Bucher’s ambitious sibling. Many trials, many treatments, one connected web, all compared at once.[2]

Anchored, with patient-level data on one side (population adjustment):

  • MAIC (matching-adjusted indirect comparison): reweight your patients so their characteristics match the average patient in the other trial, then compare.[5]
  • STC (simulated treatment comparison): build an outcome model from your patients, then predict what would have happened in the other trial’s population.[1]
  • ML-NMR (multilevel network meta-regression): the general form. It combines patient-level and aggregate data across a whole network and adjusts for the differences between them.[6]

Unanchored (no common comparator):

  • The same MAIC, STC and ML-NMR machinery can run unanchored, but now it carries a far heavier load of assumptions (see “how”).[1,3]
  • ML-UMR (multilevel unanchored meta-regression): a newer Bayesian method built specifically for disconnected evidence. It extends ML-NMR into the unanchored world and models patient-level and aggregate data together. Usefully, it separates the assumptions needed to identify a treatment effect from those needed to transport it to a different population. It does not make the assumptions disappear. It makes them explicit.[7]

And the blunt instrument:

  • The crude (naïve) comparison: line up outcomes from two studies and compare, with no adjustment at all. Mostly discouraged, for the obvious reason that any difference might be the treatment, or might just be that the patients differ. NICE guidance is explicit that naïvely pooling this kind of evidence is not recommended, and EU guidance advises against aggregate-data indirect comparisons in disconnected networks.[2,3] But it is not always dismissed, and the FOCUS trial is a clean example of when it flies. The melphalan liver-directed therapy (HEPZATO in the US, CHEMOSAT in Europe) was tested single-arm in 91 patients with metastatic uveal melanoma. Its success was defined in advance against a benchmark: a meta-analysis of historical therapies for the disease, pooling a 5.5% response rate (95% CI 3.6–8.3%). The trial came in at 36.3%. The decisive part: the lower bound of its confidence interval (26.4%) sat well above the upper bound of the benchmark’s (8.3%). Non-overlapping intervals, by design. The study met its primary endpoint on exactly that basis, and the FDA approved it in 2023.[9,12] The lesson: a crude comparison can clear a very high bar, but only when the signal is so loud that population differences cannot plausibly explain it. That is the exception, not the rule.

(A note on the map: van Beekhuizen and colleagues’ supplementary overview lays these methods out neatly, though it stops at MAIC, STC and ML-NMR; ML-UMR is newer than that list.)[7,8]

When: regulators dip in, payers depend on it

Two audiences, two very different appetites.

Regulators (FDA, EMA) can accept a single-arm trial as pivotal evidence, but they are wary, and it is worth understanding why. The central problem of a single-arm trial is attribution: with no randomised control, it is hard to be sure the outcome was caused by the treatment rather than by the natural course of the disease, the way patients were selected, regression to the mean, confounding, or plain chance.[13,14] Randomisation is what normally rules those out. A single arm leaves them all room to hide.

So the bar is high, and it tends to be cleared in one of two ways. First, a dramatic, rapid effect unlikely to have happened on its own. The EMA names this exception explicitly,[13] and it is why response rate works as an oncology endpoint at all: tumours rarely shrink without help.[14] That is the FOCUS logic above. Second, a carefully built external control. Here regulators lean toward patient-level comparator data: the FDA’s guidance on externally controlled trials is written around patient-level external controls, explicitly sets aside summary-level ones, and insists the control arm and the analysis be fixed before the trial reports.[14] The aggregate-data indirect comparisons at the heart of this article, the MAIC and STC world, are more an HTA instrument than a regulatory one.

HTA bodies and payers are where ITCs really live. Their entire job is comparative: is this better than what we already fund, and by how much? When no head-to-head trial exists, an ITC is often the only way to answer.[2,8]

And the bar is rising. Under the EU Joint Clinical Assessment, live since January 2025, a single dossier can face comparator questions from up to 27 countries at once. Many of those comparisons have no head-to-head trial behind them, so indirect comparisons move from the appendix to centre stage.[8] (We covered the JCA and its 100-day clock in our first insight.)

How: the difference between evidence and advocacy

Here is where months of building a blueprint earn their keep. The methods are well documented, software exists, and the maths is not the hard part. What separates a robust ITC from a fragile one is a handful of things people routinely underweight.

Know which assumption you are signing. Anchored comparisons need the effect modifiers to be balanced or adjusted for, because the shared arm handles the rest. Unanchored comparisons need every prognostic factor and every effect modifier accounted for, because there is no anchor to cancel them out. TSD 18 calls this assumption “very strong, and largely considered impossible to meet.”[1] The EU guideline is just as blunt: unanchored comparisons rest on “conditional constancy of absolute effects,” which is “very unlikely” to hold.[2] You cannot verify it, so you plan around it.

Match on the right variables, all of them when unanchored. Both guidelines say the same thing: population adjustment only works if you have measured every covariate that matters, and that set is “often unverifiable and unattainable.”[1,2] A missing prognostic factor does not announce itself. Deciding the covariate list up front, from clinical reasoning rather than convenience, is half the battle.

Watch how much of your trial survives the adjustment. When you reweight patients in a MAIC, some count for a great deal and some for almost nothing. The effective sample size, meaning how many patients your comparison effectively rests on, can collapse far below the number you started with. In practice it is driven mostly by whichever matching variable is scarce in your trial but common in the comparator population, and no weighting scheme does much to escape that, so you can often anticipate which variable will strain the comparison before you fit a single model. A result built on an effective handful of patients is fragile, however tidy the point estimate looks. Report the weights, the population overlap, and the effective sample size. Every time.[1,3]

Be most suspicious of the diagnostics that look best. Here is the trap that catches people. If the other trial contains types of patient yours does not contain at all, that gap leaves no fingerprint: the weights look sensible, the effective sample size looks fine, and the balance table looks immaculate, while the comparison quietly extrapolates into a void.[11] So check overlap jointly, cross-tabulating the patient characteristics rather than eyeballing them one at a time. And treat the tidy post-matching balance table with suspicion: matching forces it to look perfect by construction, so it is decoration, not proof.[1]

Pick an outcome measure that can actually hold. Hazard ratios assume the two survival curves stay in fixed proportion over time, and that assumption often does not hold in practice. A common alternative is restricted mean survival time, the average survival up to a set time horizon. It does not rest on the proportional-hazards assumption, and it reads in plain months rather than as a ratio.[10]

Quantify the bias you cannot remove. Because the central assumption is unverifiable, the honest move is to ask a different question: how large would an unmeasured factor have to be to overturn the result? That is a quantitative bias analysis, and it converts “we hope nothing is missing” into a number a reviewer can actually weigh.[1]

Prespecify, then resist improvising. Regulators and HTA guidance both want the method, the covariates and the sensitivity analyses fixed before you see the answer.[2,14] Analytical freedom after the fact is exactly how a comparison talks itself into whatever conclusion it was hoping for.

None of these are exotic. They are simply the difference between an ITC that reads as evidence and one that reads as advocacy.

Take home

  • When there is no head-to-head trial, an indirect comparison borrows evidence across studies to fill the gap.
  • The first fork is anchored (a shared comparator, randomisation preserved) versus unanchored (no shared arm, randomisation broken). Unanchored is far more demanding.
  • The toolbox runs from Bucher and NMA, through MAIC, STC and ML-NMR, to newer methods such as ML-UMR, plus the crude comparison, which only convinces when the effect is enormous.
  • Regulators occasionally approve on indirect or single-arm evidence. HTA bodies and payers rely on it routinely, and the EU JCA is pushing it centre stage.
  • A robust ITC lives or dies on its assumptions, its covariate list, its effective sample size, its outcome measure, its honesty about the bias it cannot remove, and a healthy distrust of diagnostics that look too clean.

This is the kind of work we do at Evidax: choosing the right comparison, building it on assumptions we can defend, and stress-testing it before an assessor does. If your evidence base is missing its head-to-head, that gap is worth understanding early, because everything downstream depends on it.

References

  1. Phillippo DM, Ades AE, Dias S, Palmer S, Abrams KR, Welton NJ. NICE DSU Technical Support Document 18: Methods for Population-Adjusted Indirect Comparisons in Submissions to NICE. NICE Decision Support Unit; 2016. https://www.sheffield.ac.uk/nice-dsu/tsds/population-adjusted

  2. Member State Coordination Group on Health Technology Assessment. Methodological Guideline for Quantitative Evidence Synthesis: Direct and Indirect Comparisons. 8 March 2024. https://health.ec.europa.eu/publications/methodological-guideline-quantitative-evidence-synthesis-direct-and-indirect-comparisons_en

  3. Welton NJ, Phillippo DM, Owen R, et al. CHTE2020 Sources and Synthesis of Evidence: Update to Evidence Synthesis Methods. NICE Decision Support Unit; 2020. https://www.sheffield.ac.uk/nice-dsu/methods-development

  4. Bucher HC, Guyatt GH, Griffith LE, Walter SD. The results of direct and indirect treatment comparisons in meta-analysis of randomized controlled trials. Journal of Clinical Epidemiology. 1997;50(6):683-691. https://doi.org/10.1016/S0895-4356(97)00049-8

  5. Signorovitch JE, Wu EQ, Yu AP, et al. Comparative effectiveness without head-to-head trials: a method for matching-adjusted indirect comparisons applied to psoriasis treatment with adalimumab or etanercept. PharmacoEconomics. 2010;28(10):935-945.

  6. Phillippo DM, Dias S, Ades AE, et al. Multilevel network meta-regression for population-adjusted treatment comparisons. Journal of the Royal Statistical Society: Series A. 2020;183(3):1189-1210. https://doi.org/10.1111/rssa.12579

  7. Chandler C, Ishak J. Anchors Away: Navigating Unanchored Indirect Comparisons with Multilevel Unanchored Meta-Regression (ML-UMR). Preprint. 2026. arXiv:2606.20341.

  8. van Beekhuizen S, Che M, Monfort L, et al. Indirect treatment comparisons in EUnetHTA relative effectiveness assessments: learnings and recommendations for the implementation of EU joint clinical assessments. PharmacoEconomics Open. 2025;9:597-609. https://doi.org/10.1007/s41669-025-00575-1

  9. US Food and Drug Administration. FDA approves melphalan for liver-directed treatment of uveal melanoma. August 2023. https://www.fda.gov/drugs/resources-information-approved-drugs/fda-approves-melphalan-liver-directed-treatment-uveal-melanoma

  10. Royston P, Parmar MKB. Restricted mean survival time: an alternative to the hazard ratio for the design and analysis of randomized trials with a time-to-event outcome. BMC Medical Research Methodology. 2013;13:152. https://doi.org/10.1186/1471-2288-13-152

  11. Petersen ML, Porter KE, Gruber S, Wang Y, van der Laan MJ. Diagnosing and responding to violations in the positivity assumption. Statistical Methods in Medical Research. 2012;21(1):31-54. https://doi.org/10.1177/0962280210386207

  12. Zager JS, Orloff M, Ferrucci PF, et al. Efficacy and safety of the melphalan/Hepatic Delivery System in patients with unresectable metastatic uveal melanoma: results from an open-label, single-arm, multicenter phase 3 study. Annals of Surgical Oncology. 2024;31(8):5340-5351. https://doi.org/10.1245/s10434-024-15293-x

  13. European Medicines Agency, Committee for Medicinal Products for Human Use. Reflection Paper on Establishing Efficacy Based on Single-Arm Trials Submitted as Pivotal Evidence in a Marketing Authorisation. Draft for consultation. 21 April 2023. https://www.ema.europa.eu/en/documents/scientific-guideline/draft-reflection-paper-establishing-efficacy-based-single-arm-trials-submitted-pivotal-evidence-marketing-authorisation_en.pdf

  14. US Food and Drug Administration. Considerations for the Design and Conduct of Externally Controlled Trials for Drug and Biological Products. Draft Guidance for Industry. February 2023. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/considerations-design-and-conduct-externally-controlled-trials-drug-and-biological-products

If your pivotal trial leaves HTA and payer questions unanswered, we should talk.