External control arms: how and when
Your asset was tested without a control arm. Or there is randomised data, but not against the treatment patients actually receive in the market you plan to enter.
This is common. In rare disease and oncology, a randomised controlled trial can often be unethical or not feasible.
Either way, you are left without the comparison that regulators, HTA bodies and payers want to see. So you build one using data from outside the trial. That is an external control arm.
The first decision shapes everything that follows: what data do you build it from, patient-level data or summary data?
Regulators and HTA bodies accept external controls, with conditions
An external control arm compares outcomes in your treated patients against a group of people from outside your trial. The group can be treated or untreated, historical or concurrent, drawn from a previous clinical trial or from real-world data.[1]
The idea of an external control arm is long established. ICH E10, the foundational guidance on control groups, has recognised external controls for decades. It also warns they are usable only in unusual circumstances, because it is hard to prove the two groups are comparable.[2]
Regulators have since gone further. The FDA has a dedicated guidance on externally controlled trials.[1] The EMA addresses them in its reflection paper on single-arm trials,[3] and is currently developing a reflection paper specifically on the use of external controls.[4] The MHRA also has a draft guideline on real-world-data external controls.[5] All three accept the design. All three treat it as a source of bias to be engineered out.
HTA bodies also use external controls, because they must judge a medicine against relevant comparators. When no head-to-head trial exists, an external comparison can be the only way to answer the question they must answer. NICE’s real-world evidence framework covers single-arm trials with external controls directly.[6] And under the EU Joint Clinical Assessment (JCA), one dossier can face comparator questions from many countries at once, most without a head-to-head trial behind them.[7] We covered the JCA and its 100-day clock in a previous insight.
Support is conditional. An external comparison is most convincing when the treatment effect is large, the endpoint is objective, and the natural course of the disease is well understood.[2,8]
There are two ways to build one
An external control can be built from patient-level data or from summary data.
Patient-level data means a record for each individual patient: their characteristics, their treatment, their outcome. Summary data means the published averages for a comparator group: the median survival, the response rate, the average age.
The two lead to different methods, different assumptions, and different levels of scrutiny. The choice is not a formality. It sets the ceiling on what your comparison can prove.
Patient-level data: adjust, then compare
Patient-level data usually comes from real-world sources. Disease registries, electronic health records, and insurance claims. Sometimes it comes from the control arm of a previous trial, though such trials are often sponsored by a competitor, so patient-level data are rarely released.
With individual records for both groups, you can adjust for the differences between them. You measure the characteristics that affect the outcome, then use them to balance the two groups. Common methods are propensity score matching, weighting, and doubly robust estimation, which combines two models so that only one has to be correct.[9,10]
The goal is easy to state and hard to reach. You want any difference in outcome to be caused by the treatment, not by a difference between the patients. Get that wrong, and the difference reflects the confounders, not the drug.[9]
This is why regulators and HTA bodies often prefer fit-for-purpose patient-level data. It allows adjustment using individual patient characteristics and may produce more reliable estimates than aggregate data, although it does not remove unmeasured confounding.[1,11] The FDA’s externally controlled trials guidance does not cover summary-level comparisons at all.[1]
Summary data: match to the average
Often you cannot get patient-level data for the comparator. Published trials report averages, not individual records. This is the normal situation, not the exception.
For time-to-event outcomes, pseudo patient-level data can be reconstructed from a published Kaplan-Meier curve.[12] But it recovers only the survival data, not the baseline characteristics or effect modifiers needed to adjust the comparison.
The usual setup is asymmetric: you hold patient-level data for your own trial, but only summary data for the comparator. The established methods for this are matching-adjusted indirect comparison and simulated treatment comparison. You reweight or model your own patient-level data so your population matches the average characteristics reported for the comparator. We explained how these work, and where they break, in our insight on indirect treatment comparisons.
These methods are legitimate and widely used, especially for HTA and payer submissions, where published comparator data is often all that exists. But in the single-arm setting they carry a heavy assumption. With no shared comparator to anchor the comparison, they require that every factor affecting the outcome has been measured and matched. Miss one, and the result is biased. This is why summary-data comparisons attract more scrutiny, and why they are usually presented as one part of the evidence rather than the whole of it.[7,13]
Summary against summary: anchored or naive
Sometimes you have only summary data on both sides. What you can do then depends on whether a common comparator links the trials.
If a common comparator links the trials, an anchored comparison such as a Bucher indirect comparison or network meta-analysis may be possible. The shared comparator preserves the randomisation within each trial, so you are not just placing two averages side by side. But it is only valid if the trials are sufficiently similar in the effect modifiers that matter.
Without a common comparator, as in a single-arm trial, you are left with a naive comparison. You take your trial’s result and set it directly against the comparator’s published result, making no adjustment for the differences between the two groups of patients. This is generally discouraged, because any difference might be down to the treatment, or might simply be that the patients differ.[11]
It is not always wrong. When the effect is enormous, population differences cannot plausibly explain it. The FOCUS trial is a clean example. A liver-directed melphalan therapy was tested single-arm in metastatic uveal melanoma against a pre-specified benchmark: a pooled response rate of 5.5% from historical therapies, with a 95% confidence interval of 3.6 to 8.3%. The trial came in at 36.3%, with a 95% confidence interval of 26.4 to 47.0%. The lower bound of its confidence interval sat well above the upper bound of the benchmark. The signal was too loud for confounding to explain it, and the FDA approved the therapy.[14]
The lesson is narrow. A naive comparison can clear a very high bar, but only when the effect is overwhelming. That is the exception, not the rule.
Patient-level data is only as strong as its source
So fit-for-purpose patient-level data is usually the stronger choice. But being able to build a patient-level control is not the same as being able to trust it. Real-world data is collected to treat patients, not to run trials, and it often fails the standards a credible comparison needs.
A few problems come up again and again.
Time zero. In a trial, follow-up starts at a defined moment. In real-world data, the equivalent start date can be assigned in several ways. A poor choice creates immortal time: a stretch where the outcome could not have occurred in one group, which makes the treatment look better than it is.[1]
Assessment time. Progression-free survival depends on when you look. Trials scan on a fixed schedule. Routine care scans when a clinician decides to. Because progression is only recorded at the next scan, the two groups measure the same event on different clocks. This is assessment time bias, and it can move the result on its own.[1]
Endpoint definition. A trial endpoint like RECIST progression is rarely recorded in an electronic health record. You have to rebuild it from proxies, and a loose definition misclassifies patients.[15]
Intercurrent events. Later treatments, dose changes and switches are recorded carefully in a trial and patchily in routine care. If you cannot see them, you cannot account for them.[1]
Unmeasured confounding. This is the one that cannot be fixed. Adjustment only works for the factors you have measured. If an important prognostic factor is missing, no method recovers it. Even doubly robust estimation “does not obviate the need to measure all confounders.”[10]
Whichever choices you make here, stress-test them. A sensitivity analysis reruns the comparison under different reasonable assumptions, such as a different index date or adjustment method, to show whether the result holds or depends on the choices made.[1]
None of this rules out patient-level data. It means the data has to be assessed honestly before you commit to it. A patient-level control built on data that cannot support these checks is weaker than a well-conducted summary-data comparison.
When summary data is the right call
Summary data is not a consolation prize. It is the right choice in common situations.
When the comparator exists only as published aggregate data, a well-conducted matching-adjusted indirect comparison or simulated treatment comparison is more defensible than a patient-level control stitched together from data that cannot support it. When speed matters, summary data is often already available. And for many HTA and payer questions, the comparator evidence is aggregate by nature.
The point is to choose deliberately. Match the method to the data you can actually stand behind, not to a hierarchy on paper.
Scrutinised, and still used
Two NICE appraisals show how this plays out in practice.
In its 2015 appraisal of axitinib for advanced kidney cancer, the company compared axitinib against best supportive care through indirect comparisons, including a simulated treatment comparison and anchored comparisons via a shared drug. The committee judged the comparisons highly uncertain and found that some outputs from the simulated treatment comparison lacked clinical plausibility. Even so, the indirect comparison still contributed to the evidence review, and axitinib was recommended.[16]
In its 2023 appraisal of dabrafenib plus trametinib for BRAF-positive lung cancer, the pivotal study was single-arm, and the company compared it against a competitor trial using an unanchored matching-adjusted indirect comparison. The committee listed the familiar limitations: a reduced effective sample size, and the unanchored assumption that every effect modifier and prognostic factor had been captured. Even so, it concluded that “despite the limitations of the MAIC, it was an acceptable source of comparator clinical efficacy evidence and was the committee’s preferred source for decision making.”[17] Read that again. The indirect comparison was the committee’s preferred source of evidence.
Neither comparison was treated as clean proof. Both were heavily scrutinised, qualified, and still used. That is the realistic bar. Not a flawless comparison, but one built and reported well enough that an assessor can rely on it.
Take home
-
When your trial has no suitable control, you can build one from outside it. The first choice is the data: patient-level or summary.
-
Fit-for-purpose patient-level data allows adjustment at the individual level, so regulators and HTA bodies generally regard it as the stronger basis for comparison.
-
But real-world patient-level data often fails the tests that make a comparison credible: time zero, assessment time, endpoint definition, intercurrent events, and unmeasured confounding.
-
Summary-data methods like matching-adjusted indirect comparison and simulated treatment comparison are often necessary, especially for HTA and payers, when patient-level comparator data does not exist.
-
A naive, unadjusted comparison is the blunt instrument. It convinces only when the effect is overwhelming.
-
The right answer is chosen case by case, from what the data can actually support.
This is the work we do at Evidax. We run a full feasibility assessment of both data sources before committing to either. Then we build the comparison with the method the data can support, using purpose-built workflows, and stress-test it under differing assumptions before an assessor does. Sometimes that means a patient-level external control. Sometimes it means a summary-data comparison. Often it means both, each answering the question it is best placed to answer.
If your pivotal evidence does not have a comparator, that evidence gap is worth identifying early. Everything downstream depends on it.
References
-
US Food and Drug Administration (CDER, CBER, Oncology Center of Excellence). Considerations for the Design and Conduct of Externally Controlled Trials for Drug and Biological Products. Draft Guidance for Industry. February 2023. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/considerations-design-and-conduct-externally-controlled-trials-drug-and-biological-products
-
International Council for Harmonisation. ICH E10: Choice of Control Group and Related Issues in Clinical Trials. 2000. https://www.ich.org/page/efficacy-guidelines
-
European Medicines Agency, Committee for Medicinal Products for Human Use. Reflection Paper on Establishing Efficacy Based on Single-Arm Trials Submitted as Pivotal Evidence in a Marketing Authorisation. Draft, EMA/CHMP/564424/2021. 21 April 2023. https://www.ema.europa.eu/en/documents/scientific-guideline/draft-reflection-paper-establishing-efficacy-based-single-arm-trials-submitted-pivotal-evidence-marketing-authorisation_en.pdf
-
European Medicines Agency, Committee for Medicinal Products for Human Use. Draft Concept Paper on the Development of a Reflection Paper on the Use of External Controls for Evidence Generation in Regulatory Decision-Making. EMA/CHMP/225255/2025. 25 July 2025. https://www.ema.europa.eu/en/documents/scientific-guideline/draft-concept-paper-development-reflection-paper-use-external-controls-evidence-generation-regulatory-decision-making_en.pdf
-
Medicines and Healthcare products Regulatory Agency. MHRA Draft Guideline on the Use of External Control Arms Based on Real-World Data to Support Regulatory Decisions. May 2025. https://www.gov.uk/government/consultations/consultation-on-mhra-draft-guideline-on-real-world-data-external-control-arms
-
National Institute for Health and Care Excellence. NICE Real-World Evidence Framework. Corporate document ECD9. June 2022. https://www.nice.org.uk/corporate/ecd9
-
van Beekhuizen S, Che M, Monfort L, et al. Indirect treatment comparisons in EUnetHTA relative effectiveness assessments: learnings and recommendations for the implementation of EU joint clinical assessments. PharmacoEconomics Open. 2025;9(4):597-609. https://doi.org/10.1007/s41669-025-00575-1
-
Nuño MM, Pugh SL, Ji L, et al. On the use of external controls in clinical trials. Journal of the National Cancer Institute Monographs. 2025;2025(68):30-34. https://doi.org/10.1093/jncimonographs/lgae046
-
Lambert J, Lengliné E, Porcher R, et al. Enriching single-arm clinical trials with external controls: possibilities and pitfalls. Blood Advances. 2023;7(19):5680-5690. https://doi.org/10.1182/bloodadvances.2022009167
-
Funk MJ, Westreich D, Wiesen C, et al. Doubly robust estimation of causal effects. American Journal of Epidemiology. 2011;173(7):761-767. https://doi.org/10.1093/aje/kwq439
-
Hogervorst MA, Soman KV, Gardarsdottir H, et al. Analytical methods for comparing uncontrolled trials with external controls from real-world data: a systematic literature review and comparison with European regulatory and health technology assessment practice. Value in Health. 2025;28(1):161-174. https://doi.org/10.1016/j.jval.2024.08.002
-
Guyot P, Ades AE, Ouwens MJNM, Welton NJ. Enhanced secondary analysis of survival data: reconstructing the data from published Kaplan-Meier survival curves. BMC Medical Research Methodology. 2012;12:9. https://doi.org/10.1186/1471-2288-12-9
-
Aballéa S, Toumi M, Wojciechowski P, et al. Between rigor and relevance: why the EU HTA guidelines on indirect comparisons miss the mark. Journal of Market Access & Health Policy. 2026;14(2):30. https://doi.org/10.3390/jmahp14020030
-
Zager JS, Orloff M, Ferrucci PF, et al. Efficacy and safety of the melphalan/Hepatic Delivery System in patients with unresectable metastatic uveal melanoma: results from an open-label, single-arm, multicenter phase 3 study. Annals of Surgical Oncology. 2024;31(8):5340-5351. https://doi.org/10.1245/s10434-024-15293-x
-
US Food and Drug Administration (CDER, CBER, Oncology Center of Excellence). Real-World Data: Assessing Electronic Health Records and Medical Claims Data to Support Regulatory Decision-Making for Drug and Biological Products. Guidance for Industry. 2024. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/real-world-data-assessing-electronic-health-records-and-medical-claims-data-support-regulatory
-
National Institute for Health and Care Excellence. Axitinib for Treating Advanced Renal Cell Carcinoma After Failure of Prior Systemic Treatment. Technology appraisal guidance TA333. 25 February 2015. https://www.nice.org.uk/guidance/ta333
-
National Institute for Health and Care Excellence. Dabrafenib Plus Trametinib for Treating BRAF V600 Mutation-Positive Advanced Non-Small-Cell Lung Cancer. Technology appraisal guidance TA898. 14 June 2023. https://www.nice.org.uk/guidance/ta898