RMST: survival benefit HTA committees can use
A hazard ratio of 0.70 is familiar territory in a clinical trial. It compares the rate at which an event occurs in the treatment group with the rate in the control group.
It is often read as a 30% reduction in risk. That is not quite what it means.
The hazard ratio compares only the patients who are still alive or free from the event at each moment. That group changes throughout follow-up. Treatment itself can change who remains in it.[1,2]
There is another way to describe the survival benefit. Ask how many additional months patients gained over a clinically relevant period. Then ask how many more were alive or progression-free at the time that matters.
Those answers come from restricted mean survival time and fixed-time survival rates. They put survival benefit on an absolute scale that committees, clinicians and patients can use.
A hazard ratio compares the patients still at risk
The hazard is the event rate at a particular moment among patients who have reached that moment without the event.
For overall survival, that means the patients still alive. For progression-free survival, it means those who remain alive without progression.
Randomisation balances the treatment groups at the start of a trial. It does not guarantee that the patients still at risk remain comparable later.[1,2]
Some patients have a worse prognosis than others. They are more likely to experience the event early. If treatment prevents or delays those events, it changes the mix of patients who remain at risk in that arm.
The hazard ratio later in follow-up is therefore comparing two selected groups. They are no longer necessarily the same kinds of patients who were randomised at the start.
A late change in the ratio may be selection, not biology
This creates what has been called the built-in selection of the hazard ratio.[1,2]
The Women’s Health Initiative provides a clear example. The hazard ratio for coronary heart disease among women assigned hormone therapy was harmful early in follow-up, but fell below one after year five. Read literally, the late result appeared protective.[2]
One explanation is that women who were particularly susceptible to harm had already experienced events in the treatment group. The women who remained were a more selected, less susceptible group. The late hazard ratio could fall even if hormone therapy protected nobody at that point.[2]
The same issue can make a late hazard ratio rise. A changing ratio may represent a changing treatment effect. But it may also reflect who remains at risk.
A late ratio should not be read as evidence that a treatment has started working, stopped working or reversed without looking at the survival curves and the patients still contributing to them.
Hazard ratios move when the population changes
There is a second problem. Hazard ratios are non-collapsible.[1,3,4]
That means the hazard ratio for a whole population is not an average of the hazard ratios within its patient groups.
Imagine two prognostic groups. The treatment has the same hazard ratio within each. Combine them and the overall hazard ratio can still be different. It can change over time and can even fall outside the two subgroup results.[1]
This is not necessarily confounding. A variable that predicts survival can move the population hazard ratio even when it does not change anyone’s response to treatment.[1,3]
The problem matters in indirect comparisons. Different trials recruit different patients. Their prognosis, follow-up and censoring may differ. Two studies can therefore report different population hazard ratios even if the underlying treatment effect within comparable patients is the same.[3,4]
A relative effect that moves with the population is a difficult foundation for a population decision.
RMST turns the survival curve into months
Restricted mean survival time, or RMST, is the area under a survival curve up to a chosen time.[5-7]
That definition is correct and not especially helpful on its own.
In plain English, RMST is the average time patients are alive or free from the event during a defined period.
Suppose a trial is assessed over 24 months. The RMST is the average number of those 24 months that patients in each treatment group remain alive. Subtract one treatment’s RMST from the other’s and the result is the average survival time gained during that period.
If the RMST difference is 2.4 months at 24 months, patients receiving one treatment accumulated an average of 2.4 additional months of survival during the first 24 months.
This does not mean every patient lives 2.4 months longer. It does not mean lifetime survival increases by 2.4 months. It is the average difference accumulated during the first 24 months.
The unit is time. Months gained is a result that can be discussed directly.
RMST can be averaged across patient groups
RMST difference has another important property. It is collapsible.[5,6]
The reason is easier to understand from the survival curves.
At every point in time, the vertical distance between two curves is the absolute difference in survival. That difference can be averaged properly across patient groups.
RMST adds those survival differences across the full period. Adding absolute differences preserves that property. The RMST difference for the population is therefore an appropriate weighted average of the differences within its patient groups.[5]
If two equally sized groups gain an average of one month and three months, the population gain is two months. Hazard ratios do not combine in this straightforward way.
Collapsibility does not remove confounding. It does not correct a poorly designed comparison. It means that the effect scale behaves sensibly when results are averaged over a population.
Clinicians want survival at a time that matters
RMST describes the survival benefit accumulated over a period. Sometimes the decision requires an answer at a particular time.
A 2015 NHS England evidence review of chemosaturation for liver metastases from ocular melanoma made this explicit. Its prespecified outcomes marked “critical for decision-making” included progression-free survival and overall survival at one and five years.[8]
Those questions are direct. Of 100 patients starting treatment, how many are alive at one year? How many remain free from progression? What about at five years?
The answer is a fixed-time survival rate. If 62% of patients are alive at two years with one treatment and 50% with another, the absolute survival difference is 12 percentage points.
The CHEMOSAT review found that the available evidence was too limited to provide firm answers. But the outcomes it selected show why fixed-time survival matters to clinicians and commissioners.[8]
RMST uses the journey. Fixed-time survival marks the destination
RMST and fixed-time survival are complementary.
RMST uses every part of the survival curve up to the chosen time. It captures benefit accumulated before that point.[7]
A fixed-time rate reads the survival curve once. It shows where the treatments stand at a clinically important time, but not when the difference appeared.[7]
Two treatments could have the same survival rate at 24 months after taking very different paths to get there. One might offer an early benefit that later disappears. The other might separate only after a delayed response.
The full survival curves show that pattern. RMST summarises the accumulated difference. Fixed-time rates answer the question at selected moments.
Committees should see all three, with confidence intervals and the numbers of patients still at risk.
Choose the time before you see the answer
An RMST difference at 12 months and an RMST difference at 36 months are different treatment effects.
The same is true of survival at one year and survival at five years. The time horizon is part of the question being answered.[5,7,9]
There is no universal best horizon. The papers in this field give a consistent set of principles.
The time should matter clinically. It should apply equally to every treatment. It should be supported by reliable follow-up. It should preferably be chosen before the results are inspected.[5,7,9]
Late time points need particular care. Fewer patients remain at risk, so the survival estimate becomes less precise. A single RMST value does not show how uncertainty changes across the curve. That is why it should be reported with confidence intervals and numbers at risk.[7]
Analyses chosen after looking at the curves may also be seen as selecting a favourable answer. Monnickendam and colleagues found that RMST analyses in HTA were often post hoc, which may limit the weight assessors give them.[7]
If several horizons matter, report a prespecified profile. For example, RMST and survival differences at 12, 24 and 36 months. That shows how the benefit develops without turning every time point into another chance to find a positive result.
Define the question before choosing the method
The exact treatment question is called the estimand.
For an RMST comparison, it might be:
What is the average difference in overall survival accumulated over 24 months if patients in the target population receive treatment A rather than treatment B?
For a fixed-time comparison, it might be:
What is the difference in the proportion of patients alive at 24 months?
Those questions specify the population, treatments, outcome, horizon and summary measure. They should also specify how treatment switching, discontinuation and subsequent treatment are handled.
Remiro-Azócar and colleagues show why this comes first. Population adjustment methods can target different effects. MAIC usually targets a marginal effect for a population. Conventional regression adjustment often produces a conditional effect for patients with specified characteristics. With a non-collapsible measure such as the hazard ratio, those effects are not interchangeable.[4]
Choose the question first. Then choose a method that estimates that question.
Indirect comparisons need one target population
Raw survival curves from different studies should not simply be placed beside each other.
The studies may contain patients with different prognosis, prior treatment, eligibility criteria or follow-up. RMST and fixed-time survival do not correct those differences. They are summaries calculated after the treatment evidence has been aligned.
The aim is to estimate one survival curve for every treatment in the same decision target population. The comparisons then follow naturally.
-
The RMST difference is the area between the target-population curves up to the common horizon.
-
The fixed-time survival difference is the vertical distance between the curves at a chosen time.
Phillippo and colleagues identify population-average survival curves as the quantities needed for decision making and economic models. MAIC and STC can estimate survival curves for the population represented by the aggregate-data study. ML-NMR can potentially estimate curves for all treatments in another target population when the necessary covariate and baseline survival information are available.[3]
This changes the sequence of the analysis.
Do not start by asking which statistic can be extracted from each paper. Start by defining the decision population and the survival question. Then reconstruct, adjust or standardise the treatment evidence to that population. Calculate RMST and fixed-time differences last.
Published Kaplan-Meier curves can help. A validation across 39 randomised trials found that RMST calculated from reconstructed patient-level survival data closely matched RMST calculated from the original trial data. Accuracy was poorer when numbers at risk were unavailable.[10]
Reconstruction recovers the survival pattern. It does not recover the missing patient characteristics needed for population adjustment.
RMST remains interpretable when hazards are not proportional
A constant hazard ratio assumes that the relative event rate remains proportional over time.
That assumption is often difficult to defend when treatments have delayed effects, when survival curves cross, or when a small group of patients experiences durable long-term benefit.
RMST and fixed-time survival differences do not require proportional hazards. They remain interpretable because they are calculated from the survival curves themselves.[1,5,7]
This is particularly relevant in immuno-oncology. Liang and colleagues reviewed 25 randomised trials of immune checkpoint inhibitors involving 12,870 patients. RMST-based measures and reported hazard ratios agreed on the direction of treatment effect in every trial, but disagreed on statistical significance in two. The hazard ratio produced a larger effect estimate in all 25 trials.[11]
That does not mean RMST will always be smaller, more conservative or more powerful. It means that RMST continues to answer a clear question when one constant hazard ratio no longer describes the survival pattern well.
NICE’s DSU separates observed months from modelled months
NICE-commissioned Decision Support Unit documents draw a clear line between observed RMST and lifetime survival.[12,13]
RMST can summarise survival within a defined period. But TSD 14 recommends restricted mean approaches for economic models only when data are almost entirely complete. When trial data are immature, lifetime survival must be extrapolated beyond follow-up.[12]
Those later months are modelled, not observed. They depend on the selected survival model, external evidence and assumptions about how long the treatment effect lasts.
TSD 21 found that many survival models produced relatively little bias when estimating RMST to the end of trial follow-up. The same models could produce substantial bias when extrapolated to lifetime mean survival.[13]
An RMST calculated from observed data is therefore not automatically a lifetime survival gain. An extrapolated RMST is not model-free.
A better survival comparison is a package, not one replacement number
The answer is not to ban hazard ratios.
They remain useful relative summaries when their assumptions and interpretation are clear. The problem is asking one ratio to carry the whole survival story.
A decision-ready survival comparison should show:
-
The survival curves with confidence intervals and numbers at risk
-
RMST for each treatment and the difference between them at a prespecified horizon
-
Survival for each treatment and the absolute difference at clinically relevant times
-
The hazard ratio as a complementary relative measure
-
The target population, estimand and handling of treatment changes
-
Sensitivity to the horizon, population adjustment and any extrapolation assumptions
For indirect comparisons, producing that package requires more than reading values from published papers. The evidence must be identified, the survival data may need to be reconstructed, the target population must be defined, and the curves must be estimated on a common basis before the summaries are calculated.
That is where Evidax fits. We design the survival question first, assess what the available patient-level and published data can support, and build the comparison around the target population the decision is actually for.
Take home
-
A hazard ratio compares event rates among the patients who remain at risk. Treatment changes who those patients are.
-
A late change in the hazard ratio may reflect selection rather than a treatment starting to work, wearing off or reversing.
-
Hazard ratios are non-collapsible. RMST differences and absolute survival differences can be averaged properly across patient groups.
-
RMST reports the average survival time accumulated over a chosen period. It does not tell every patient how much longer they will live or automatically estimate lifetime survival.
-
Fixed-time survival answers how many patients are alive or event-free at a clinically important time. RMST describes the journey to that point.
-
The horizon is part of the estimand. Choose it for clinical relevance, apply it equally across treatments, and support it with reliable follow-up.
-
In an indirect comparison, estimate every treatment’s survival curve in the same target population before calculating RMST or fixed-time differences.
-
Show the curves, months gained, fixed-time rates, uncertainty and assumptions together. No single number should carry the whole survival story.
References
-
Dumas E, Stensrud MJ. How hazard ratios can mislead and why it matters in practice. European Journal of Epidemiology. 2025;40:603-609. https://doi.org/10.1007/s10654-025-01250-9
-
Hernán MA. The hazards of hazard ratios. Epidemiology. 2010;21(1):13-15. https://doi.org/10.1097/EDE.0b013e3181c1ea43
-
Phillippo DM, Remiro-Azócar A, Heath A, et al. Effect modification and non-collapsibility together may lead to conflicting treatment decisions: a review of marginal and conditional estimands and recommendations for decision-making. Research Synthesis Methods. 2025;16:323-349. https://doi.org/10.1017/rsm.2025.2
-
Remiro-Azócar A, Heath A, Baio G. Conflating marginal and conditional treatment effects: comments on “Assessing the performance of population adjustment methods for anchored indirect comparisons: a simulation study”. Statistics in Medicine. 2021;40:2753-2758. https://doi.org/10.1002/sim.8857
-
Ni A, Lin Z, Lu B. Stratified restricted mean survival time model for marginal causal effect in observational survival data. Annals of Epidemiology. 2021;64:149-154. https://doi.org/10.1016/j.annepidem.2021.09.016
-
Lin Z, Ni A, Lu B. Matched design for marginal causal effect on restricted mean survival time in observational studies. arXiv:2205.02241. 2022. https://arxiv.org/abs/2205.02241
-
Monnickendam G, Zhu M, McKendrick J, Su Y. Measuring survival benefit in health technology assessment in the presence of nonproportional hazards. Value in Health. 2019;22(4):431-438. https://doi.org/10.1016/j.jval.2019.01.005
-
Solutions for Public Health. Evidence Summary Report: Chemosaturation for Liver Metastases from Ocular Melanoma. Prepared for NHS England. 10 December 2015.
-
Oakley JE, Ren S, Forsyth JE, et al. NICE DSU Technical Support Document 26: Expert Elicitation for Long-Term Survival Outcomes. Decision Support Unit. 2025. http://www.nicedsu.org.uk
-
Everest L, Blommaert S, Tu D, et al. Validating restricted mean survival time estimates from reconstructed Kaplan-Meier data against original trial individual patient data from trials conducted by the Canadian Cancer Trials Group. Value in Health. 2022;25(7):1157-1164. https://doi.org/10.1016/j.jval.2021.12.004
-
Liang F, Zhang S, Wang Q, Li W. Treatment effects measured by restricted mean survival time in trials of immune checkpoint inhibitors for cancer. Annals of Oncology. 2018;29(5):1320-1324. https://doi.org/10.1093/annonc/mdy075
-
Latimer N. NICE DSU Technical Support Document 14: Undertaking Survival Analysis for Economic Evaluations Alongside Clinical Trials: Extrapolation with Patient-Level Data. Decision Support Unit. 2011, last updated March 2013. http://www.nicedsu.org.uk
-
Rutherford MJ, Lambert PC, Sweeting MJ, et al. NICE DSU Technical Support Document 21: Flexible Methods for Survival Analysis. Decision Support Unit. 2020. http://www.nicedsu.org.uk