Randomized non-comparative trials: an oxymoron?

The more I think about RNCTs the more I conclude that statisticians who endorse these need to have their ethics examined.

2 Likes

Maybe this is how events play out (?):

  • Researchers really don’t want to risk missing an efficacy signal for a new therapy in a rare disease;

  • They know that the most reliable way to assess whether one therapy is superior to another is by conducting an RCT i.e., randomizing some patients to the new drug and other patients to the standard of care, then comparing their clinical trajectories;

  • But researchers also know that a small RCT will be less capable of detecting a new therapy’s superiority over standard of care than a large RCT. Since few people will remain alive at each time point after randomization (due to the small number of trial participants), between-arm differences in survival rates at any given time point will be hard to interpret;

  • Therefore, they worry that the small size of their trial could cause them to miss the signal for superiority that they are seeking from the new drug and which might actually be present;

  • To address this fear, they propose a backup plan- an alternate way to search for a potential superiority signal for the new therapy, should the within-trial between-arm comparison fail to detect one;

  • They decide, when designing their study, that if the between-arm comparison within their small RCT ends up being non-informative, they will instead look for an efficacy signal by comparing the clinical trajectory of patients treated with the new drug in their current trial with the clinical trajectory of patients treated with standard of care therapy in historical trials. But this type of comparison, with historical controls, is no longer consistent with an “RCT” design. Rather, it’s a type of observational cohort study (?), with the progress of the two arms potentially being observed in different time periods and using patients who are members of different convenience samples;

  • Researchers settled on the term “Randomized Non-Comparative Trial” because they couldn’t justifiably call their study an “RCT” if they weren’t, realistically, planning to focus their statistical analysis on the randomized between-arm comparison within their single convenience sample. They would plan to “eyeball” the within-trial between-arm comparison behind the scenes, but then ignore it in their subsequent publication if they didn’t see any obvious separation of the trial arms. By randomizing, they would preserve their ability to 1) be happily surprised by unexpected detection of a superiority signal for the new drug; or 2) contribute their trial’s data to a future meta-analysis of RCTs involving the same treatment;

  • The problem: If they don’t end up reporting the within-trial between-arm comparison (because they don’t see a superiority signal when they eyeball their results) and instead end up focusing their analysis on the comparison with a historical control, any failure to admit, explicitly, that their main treatment comparison has become observational in nature (rather than randomized) will mean that readers might not recognize that they haven’t met the standards required for a valid observational treatment comparison. Readers who don’t recognize that they are actually reading about an observational treatment comparison (because the word “randomized” is used in the study’s title) might not register that the authors haven’t tried to identify important between-trial, between-arm untreated prognostic distribution differences. These between-trial differences need to be addressed because the patients in the arms being compared will come from two different convenience samples (the current trial’s convenience sample and the historical study’s convenience sample)- not from a single convenience sample, as would have been the case if the authors had compared the trial’s two randomized arms;

RNCT authors need to take the precautions required for a valid observational treatment comparison- the type of comparison that will usually end up being the main focus of an "R"NCT.

1 Like

Yup. But if RNCTs are as absurd as we think they are then those who dare to compare should gain a meaningful advantage. We have done our part raising awareness. Leading by example is the strategy we are henceforth predominantly adopting to counter RNCTs.

2 Likes

@ESMD I think your layout of the logic that is likely being used by proponents of RNCTs is excellent. At the core of their misunderstandings is mean squared error of treatment effect estimates. Even a small bias from historical controls will render their perceived variance reduction moot, as exemplified here - EHRs and RCTs: Outcome Prediction vs. Optimal Treatment Selection – Statistical Thinking.

1 Like

Compare this with the esotericism in Project Optimus, as discussed in this PubPeer post,

My conjecture now is that FDA OCE ultimately came to view ‘dose optimization’ as fundamentally an esoteric practice—one wholly driven by idiographic considerations, and utterly inaccessible to a nomothetic perspective. Thus, ultimately, Project Optimus degenerated into a syncretistic doctrine that embraced and tolerated contradictions (Eco 1995, Popper 1940).

To reveal the origins of this insight, here’s a bit of my interaction with an FDA OCE Reviewer 2 of my 2023 PSP paper, How large must a dose-sptimization trial be?:

Apropos of Erin’s “eyeball” specifically, compare the concept of gestalt randomization I helpfully offer a bit later on in that exchange:

1 Like

Again apropos of Erin’s “eyeball”, it occurs to me that the isotonic transformation applied in the MERIT design I’ve criticized in another thread could be seen as a way to indulge the illusion so vital to RNCT’s, that this pesky randomness statisticians are always going on about can somehow be swept away without their assistance:

In any case, the upshot of these last 2 posts is that, if we are looking for empirical evidence on the mental substrate for RNCT’s, then Project Optimus provides plenty of material. Indeed, FDA OCE developed and advocated its “randomized, parallel [yet non-comparative] dose–response trials” out in the open (on YouTube, in a ‘white paper’, in JCO and the NEJM), much as Project 2025 did for the current US Administration.

2 Likes

A 2019 publication seems to speak to the problem at the heart of RNCTs and focuses on how variance is affected not only by non-concurrence in time of compared groups, but also non-concurrence in convenience sample.1 The most crucial point (bold and italic fonts are mine):

For example, assuming 𝜎*2=1, and a scenario where 10 historical control arms including 100 patients each with a between-study variance* 𝜏*2=0.02, running a randomised clinical trial with 130 patients, of which half of them would be dedicated to the concurrent control arm, would be sufficient in order to achieve the same variance.*

The authors use the previously-published TARGET trial to illustrate their key messages. The TARGET trial compared outcomes among patients treated with either lumiracoxib or ibuprofen. However, this was effectively two trials in one, as subjects were first recruited within each centre and only then were randomly allocated to receive either lumiracoxib or ibuprofen. The authors seem to highlight that this design (i.e., TWO separate randomization procedures, one for each of TWO convenience samples) must be analyzed in a different way than a trial in which randomization occurred at a point “upstream” from the centres participating in the overall trial (i.e., ONE randomization procedure for a SINGLE recruited convenience sample).

Crucially, in the TARGET trial, some outcomes among patients randomized to exactly the same therapy during exactly the same time period, but from within two different recruited convenience samples (i.e., the two substudies), were different. It followed that any subsequent analysis that attempted to “pool” the two substudies and treat them as though they were part of a single larger study (involving a single convenience sample) would need to acknowledge the contribution of between-convenience-sample/“between-study” variance in the final analysis. This variance would be greater than the variance that would be generated from a study that had randomized the same total number of randomized subjects, but from within a single convenience sample.

The authors also explore the implications of their findings for the standard errors presented in the analysis of observational studies (bold type is mine).

This raises an interesting issue for observational studies themselves. These are conventionally analysed as if they were less-than-perfect parallel group trials with adjustment for ‘confounders’ being the solution for dealing with the imperfection. That is to say, compared to a clinical trial, a penalty is paid for the loss of orthogonality that confounding brings,23 but otherwise the variance term is treated as if a parallel group trial were appropriate. Many cohort studies are analysed exactly like this. In other words, the problems we have described raise the following possibility, namely that confounding is not the only problem with observational studies. A further problem is the implicit assumption of conditional independence of observations given adjustment. Of course, from one point of view, this is simply another form of bias, and a recent important paper by XiaoLi Meng24 draws attention to the dangers of naively assuming that accuracy can be determined solely by using traditional standard errors inversely proportional to the square root of the number of data points.

The authors conclude (bold type is again mine):

There are several problems in moving from randomised concurrently controlled trials to using historical data. One that has received increasing attention is that variation between studies is an important factor that must be accounted for. This means that not only is it inappropriate in judging the amount of information available to concentrate on the number of patients available to use as controls, it is naïve to suppose that some constant discounting will deal with the fact that the information they provide is not as good as that provided by concurrent controls. Instead, inter-study variation provides an upper bound on the amount of information that historical controls can provide.

Appreciating that the closest randomised analogy is not a parallel group trial but a trial in which treatments are allocated to clusters, immediately focuses on the problem of using historical data. The randomisation analogue involves k control observations (the historical trials) versus 1 experimental observation (the current trial). Not only does the relevant variance term not go to zero as the number of patients increases, because of the behaviour of the second term in (3), it does not do so as the number of historical trials increases. Note also that as discussed above, the model is one that would apply, if studies could be regarded as exchangeable. Since no randomisation is involved, this is a strong assumption.

As far as I can tell (?), authors of “RNCTs” would be wise to take the following message from this paper: “You are mistaken if you believe that a comparison with historical control might represent an acceptable “backup plan” if your within-sample between-arm comparison ends up being uninformative. Any methodologically-sound attempt at such a comparison with historical control would need to acknowledge such a massive increase in variance (reflecting not just variance inflation due to time trend considerations, but also between-sample considerations) that any benefit that you might hope to have achieved from studying a larger number of subjects would be rendered inconsequential.”

1 Collignon O, Schritz A, Senn SJ, Spezia R. Clustered allocation as a way of understanding historical controls: Components of variation and regulatory considerations. Stat Methods Med Res. 2020 Jul;29(7):1960-1971. doi: 10.1177/0962280219880213. Epub 2019 Oct 10. PMID: 31599194.

3 Likes

Exactly, hence why the “original sin” in RCTs were treatment group-specific inferences. From there, it takes only one more jump to fully degenerate into RNCTs by disallowing comparative treatment inferences and focusing only on treatment group-specific inferences .

3 Likes

Yes, the practice of treatment group-specific inference potentially reflects two distinct misunderstandings:

  1. Failure to appreciate that the absence of random sampling in clinical RCTs will severely hamper the generalizability of any within-group descriptive statistics beyond the convenience sample at hand; and
  2. Failure to appreciate that within-group outcomes don’t, with certain exceptions (e.g., oncology, where patients’ “counterfactual” disease trajectory, in the absence of exposure to an intrinsically efficacious therapy, is known to be one of steady worsening), permit inferences related to the causal effect of a therapy for individual patients. Individual effects are causally non-identifiable in a parallel group RCT if the disease has a fluctuating natural history and the trial lacks crossover periods.

In short, neither of the ostensible aims of a within-group analysis (descriptive OR causal) can realistically be achieved for most clinical RCTs.

3 Likes

It may be the case that adding a bias parameter in joint Bayesian models of randomized + observational data handles this automatically. This is instead of using the usual discounted prior approach that takes the posterior distribution of a parameter from historical data at face value. See Incorporating Historical Control Data Into an RCT – Statistical Thinking

2 Likes