Randomized non-comparative trials: an oxymoron?

So people are clearly onto this ever-widening focus-on-within-arm-change scam:

https://www.tandfonline.com/doi/full/10.1080/09332480.2025.2510165#d1e121

An article I read the other day caused me to view this trend in a much more cynical light than I had up to this point. Is it possible that this lipstick-on-a-pig trend with non-comparative randomized trials is being driven not so much by methodologic naïveté, but rather by mercenary interests of companies that depend on private investment ?

I understand less than nothing about the financial world. But aren’t there people who would stand to benefit from methodologically nightmarish mid-phase efforts to turn chicken *!@# into chicken salad?

Maybe this angle is unfairly uncharitable. But why would the biotech industry be so different from the tech industry?

https://www.politico.com/newsletters/digital-future-daily/2023/12/01/5-questions-for-meredith-whittaker-00129677

“It’s not simply that one piece of technology is overhyped, it’s that hype is a necessary ingredient of the current business ecosystem of the tech industry. We should examine how often the financial incentive for hype is rewarded without any real social returns, without any meaningful progress in technology, without these tools and services and worlds ever actually manifesting. That’s key to understanding the growing chasm between the narrative of techno-optimists and the reality of our tech-encumbered world.”

3 Likes

It is never that simple. In some cases industry (either big pharma and/or biotech; each of which are vastly different in motivations and practices) will indeed push quality down, whereas in others it will maintain higher standards than academics. In oncology, for example, while RNCTs are indeed more frequently industry-sponsored, industry is more likely than academia to better represent survival plots, show covariate adjusted results, avoid the Table 1 fallacy, and provide more complete adverse event reporting.

3 Likes

This might be taking the discussion in a direction that deserves a thread of its own…

Part of the motivation behind “randomised non-comparative trials” may be the over-use and inappropriate use of single-arm trials. I have become pretty convinced that in oncology in particular, and maybe other areas of medicine too, there is a problem with what I’ve called “single arm thinking” which is the idea that it’s possible to evaluate treatment effectiveness using single arm trials.

Actually, that is probably not completely wrong, but to do so requires a lot of careful work in understanding/modelling the outcomes that would be expected under the control treatment, or constructing external control groups etc etc, not to mention being very realistic about the extra uncertainty involved in such methods.

That isn’t what almost all single-arm trials do though. Repeatedly I see single-arm trials that say they aim to “evaluate a therapy” or “estimate the response rate (or whatever) of a particular therapy,” with no acknowledgement that the results you get from a single arm are going to be VERY dependent on the patients that get recruited. Patient heterogeneity and selection bias are real issues! If there is a comparison, it is almost always with some assumed historical rate, with no guarantees that this applies to the patients in the trial.

Of course, single-arm trials have legitimate uses – but the fact that they can be useful in some circumstances seems to have led to serious mission creep, so they now get used in many situations where they are just not suitable. I get that they have attractive features – everyone gets the exciting new therapy and there is no need to explain randomisation. But that isn’t worth the cost of making the trial useless.

This was brought home to me by a recent experience with a randomised multi-armed oncology trial, which made it all the way through to being funded, without specifying what comparisons it was going to make – because apparently everyone was just thinking about each arm in isolation and assuming that the results of each single arm were meaningful in themselves.

I don’t know if I’m overstating this issue – maybe I’ve just had some bad experiences – but it almost feels like there is a collective conspiracy to delude ourselves.

Interested in others’ views!

6 Likes

Realise I’ve said some of this earlier but thought it was worth directly addressing the issue of single arm trials and their problems.

3 Likes

This was brought home to me by a recent experience with a randomised multi-armed oncology trial, which made it all the way through to being funded, without specifying what comparisons it was going to make – because apparently everyone was just thinking about each arm in isolation and assuming that the results of each single arm were meaningful in themselves.

Indeed. It is also related to the frequent practice of making treatment group-specific inferences by showing 95% CI for each treatment group in RCTs. This indirectly equates them with the far more reliable 95% CI for comparisons.

On a related note, just gave last week an ISBS webinar lecture to applied biostatisticians raising awareness of RNCTs (from 40:50 onwards). The beginning of the lecture emphasizes that single-arm trials can be helpful in certain situations leading, e.g., to the recent improvements in survival outcomes of patients with renal medullary carcinoma, a clinically and molecularly homogeneous highly aggressive malignancy. But in many other contexts we should randomize and then dare to compare.

3 Likes

We just published another letter at JCO (the flagship journal of the American Society of Clinical Oncology which has been publishing RNCTs almost every other week) pointing out another hidden RNCT example not evident in the title or abstract. The authors’ reply is here.

At this point we have more than done our due diligence highlighting this issue across medicine. It is time for others, including statistical and clinical peer reviewers and editors, to step up and contain this torrent.

2 Likes

This is the authors’ stated justification for using the RNCT design (bolding is mine):

“The randomized, noncomparative design was selected over performing two sequential, single-arm phase Il studies to distribute unknown prognostic characteristics and to reduce time trend bias related to the emergence of human papillomavirus-associated disease in HNC clinical trials…”

I don’t understand what their bolded phrases mean- do you?

Also- can you summarize the purpose of Phase II trials in oncology? Specifically, can you explain what function randomization serves in phase II in fields like oncology, where biological activity of a therapy can be discerned without implementing a trial design with more than one arm (since tumours don’t spontaneously resolve over time)?

Addendum- Uggh #2- Never mind Pavlos. I just came across a 2019 publication on exactly this topic: Grayling M et al; Review of Perspectives on the Use of Randomization in Phase II Oncology Trials. Clearly this is much too big a question to be tossed out so casually. What a dog’s breakfast (the topic, not the article)…Also managed to find the very first published RNCT- the authors were Frankenstein, V et al…

https://doi.org/10.1093/jnci/djz126

2 Likes

And is there anything in the authors’ response that would justify RNCT? I have my doubts.

1 Like

Not at all unfortunately…

Nothing that I could find. Hopefully they will think twice about this next time…

1 Like

This is the problem: in many single-arm trials you can’t confidently ascribe biological activity to a new treatment, because patients receive other treatments. Very often the question is about adding a new therapy to existing treatments, which requires randomisation (or some hard work to produce an appropriate external control group).

3 Likes

This “Multiple opportunities exist…” section seems particularly strong to me, with so many constructive [and well-referenced] alternative methods. Yet Ferris, Bauman, Wang & Zandburg tellingly ignore them in their reply.

This effort of yours reminds me of #4 in Timothy Snyder’s list:

4. Take responsibility for the face of the world. The symbols of today enable the reality of tomorrow. Notice the swastikas and other signs of hate. Do not look away, and do not get used to them. Remove them yourself and set an example for others to do so.

I recently read Mark Bray’s Antifa book, and was impressed by the story of one European antifascist activist who just kept plastering over nazi graffiti for months on end until the nazi simply gave up.

2 Likes

Just re-read that excellent paper. Two things jumped out at me:

  1. the number of spurious arguments against randomisation that have been put forward over the years;
  2. the huge number of randomised non-comparative trials: the review they quote found 76/193 randomised trials were “non-comparative.”
4 Likes

Yup. We even just published a practice-informing randomized comparative phase II trial (I don’t think we knew RNCTs existed when we designed it) with n=90 patients which is in the sample size range of all these RNCTs. There is no excuse to avoid prespecifying comparisons.

3 Likes

The more I think about this thread, the more I wonder if the whole problem could simply reflect a failure to understand the advantages of concurrent versus historical control (?) I suspect the issue might be a lot more complicated than this though, since I have trouble believing that biostatisticians wouldn’t understand the importance of concurrent control (?)…

I don’t know whether this very crude graph could help to illustrate the problem or not (I realize this isn’t exactly how survival curves look).

The meaning of the legend terms is as follows:

CS1 4y SC- survival curve for Convenience Sample #1, from 4 years ago, for subjects treated with “Standard of Care” therapy

CS2 4y SC- survival curve for Convenience Sample #2, from 4 years ago, for subjects treated with “Standard of Care” therapy

CS3 T SC- survival curve for Convenience Sample #3, from Today, for subjects treated with “Standard of Care” therapy

CS4 T SC- survival curve for Convenience Sample #4, from Today, for subjects treated with “Standard of Care” therapy

CS5 T ND- survival curve for Convenience Sample #5, from Today, for subjects treated with “New Drug” therapy

CS6 T SC- survival curve for Convenience Sample #6, from Today, for subjects treated with “Standard of Care” therapy

CS6 T ND- survival curve for Convenience Sample #6, from Today, for subjects treated with “New Drug” therapy

Key points conveyed by the graph:

  • If concurrent control is available (as in an “RNCT” design), it doesn’t make sense to compare the survival trajectory for patients taking a New Drug (brown line) with the survival trajectory of patients treated with Standard of Care therapy in a non-concurrent treatment arm (i.e., patients from studies that used other convenience samples, either historical or current day- Blue, Grey, Green, or Yellow lines);
  • Different convenience samples will generate different survival trajectories; this will be true for patients treated with both Standard of Care therapy AND those treated with a New Drug. This phenomenon occurs because the manner in which the available “covariate space” defined by a trial’s inclusion criteria will end up being populated by patients who enrol in trial will differ from trial to trial, even when inclusion criteria are the same. For example, consider two trials with identical designs and inclusion criteria, each open to patients with a particular disease who are between the ages of 18 and 65. In the first trial, 75% of patients who enrol might end up being over the age of 50 and none might be under the age of 40, while in the other trial, 75% of patients enrolling might end up being under the age of 30 and none might be over the age of 50. If age is an important prognostic factor for the disease in question, then the survival trajectories for patients treated with Standard of Care therapy in these two trials could end up looking quite different;
  • If researchers choose to compare the survival trajectory of patients treated with a New Drug in their present-day convenience sample (brown line) with the trajectory of patients treated with the Standard of Care therapy in another study (either past or present-day- Blue, Grey, Green, or Yellow lines), their inferences regarding the relative efficacy of the New Drug could be quite different than if they had relied on comparison with their concurrent control (pink line);
  • In the absence of extensive efforts to identify, and adjust for, between-convenience-sample differences in untreated prognosis, the only way to reliably isolate the relative intrinsic efficacies of the therapies being compared is to perform the between-arm comparison within the same convenience sample (brown compared with pink). Only in this way can researchers eliminate the possibility that factors OTHER than superior New Drug efficacy might explain the differing trajectories in their between-arm comparison.
3 Likes

Great points, excellently put! From my experience the second point is key, and under-appreciated. I find myself saying a lot that there is always selection bias in the patients that end up in a trial - this is obvious, we have eligibility criteria so they can’t be a random selection - and it’s different for every trial (even if the eligibility criteria are the same). So doing anything apart from comparing with a concurrent group that is randomly allocated (or matched in another way) is weakening any inferences, potentially to the point that they are meaningless. That’s not to say historical data can’t be useful, but we need to be careful and honest about their use.

2 Likes

And keep in mind @Stephen 's often made point that selection is a two-phase filter. First you select on inclusion criteria then patients further select themselves by having to agree to be randomized.

4 Likes

This 2022 publication on “Augmented” RCTs seems like it might be describing a precursor to the study design that’s now being called the “Randomized Non-Comparative Trial (RNCT)” (?)

https://pmc.ncbi.nlm.nih.gov/articles/PMC10089586/

Freidlin B, Korn EL. Augmenting randomized clinical trial data with historical control data: Precision medicine applications. J Natl Cancer Inst. 2023 Jan 10;115(1):14-20.

An excerpt:

“To mitigate concerns associated with nonrandomized studies that rely solely on historical controls, Pocock (18) suggested a design in which data from the control arm of a RCT is augmented with data from an “acceptable” historical control cohort. (“Historical control cohort” is used throughout as a shorthand for a nonrandomized control cohort.) The criteria given for when the historical data would be acceptable were quite stringent (Box 1)…In general, however, the Pocock criteria would rarely be satisfied. In particular, any time trends in the patient population, ancillary care, or diagnostic/response methodology can lead to a violation of Pocock criteria and misleading results (20).”

Reference #18 seems like it might represent the origin of the practice (?):

18. Pocock SJ. The combination of randomized and historical controls in clinical trials. J Chronic Dis. 1976;29(3):175-188.

I don’t understand statistics well enough to be able to highlight differences between what Pocock was proposing back in 1976 and what RNCT authors are actually doing today. It seems like Pocock was suggesting that, if certain strict criteria were met, it might be possible to “augment” the control arm of an RCT with historical data. But this is different than ignoring the RCT control arm data altogether and, rather than borrowing data from a historical control, simply comparing a new drug arm with a control arm from a historical study (?) Is this what RNCT authors are doing?

3 Likes

Correct. Borrowing historical information to inform RCT comparisons is defensible. Comparing with historical controls also actually has a rationale. What RNCTs do is in a league of its own. For example, some RNCTs will randomize patients to treatment or placebo and then not compare the two. Instead, they benchmark each separately against external or prespecified values. It is just so odd and inverts the usual logic by Pocock etc.: a historical-control design reaches for an external comparator because it has no concurrent one, whereas the RNCT discards the better (concurrent, randomized) control it already paid to obtain in favor of a weaker external benchmark. What is the point of the placebo arm if not to be compared with treatment?

Thus, the mental model is not that “RNCT = swap a historical control in for the RCT control arm”. RNCTs are actually stranger than that: they usually keep a concurrent randomized control and then decline to use it…

3 Likes

The best defence of the randomised non-comparative design that I’ve heard was a case where they said they just wanted to split the participants into two groups, without any intention of comparing them, and randomisation was the fairest way of doing that.

But that’s still a bit weird to me because the study was looking at two alternative treatments, so comparing them would seem a logical thing to do. Also they still described the study as a randomised controlled trial - which, without that comparison, it isn’t.

2 Likes