Oseltamivir (Tamiflu) again proves the need for RCT

This Lancet study (still in preprint) produced a crude number needed to kill of 17 in the guideline standard treatment of Influenza with critical illness (under guidelines for 17 years). This is more shocking then the CAST trial which changed the RCT design away from surrogate endpoints in the field of cardiology.

While I have been critical of RCT design, particularly the cause-agnostic RCT (CAR) when safe transport has not been considered. However, this RCT shows the consequences of failing to perform an RCT at all in the target illness state and instead simply transporting the results from another quite different state.

This is a CIR (causal-integrity RCT) : laboratory-confirmed influenza → critical-illness/organ-support restriction → randomization → oseltamivir vs no antiviral.

tauCIR(influenza)=E[Y|do(oseltamivir),D=influenza]
minus E[Y|do(no antiviral),D=influenza]

The original trial reversed by the CAST (which had a NNK of about 20) used a surrogate. The original Oseltamivir trial was in a different population and transported to critical care.

The guidelines in Australia have already been changed in red letters but let’s hope this REMAP CAP study is incorrect. 17 years of a crude NNK of 17 due to what appeared to be quite well reasoned guideline committee transport decisions would be too difficult to accept.

https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7172531

Here is a discussion about the REMAP CAP trial for Oseltamivir.

This is a VERY popular blog site with truly expert clinicians. I thought it would be useful for statisticians to read this blog to begin a discussion.

Note: This is not about cause agnostic RCT (which I understand no one want to discuss). This is a discussion is centered on the present core of clinical trial science. I hope to see robust discussion of this RCT.

Not a comment on REMAP-CAP, but there’s a fairly serious misunderstanding in that blog (though quite a common one I fear). It says:

“it can be valid for RCTs to adjust for confounding variables. This is a technique which is occasionally utilized to reduce noise and decrease the number of subjects required in the study. So it’s not necessarily wrong.”

They’re not “confounding variables” in an RCT. It’s not “occasionally utilized” - inclusion of relevant covariates should be standard. And “it’s not necessarily wrong” is just wrong! (Or maybe it’s getting into the realms of “not even wrong.”)

Full disclosure: I’m involved in REMAP-CAP as a member of the DSMB so I will not be making any comment on that trial.

3 Likes

Yes, I noticed that. But why wouldn’t you comment on the trial?

Does participation in REMAP-CAP constitute a conflict of interest? It certainly warrants disclosure, but it should not require silence. You are exactly the kind of academic we need examining this result.

If this trial is correct, oseltamivir may have caused iatrogenic death on an industrial scale. Yet transport to critical care certainly seemed reasonable. This is a lesson in the need for more and better focused RCT. Yet you cannot discuss the trial because you are part of REMAP-CAP? Where does that leave the clinician trying to decide how to treat the next patient?

I once read that @Stephen took considerable heat for identifying problems with “measurement” in trials, as though “measurement” was somehow outside his lane. My first thought was, “Why would he care?” “He is an established genius. Why should he allow other scientists to define his lane, especially when patient safety is at stake?”

I can tell you that except to learn from them, I don’t care what other scientists think about my carefully reasoned commenting choices relating to science. IMO no scientist or physician should. I understand grant dependency creates social constraints but it should not.

We had the same response on REMAP CAP corticosteroid domain in CAP from Prof. Angus. He simply defaulted to the statement that he hated “that trial”.

Please help us understand these trials.

There’s a Berry Consultants video on this trial/paper - obviiously this will give the view from the trial investigators and analysts. I haven’t watched it yet.

1 Like

Kert Viele of Berry Consultants has an excellent blog article here. One of several issues discussed is the relatively large effect of covariate adjustment on strengthening evidence for harm. We should remember that when there is outcome heterogeneity within treatment group, the assumptions of most statistical methods (including Fisher’s “exact” test and the more accurate Pearson \chi^2 test) are violated. These methods assume that every patient within a treatment group has the same probability of response. Unadjusted methods are intended to be used for homogeneous cases where the inclusion criteria are exceedingly tight. BTW I’ve never seen a clinical trial example where we expect to have no outcome heterogeneity.

3 Likes

Kert’s piece is (predictably) absolutely superb!

Presumably (which is kind of the same point about outcome heterogeneity) noncollapsibility of odds ratios may also be a contributory factor to the differences between the unadjusted and covariate-adjusted estimates? Inclusion of covariates pushes the estimate further away from 1.

Josh Farkas’s blog seems to illustrate a thing I see quite a lot (apologies if I’m maligning him here!), which is a general distrust of “adjusted” analyses. I think many people (still) do not understand why this is done and it’s often viewed as a form of statistical cheating, and certainly inferior to a “clean” unadjusted analysis - whereas it’s actually about using the information we have in an informative way. I try to avoid talking about “adjustment” as that sounds too close to shady manipulation, but instead talk about “models including covariate information” or something like that, which admittedly is a bit of a mouthful.

2 Likes

These misunderstandings about covariate adjustment are highly frustrating. Instead of talking about non-collapsibility of odds and hazards ratios we should be talking about the simple fact that unadjusted odds ratios were designed to be applied only to the case of homogeneous within-exposure-group outcome tendencies.

For Cox models, adjustment for baseline variables prevents power loss caused by the model’s assumptions being violated when easily explainable outcome variation is refused to be explained. A nice demonstration from the literature may be found here.

3 Likes

I’ve also (usually unsuccessfully) tried to convince people that including covariates in the model allows us to learn more about the treatment effect, because we include more information that explains the outcome (outcome heterogeneity).

I think this stems from people believing that you are messing with the data to find a positive result, which is of course an non issue if the analysis code is prespecified…

Also some trialists seem to be more averse to adjustment than clinicians. Idk why.

1 Like

These are longstanding problems in statistical education of non-statisticians. Besides getting the model more correct by not assuming outcome homogeneity, covariate adjustment provides THE basis for examining differential treatment effects through interactions.

1 Like

I think Dr. Farkas is thinking at a deeper level than you perceive. I can’t speak for him, but experienced critical care physicians inevitably apply Bayesian reasoning to almost everything.they analyze, not only to the oseltamivir trial itself, but to the meta-level reliability of the critical-care RCT literature itself

That matters when interpreting the REMAP-CAP oseltamivir raw result and the output derived from adjustment based on so many diverse covariates. A finding suggesting that a long-standing guideline treatment may actually be harmful naturally raises an extraordinary implications including potentially substantial iatrogenic mortality from years of recommended treatment.

An experienced intensivist is unlikely to accept an implication of that magnitude from a single RCT without considering the prior clinical trial history of the field.

And that history is abysmally sobering. Critical care physicians have lived through decades of RCT-based guideline reversals, many involving harm. Importantly, many of those trials were cause-agnostic RCTs where covariate adjustments comprise mathematical window dressing. Since statisticians generally do not distinguish them from cause-integrity RCTs, to clinicians they have all simply been presented as decades of “gold standard RCTs…reversed for harm”.

REMAP-CAP itself provides a striking example. Its recent corticosteroid domain tested another guideline-standard treatment and reported approximately an 89% posterior probability of harm. The results have been largely swept under the rug and the guidelines remain unchanged. This recently prior REMAP CAP trial (which also included influenza) and its covariate structure are discussed here.

https://pubmed.ncbi.nlm.nih.gov/42464342/

Of course the relevant prior is not that “Critical care RCTs are wrong.”, It is that the probability that any single critical-care RCT treatment effect, especially one from the REMAP CAP platform will prove clinically durable and transportable into guidelines may be relatively low.

That prior warrants caution before concluding from the oseltamivir trial that previous guidelines caused deaths on a massive scale AND it also makes it difficult to defend those previous guidelines simply because earlier evidence supported them especially since that “evidence” was transported from a different state of the disease. However failure mode analysis is not a part of clinical RCT based guideline culture.

There is, of course, a deeper structural explanation for the history of so many reversals in critical care. . Cause-agnostic trials can estimate different mixture-weighted effects while appearing to test the same clinical question. So the intensivist’s skeptical prior is not merely experiential, part of it has a structural mathematical basis. Previous broad defense by statisticians of widely discretionary covariate adjustment (as long as pre specified) have been made as if critical care RCTs are structurally similar. They are not. This may be true for cause integrity trials such as the oseltamivir trial but it is not true for much of the standard critical care RCT structures.

So when an expert critical care physician hesitates before accepting the oseltamivir result into a claim of massive historical iatrogenic harm, that hesitation is entirely compatible with deep, thoughtful and informed Bayesian reasoning. She is applying Bayes one level above the individual trial.

That’s all pretty reasonable! The implications for priors for trials in this area would be an interesting thing to discuss (not getting into that now though).

It’s also interesting that Dr Farkas is very critical of Bayesian methods, but then uses them to make one of his points:

But the pre-test probability that a 5-day course of oseltamivir is causing a big increase in mortality should be exceedingly low. Prior to this study, I might have guessed with maybe 95% certainty that oseltamivir would have a neutral impact on mortality. The results of this trial aren’t sufficiently robust to change my mind.

Maybe not surprising - Bayesian methods aren’t so different from the way we often think.

Some other points:

The prior probability is generally an issue (e.g., even “neutral” prior probability distributions don’t assign enough weight to the null hypothesis).

Two comments: first. giving special weight to a treatment effect of exactly zero (much more than to tiny positive or negative effects) just seems weird. Second, he’s missing a major elephant in the room, which is that traditional frequentist testing methods give WAY too much weight to prior values that are completely implausible.

Another issue is that the posterior probability of harm equates small harm (e.g., 0.0001% mortality risk) with a massive harm (90% mortality risk). This artificially forces the study to generate a binary result (oseltamavir either saves lives or kills people) – when in reality oseltamavir is probably doing nothing.

This is such a bizarre thing to say. One of the main issues with traditional NHST is that it leads to dichotomisation of results, and Bayesian methods actually give us ways of avoiding that.

2 Likes

Respectfully, what is actually interesting is NOT that critical-care physicians use Bayesian reasoning while questioning its application in critical-care trials. The interesting thing is that…after decades of RCT reversals, many for harm, trialists and statisticians expect more trust and mathematical discretion over priors, shrinkage, borrowing, covariate adjustment, and complex estimands without engaging in substantive failure-mode analysis or a corresponding increase in causal discipline. Indeed, we have already experienced such “strapped-on” Bayesian discretion in the REMAP-CAP cause agnostic RCT, (corticosteroid domain) which apparently no one defends.
Statisticians should embrace the healthy critical views of critical care physicians and engage in the failure mode analysis required to regain their trust.

Possibly I’ve not been paying attention, but I’m not sure what you mean by failure mode analysis. Could you elaborate?

Thanks for the question.

Failure Mode Analysis (FMA) is a systematic method for asking:

  1. *How can, or did, this system fail? (For example, an RCT or RCT derived guideline.)
  2. How can that failure be detected or prevented?

Highly advanced in aviation, FMA does not merely establish that its mathematics are correct. You map a process (eg structural casual modeling) before the trial and examine ways the design could fail.

A lengthy discussion of the “limitations of the trial” and even sensitivity analysis are not a substitute for FMA.

Failing to distinguish true FMA creates risk. Extensive analysis of downstream technical details may produce false confidence that a system has been adequately tested while leaving the underlying design failure hidden for decades.This generates a repeating pattern of “mathematical arrogance despite unmitigated implementation failure”. Repeated failures are attributed to what I call “amorphous academic excuses from the ether” such as chance, heterogeneity, implementation, or unavoidable uncertainty rather than to a reproducible structural defect.

The “RCT failure mode trilogy” is the first set of papers to combine historical and structural analysis to search for a primary failure mode underlying decades of reversals for harm of critical care trials and RCT-derived guidelines. I hope you will read them and offer critique, if interested in the causal stucture if RCTs.

Getting back to your earlier question, this helps explain why some critical care physicians hesitate to grant statisticians trust relavant greater mathematical discretion BEFORE these statisticians first formally investigate why earlier RCT derived guidelines repeatedly failed for decades. In other words, we know they don’t know why those earlier RCTs failed, and we can’t trust them to fix with math that which they have not even deeply examined.

Critical care physicians are the guardians of the critically ill and we know that flexible mathematics cannot correct an unidentified structural failure mode. We also know that, if applied to the same pathological trial structure, more complex mathematics can embellish the trial providing an “elegant pseudo-fix” thereby provided an alternative path bypassing the required FMA. We know this could bring to the field another decadal wave of failures and patient harm which can be each be eventually swept under the rug without deep structural FMA like the December 2023 Bayesian RE-MAP CAP cause agnostic RCT already was.

Here is a slide deck I am preparing for a conference teaching clinicians about RCT failure mode analysis. Note the components of RCT. Traditional discussions of RCT focus on the estimators (the statistical section in the center between the trial entry fate and the guideline application gate. This is incomplete. The instant trial (Influenza) was a causal integrity RCT. But most that have failure in critical care were cause agnostic RCT. Since statiticians don’t distinguish these species, clinicians begin to lose faith that RCTcan be trusted in the complex critical care environment in synthetic syndrome trials. There remains, of course a high degree of confidence in Causal integrity RCT but skepticism is bound to spill over when Bayesian platform trials demonstrate another reversal for harm of the standard of care.

We know that if the failure mode is structural then statistical fixes (eg (Bayesian design) will provide a false sense of beneficial action. So critical care physicians are thinking much more deeply then statistician’s perceptions of their understanding. This is why both groups have a certain degree of intellectual arrogance when dealing with the other (which arrogance emerges in these blogs). That’s a good thing. That is great science, but only if both sides have open minds.