# Information, Evidence, and Statistics for critical research appraisal

**URL:** <https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226>\
**Category:** education\
**Tags:** bayes, information, rct, interpretation\
**Created:** [November 19, 2022, 11:24pm UTC](https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226 "2022-11-19T23:24:12Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![R\_cubed](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/r_cubed/32/1518_2.png) [@R\_cubed](https://discourse.datamethods.org/u/R_cubed)\
**Post date:** [November 19, 2022, 11:24pm UTC](https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226/1 "2022-11-19T23:24:12Z")

</div>

In another post about understanding p-values, ESMD made an important observation:

> **ESMD wrote:**  
> I suspect hat much of the confusion among physicians learning critical appraisal stems from the fact that our teachers are often (?usually) practising MDs, who might not understand the historical foundations of statistics well enough to address questions like this. As a result, these types of questions get glossed over.

> [@Random sampling versus random allocation/randomization- implications for p-value interpretation](https://discourse.datamethods.org/t/random-sampling-versus-random-allocation-randomization-implications-for-p-value-interpretation/6177/9):
>
> Randomization may be random sampling or random allocation. The phrase in block quotes above serves to justify the fact that we use p values in RCTs. But Dr.Senn’s blog piece linked above emphasizes how important it is to not confuse random sampling with random allocation. In the first instance (random sampling) we seem to have access to the entire population of potential interest, whereas in the second instance (randomization that occurs in the context of an RCT), we are using a convenienc…

In that thread, @Sander_Greenland mentioned Richard Royall – the biostatistician and author of an important philosophical text **Statistical Evidence: A Likelihood Paradigm** (published in 1997) where he wrote in the Preface:

> Blockquote  
> …Standard statistical methods regularly lead to the misinterpretation of scientific studies. The errors are usually quantitative, when the evidence is judged to be stronger (or weaker) than it really is. But sometimes they are qualitative – sometimes one hypothesis is judged to be supported over another when the opposite is true. **These misinterpretations are not a consequence of scientists misusing statistics. They reflect instead a critical defect in current theories of statistics.**

The subtle differences in the Neyman-Pearson and Fisher interpretation of p-values Sander describes, is elaborated on in this article I highly recommend:

> **[Frontiers | Fisher, Neyman-Pearson or NHST? A tutorial for teaching data testing](https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2015.00223/full)**
>
> Despite frequent calls for the overhaul of null hypothesis significance testing (NHST), this controversial procedure remains ubiquitous in behavioral, social...

Royall’s monograph was published only 1 year **after** the widely cited Sackett et. al paper on Evidence Based Medicine (EBM):  
[https://www.bmj.com/content/312/7023/71?%25253Cbr](https://www.bmj.com/content/312/7023/71?%25253Cbr)

> Blockquote  
> Evidence based medicine is the conscientious, explicit, and judicious use of current best evidence in making decisions about the care of individual patients.

Later evolution of EBM involved elevating statistical fallacies into so-called “hierarchies of evidence” that are logically incoherent when looked at carefully. Philosophers have had a field day tearing apart the simplistic “hierarchies of evidence”, which is how I ended up here way back in 2019!

I independently discovered Stegenga’s argument from thinking about “evidence” as closely related to the economic problems of Welfare and Social Choice Theory, where Kenneth Arrow’s [famous result](https://plato.stanford.edu/entries/arrows-theorem) seemed applicable. Stegenga’s paper preceded my discovery by a few years, but his impossibility proof was very similar to my own thinking.

> **[An impossibility theorem for amalgamating evidence - Synthese](https://link.springer.com/article/10.1007/s11229-011-9973-x)**
>
> Amalgamating evidence of different kinds for the same hypothesis into an overall confirmation is analogous, I argue, to amalgamating individuals’ preferences into a group preference. The latter faces well-known impossibility theorems, most famously...

> **[Down with the Hierarchies - Topoi](https://link.springer.com/article/10.1007/s11245-013-9189-4)**
>
> Evidence hierarchies are widely used to assess evidence in systematic reviews of medical studies. I give several arguments against the use of evidence hierarchies. The problems with evidence hierarchies are numerous, and include methodological...

A more “statistical” critique is found here:

> [@Article: Going from evidence to recommendations: Can GRADE get us there?](https://discourse.datamethods.org/t/article-going-from-evidence-to-recommendations-can-grade-get-us-there/5445):
>
> I thought the following was worth reading: Mercuri, M, Baigrie, B, Upshur, REG. Going from evidence to recommendations: Can GRADE get us there? J Eval Clin Pract. 2018; 24: 1232– 1239. [https://doi.org/10.1111/jep.12857](https://doi.org/10.1111/jep.12857) From the abstract: Blockquote In this paper, we reveal several issues with the underlying logic of GRADE that warrant further discussion. First, the definitions of the “grades of evidence” provided by GRADE, while explicit, are functionally vague. Second, the “criteria for a…

🆕 The fundamental problem of EBM hierarchies and checklists is that they take frequentist pre-data design criteria and apply them in a post-data context. This violation of the likelihood principle is true of the Neyman-Pearson perspective, not the Fisherian one that Sander alluded to. Casella describes the problem in this paper:

**Goutis, C., & Casella, G. (1995).** Frequentist Post-Data Inference. International Statistical Review / Revue Internationale de Statistique, 63(3), 325–344. [Frequentist Post-Data Inference on JSTOR](https://doi.org/10.2307/1403483)

> Blockquote  
> The end result of an experiment is an inference, which is typically made after the data have been seen (a post-data inference). **Classical frequency theory has evolved around pre-data inferences, those that can be made in the planning stages of an experiment, before data are collected. Such pre-data inferences are often not reasonable as post-data inferences, leaving a frequentist with no inference conditional on the observed data.** We review the various methodologies that have been suggested for frequentist post-data inference, and show how recent results have given us a very reasonable methodology. We also discuss how the pre-data/post-data distinction fits in with, and subsumes, the Bayesian/frequentist distinction

These **critical defects** have so deformed the “peer reviewed” literature (as well as clinical practice guidelines based upon them) that large areas of scholarship are practicing cargo-cult (from Feynman) or “pathological” science (from psychologist and psychometrician Joel Michell).

> [@Statistics Applied Using Threshold Science Fails to Deliver Progression Toward Truth](https://discourse.datamethods.org/t/statistics-applied-using-threshold-science-fails-to-deliver-progression-toward-truth/4940):
>
> A recent study addressing a very fundamental question of considerable importance to the emergency room physician demonstrates the failure of the statistics applied using “Threshold Science” to progress the science toward an answer. Here, the question of the utility of the band count dates back tot the turn of the century. This question is asked here. The fundamental controversy is . Are the presence of an elevated absolute band count and/or elevated % bands in blood predictive of excess a…

> [@What is a fake measurement tool and how are they used in RCT](https://discourse.datamethods.org/t/what-is-a-fake-measurement-tool-and-how-are-they-used-in-rct/3955):
>
> I have linked an article below published this month. This provides more evidence that the measurement tools used for critical care RCT esp. for sepsis and mortality are often fake. Those who have studied this issue did not need more evidence but perhaps if enough of these studies are published the perceived bright star of SOFA will fade (as SIRS has faded). This study is no surprise given there has not been a single reproducibly positive sepsis RCT for 30 yrs. Now fake is a harsh word and I d…

A snarky Bayesian might describe it as “researchers pretending to know the difference between their priors and their posteriors.”

My quasi-formal analysis of this in the context of parametric vs. ordinal models was posted here:

> [@Preprint: Analysis of Likert Type Data Using Metric Methods](https://discourse.datamethods.org/t/preprint-analysis-of-likert-type-data-using-metric-methods/5536):
>
> Dennis Boos and Judy Chen (2022). Analysis of Likert-Type Data Using Metric Methods. new I edited this to bring the main points earlier, and the decision theoretic justification later in the post. Introduction The authors defend the common practice of using parametric models on ordinal data on the basis of pragmatics (ie.“easier” to for researchers understand), and perceived robustness in that parametric methods on empirical or simulated data sets are not sufficiently inaccurate that skepti…

> [@Evidence Based Decision Making and Meta-Analysis -- An Information Theoretic Critique](https://discourse.datamethods.org/t/evidence-based-decision-making-and-meta-analysis-an-information-theoretic-critique/5033):
>
> Andrew Gelman posted an interesting review of a paper that documents the [failure of international public health authorities](https://statmodeling.stat.columbia.edu/2021/10/22/how-did-the-international-public-health-establishment-fail-us-on-covid-by-explicitly-privileging-the-bricks-of-rct-evidence-over-the-odd-shaped-dry-stones-of-mechanistic-evidence/) throughout the pandemic. Blockquote It might sound silly to say that people are making major decisions based on binary summaries of statistical significance from seriously flawed randomized studies, but that seems to be what’s happening. After much reading and lurking here, I’ve come to believe that @Sander is right, and statistics needs to be taught in the context of …

**Related Threads**

> [@Definition of statistics](https://discourse.datamethods.org/t/definition-of-statistics/6191/2):
>
> I think Herman Chernoff and Lincoln Moses defined statistics perfectly on the first page of Chapter 1 of their book Elementary Decision Theory: Blockquote …today’s statistician will say that statistics is concerned with decision making in the face of uncertainty. Decision analysis is woefully neglected in basic stats. Why do we conduct research in medical science other than to make better decisions? It is much easier to test recall of the formula of the standard error than it is to disc…

> [@Definition of statistics](https://discourse.datamethods.org/t/definition-of-statistics/6191/10):
>
> Blockquote “Bayesian decision making” it’s not very common in med research, as far as I can see. And it is also not very commonly meant in intro statistics books The fact that a decision theory perspective is mostly absent from medical research and intro stat books I see as an important cause of poor research conduct and misinterpretation of statistical methods. I have not seen any rebuttal to the claim that the maximization of utility, when applied to the scientific research domain is simp…

> [@Stability of results from small RCTs](https://discourse.datamethods.org/t/stability-of-results-from-small-rcts/2287/19):
>
> Blockquote Randomization does not guarantee exchangeability. The very definition of random implies imbalances in prognostic factors could happen, even if they are more and more rare as the sample size increases. This very important point is carefully elaborated in the following: So that leads me to the following question: Are small randomized trials worth conducting? In a previous post @f2harrell wrote: Blockquote Most statisticians I’ve spoken with who frequently collaborate with …

> [@2019 Nobel Prize in Economic Sciences is about limitations of RCT](https://discourse.datamethods.org/t/2019-nobel-prize-in-economic-sciences-is-about-limitations-of-rct/2634/14):
>
> This POV is the common one in discussions regarding the philosophy or foundations of statistics. The model where an external skeptic has a more challenging prior has been taken for granted in that if you think about 2 scenarios – you flip a coin and verify heads vs. someone flips a coin and reports heads, the latter probability should be less than or equal to yours. So it seems rational to discount reports from external agents, (unless you question your own perceptions!). Therefore, it is rea…

> [@Who should define reasonable priors and how?](https://discourse.datamethods.org/t/who-should-define-reasonable-priors-and-how/2935/13):
>
> I thought I’d bump this thread after finding a valuable discussion between Andrew Gelman and Sander Greenland on this precise issue of setting priors. Specifying a prior distribution for a clinical trial: What would Sander Greenland do? [https://statmodeling.stat.columbia.edu/2008/02/05/specifying\_a\_pr\_1/](https://statmodeling.stat.columbia.edu/2008/02/05/specifying_a_pr_1/) Specifying a prior distribution for a clinical trial [https://statmodeling.stat.columbia.edu/2008/01/24/specifying\_a\_pr/](https://statmodeling.stat.columbia.edu/2008/01/24/specifying_a_pr/) My posts in this thread were some speculations I had about how 2 hone…

> [@What is the best way to improve Statistical Thinking?](https://discourse.datamethods.org/t/what-is-the-best-way-to-improve-statistical-thinking/5797):
>
> To make better decisions it’s not enough to produce accurate predictions for decision-makers and it’s not enough to use proper scoring rules, it’s also important that our decision makers will have good problalistic thinking along with their unique domain knowledge. What will be a good way to improve probabilistic thinking for decision-makers? Maybe all clinicians should spend some time playing poker sweat_smile

> [@Necessary/recommended level of theory for developing statistical intuition](https://discourse.datamethods.org/t/necessary-recommended-level-of-theory-for-developing-statistical-intuition/5648):
>
> I brought this up after Day 1’s lecture but I thought it might warrant further discussion/elaboration. (Apologies for the length!) I am of the opinion that a “strong” background/foundation in theory is important to becoming a good applied statistician. I leave “strong” in quotation marks to be purposefully vague, as I don’t think it can really be quantified. The concept may be related to “mathematical maturity” or “developing statistical intuition”. (I welcome contrasting viewpoints of course!) …

> [@Paper on European Radiology Experimental on statistical significance](https://discourse.datamethods.org/t/paper-on-european-radiology-experimental-on-statistical-significance/5298/7):
>
> new I accidentally deleted this (post 4) trying to copy the link for another thread; apologies for repost and bump. Blockquote In this article, we discuss the value of p value and explain why it should not be abandoned nor should the conventional threshold of 0.05 be modified. It seems to me they continue the confusion between \alpha and p. To continue the use of a 0.05 threshold is analogous to saying the same prior should be used in every problem, regardless of the precision of the stu…

> [@Statistical Misinterpretations and Clinical Practice Guidelines: any examples?](https://discourse.datamethods.org/t/statistical-misinterpretations-and-clinical-practice-guidelines-any-examples/4372):
>
> In this important paper on the interpretation of statistics in medicine, the authors write: Blockquote Misinterpretation and abuse of statistical tests, confidence intervals, and statistical power have been decried for decades, yet remain rampant. A key problem is that there are no interpretations of these concepts that are at once simple, intuitive, correct, and foolproof. Instead, correct use and interpretation of these statistics requires an attention to detail which seems to tax the patie…

---

<div class="post-metadata">

**Author:** ![ESMD](https://discourse.datamethods.org/letter_avatar_proxy/v4/letter/e/9e8a1a/32.png) [@ESMD](https://discourse.datamethods.org/u/ESMD)\
**Post date:** [November 20, 2022, 5:06pm UTC](https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226/2 "2022-11-20T17:06:33Z")

</div>

Hi Robert

Lots of interesting links and different ideas in your post- some related to statistical methods and others to the validity of core EBM concepts. I just wanted to pick up on one thread from what you’ve written and to offer a physician’s perspective.

Regarding evidence “hierarchies”: maybe this is an unpopular opinion, but I feel pretty strongly that they _do_ actually have a role in medicine, just not in the way they are often perceived by non-statisticians and non-clinical people. I think most practising physicians would agree that the traditional pyramid (RCT at the top) applies in assessments of treatment _efficacy_ but not necessarily in assessment of treatment _safety_. I’ve noticed that those who tend to argue against an efficacy pyramid are often those in non-clinical professions.

The idea of first “doing no harm” really is a North Star in medical practice. Since it’s possible for a treatment that lacks intrinsic efficacy to have side effects, we strive, as often as possible, to apply therapies that have proven themselves to have meaningful intrinsic efficacy. Applying “dud” therapies to patients has at best a neutral effect, but more likely a net detrimental effect. Even if a therapy lacking intrinsic efficacy were to have very few side effects, there would usually be a cost to the patient or healthcare system in terms of time and money.

Those who extol the virtues of observational evidence point out that patients enrolled in clinical trials are often quite different from those who subsequently receive the treatment in the postmarket setting. This is true. They argue that we have no “guarantee” that patients who would not have fulfilled the inclusion criteria for the pivotal trial(s) (on whose basis the therapy was approved) will benefit from the treatment. This is also true. But if our threshold for every clinical decision we made were a prior demonstration that someone _exactly like_ the patient in front of us had been documented to “respond” to the treatment, we would be waiting forever, since _no two patients are exactly the same_. We’d spend every day in the office paralyzed with indecision. Very often, the best we can do is to ensure that we are not prescribing an inert/dud therapy- this is the “bet” we make as doctors when we write a prescription. We are saying,_“since this drug has shown that it has the ability to improve outcomes in at least__ **some** _ _people for this indication, it is_ _ **not unreasonable** _ _to bet that it might also improve the outcome for **the patient in front of us**.”_ Whether the _particular_ patient in front of us will end up “responding” to the therapy is something we often can’t accurately predict in advance, and in fact _may never actually learn_ (especially in the case of many _preventive_ therapies like statins, as compared with treatments like bronchodilators, which are used to treat _symptomatic_ conditions).

In short, the _over-riding_ priority of physicians is to avoid exposing patients to dud treatments.

A well-designed RCT will 1) reveal a treatment’s intrinsic efficacy, if it exists; and 2) minimize the chance that we will _infer_ intrinsic efficacy in situations where it _doesn’t_ actually exist. In medicine, there are _many_ examples of observational assessments of efficacy (involving either non-drug therapies or approved drugs that had been used “off-label”), which did not stand up to subsequent experimental testing. When results of observational studies and RCTs of efficacy conflict, physicians will virtually _always_ preferentially trust the RCT (provided it was well-done).

Ultimately, I don’t think that there will ever come a time when observational studies are considered by clinicians to be as compelling as RCTs for demonstrating efficacy. I never really bought into the common argument: “but we can’t do an RCT for _everything_, so _some_ evidence is better than _nothing_…isn’t it?” Actually, sub-optimal evidence often IS worse than no evidence, at least where consequential decisions for patients and healthcare systems are concerned. Sub-optimal evidence is insufficiently reliable and lends an undue veneer of certainty in situations where certainty is unwarranted. In turn, unwarranted certainty can profoundly affect clinical decision-making (specifically, the mental risk/benefit calculations that are performed many times every day by practising physicians), potentially with serious consequences for the patient.

Finally, and at the risk of going down a rabbit hole that nobody wants to go down, it’s important to flag a notable exception to the above framework. These are situations in which it’s _essential_ to apply a treatment even in the _absence_ of RCT evidence. The scenarios in question are public health emergencies, during which risk/benefit calculations are often best viewed through the lens of _physics_ (and _common sense_), rather than human biology and efficacy “pyramids” (e.g., the decision to recommend masking to decrease the spread of an airborne illness).

---

<div class="post-metadata">

**Author:** ![Pavlos\_Msaouel](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/pavlos_msaouel/32/739_2.png) [@Pavlos\_Msaouel](https://discourse.datamethods.org/u/Pavlos_Msaouel)\
**Post date:** [November 20, 2022, 7:13pm UTC](https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226/3 "2022-11-20T19:13:41Z")

</div>

> [@ESMD](#):
>
> I think most practising physicians would agree that the traditional pyramid (RCT at the top) applies in assessments of treatment _efficacy_ but not necessarily in assessment of treatment _safety_. I’ve noticed that those who tend to argue against an efficacy pyramid are often those in non-clinical professions.

Will have to write a separate paper on the topic at some point but the reason why such evidence hierarchies are problematic is that RCTs have the advantage that they yield (in theory) valid p-values and frequentist CIs for the differences between cohorts but do not offer an “apples-to-apples” comparison, unless we block (the less accurate term is “stratify”) in the design and adjust for strong prognostic factors in the statistical model.

Conversely, there are meticulously detailed and accurate case reports or case series that have yielded more information, in terms of bits of surprisal, describing patients very similar to the ones I may encounter in clinic. All this of course keeping in mind that my clinical practice is highly skewed towards rare and aggressive kidney cancers. But there is certainly a fair amount of RCTs in oncology that are uninformative or as @f2harrell sometimes puts it: the money was wasted. A common reason is that trialists may falsely believe that the random treatment assignment is enough on its own to save a poorly conducted RCT. Thus key aspects of internal validity are ignored, including patient adherence, accurate data collection, as well as proper blocking and covariate adjustment.

This does not invalidate the tremendous value that RCTs can provide in certain contexts. A good strategy is to use whatever source of information works best depending on the problem at hand. For some, this may plausibly consist predominantly of RCTs.

---

<div class="post-metadata">

**Author:** ![ESMD](https://discourse.datamethods.org/letter_avatar_proxy/v4/letter/e/9e8a1a/32.png) [@ESMD](https://discourse.datamethods.org/u/ESMD)\
**Post date:** [November 20, 2022, 7:56pm UTC](https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226/4 "2022-11-20T19:56:30Z")

</div>

Yes, like any other type of study, RCTs have to be _well designed_ in order to be valuable. You are essentially highlighting the fact that RCTs aren’t valuable if they lack assay sensitivity. And if you select your trial subjects poorly, your trial will lack assay sensitivity- it will fail to identify the intrinsic drug efficacy that _is_ actually present.

If you enrol patients with tumours that are being driven by wildly _variable_ biologic pathways into a clinical trial of a therapy that is designed only to affect a _specific_ biologic pathway, you won’t be surprised if your trial doesn’t identify the intrinsic efficacy of that therapy, since only a small fraction of trial subjects _even stood a chance_ to respond. You have wasted your money and have falsely concluded that the therapy “lacks efficacy,” when the problem was instead that you didn’t think carefully enough about the _design_ of your trial. The root failure here, though, is not _randomization as a concept_, but rather a failure _to sufficiently understand the disease you are trying to treat_ before you start testing treatments for it. This was the underlying point in all those other threads about the poor track record of RCTs for critical care “syndromes.”

As you say, oncology is an outlier compared with other branches of medicine, in the sense that the spectrum of underlying pathology being treated is likely far more complex and variable than the pathology we see in other fields (e.g., cardiology). At the end of the day, though, if you, as an oncologist, have a choice between basing your decision to use a new therapy on a case report involving a single patient who survived much longer than expected, versus an RCT involving patients with similar tumour biology, in which the same therapy was tested against standard of care, I expect that you will consider the RCT results to be more compelling (?)

---

<div class="post-metadata">

**Author:** ![R\_cubed](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/r_cubed/32/1518_2.png) [@R\_cubed](https://discourse.datamethods.org/u/R_cubed)\
**Post date:** [November 20, 2022, 8:43pm UTC](https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226/5 "2022-11-20T20:43:13Z")

</div>

> [@ESMD](#):
>
> I think most practising physicians would agree that the traditional pyramid (RCT at the top) applies in assessments of treatment _efficacy_ but not necessarily in assessment of treatment _safety_. I’ve noticed that those who tend to argue against an efficacy pyramid are often those in non-clinical professions.

I’m glad you expressed this, because this common attitude encouraged by evidence hierarchies is neither ethically optimal, nor statistically correct.

Richard Royall (who I mentioned above) wrote an excellent paper discussing the ethics and scientific issues around the ECMO trial for infants. I cannot express the logic better than by offering some quotes and recommending reading it.

[https://projecteuclid.org/journals/statistical-science/volume-6/issue-1/Ethics-and-Statistics-in-Randomized-Clinical-Trials/10.1214/ss/1177011934.full](https://projecteuclid.org/journals/statistical-science/volume-6/issue-1/Ethics-and-Statistics-in-Randomized-Clinical-Trials/10.1214/ss/1177011934.full)

Some key quotes:

1. From the abstract:

> Blockquote  
> We urge that the view that randomized clinical trials are the only scientifically valid means of resolving controversies about therapies is mistaken, and we suggest that a faulty statistical principle is partly to blame for this misconception.

1. In discussing the 1985 ECMO trial where the investigators felt “compelled to conduct a prospective, randomized study” on infants with severe respiratory distress without equipoise, Royall has this to say

> Blockquote  
> This is particularly disturbing to me as a statistician, because it is we statisticians who are largely responsible for creating attitudes and assumptions that compelled this study… In the next section, I propose that the above doctrine springs at least in part, adherence to a faulty statistical principle.

1. In presenting alternatives to the randomized “play the winner” design, he says:

> Blockquote  
> Science desires randomized clinical trials, **it does not demand them.** [my emphasis] Moreover, the importance of randomization is exaggerated when historical controls are the only alternative discussed…many of the weaknesses of historical controls can be avoided by using concurrent, nonrandomized controls. And of course, comparing treatment to historical and concurrent controls can provide even greater evidence of efficacy.

4 In discussing the randomization principle (ie. only randomized experiments provides a basis for inference) he writes:

> Blockquote  
> Statistical theory explains why the randomization principle is unacceptable. It does this in terms of the concepts of conditionality and likelihood…the conditional randomization distribution is degenerate, assigning one to the actual allocation used … the only “inference” on the observed data is “I saw what I saw”

---

<div class="post-metadata">

**Author:** ![ESMD](https://discourse.datamethods.org/letter_avatar_proxy/v4/letter/e/9e8a1a/32.png) [@ESMD](https://discourse.datamethods.org/u/ESMD)\
**Post date:** [November 20, 2022, 9:01pm UTC](https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226/6 "2022-11-20T21:01:49Z")

</div>

> [@R\_cubed](#):
>
> I’m glad you expressed this, because this common attitude encouraged by evidence hierarchies is neither ethically optimal, nor statistically correct.

I think we’ll have to agree to disagree about this- that’s okay. Unfortunately, I don’t have a list of citations readily available to support my view, but I’m pretty sure that it’s not unique. The priority placed on RCT evidence as the basis for drug approval by regulatory agencies around the world for many decades now provides indirect evidence that this is not a minority opinion.

It’s always possible to find experts on either side of an issue- I guess ultimately we all have to “pick sides…”

---

<div class="post-metadata">

**Author:** ![Pavlos\_Msaouel](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/pavlos_msaouel/32/739_2.png) [@Pavlos\_Msaouel](https://discourse.datamethods.org/u/Pavlos_Msaouel)\
**Post date:** [November 20, 2022, 9:28pm UTC](https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226/7 "2022-11-20T21:28:08Z")

</div>

> [@ESMD](#):
>
> At the end of the day, though, if you, as an oncologist, have a choice between basing your decision to use a new therapy on a case report involving a single patient who survived much longer than expected, versus an RCT involving patients with similar tumour biology, in which the same therapy was tested against standard of care, I expect that you will consider the RCT results to be more compelling (?)

Excellent question at the heart of this discussion. The honest practical answer is that such a choice is unnecessary. Why not just use both instead of choosing between them? If we have a well conducted internally and externally valid RCT relevant to our patient in clinic then it would be prudent to take advantage of all the statistical inferential machinery built around RCTs to help us generate plausible uncertainty estimates. But if there is also an additional detailed case report with surprising findings refuting strong hypotheses relevant to our patient then it would be prudent to use that information as well in our final inference. We can think of case reports as patient stories backed by evidence that can be used deductively to refute theories and [are most informative when they are anomalies](http://www.stat.columbia.edu/~gelman/research/published/storytelling.pdf). This is why researchers writing such case reports should treat them with the utmost respect and rigorously provide as much empirical evidence on the details of the story as possible, especially on the aspects of the case that are surprising and incompatible with current theories.

What if the hypothetical RCT and case report are contradictory? Then we would need to probe why that is and which source of information is more relevant to our current patient.

On a theoretical / philosophical level, the dilemma you posed corresponds to one of pure randomization inference (perfect robustness) versus full conditioning (perfect patient relevance). This is a [special case of the variance / bias trade-off](https://www.annualreviews.org/doi/full/10.1146/annurev-statistics-010814-020310) and it is reasonable to think of this trade-off as part of a continuum and not a simple dichotomy. However, if strongly pressed, I have already [nailed my colors to the mast](https://www.tandfonline.com/doi/full/10.1080/07357907.2022.2084621) by claiming that patient relevance yielded by biological insights is more fundamental than robustness for patient-centered inferences. This perhaps is personal preference and is not surprising given my day job as a wet lab-based physician scientist.

---

<div class="post-metadata">

**Author:** ![ESMD](https://discourse.datamethods.org/letter_avatar_proxy/v4/letter/e/9e8a1a/32.png) [@ESMD](https://discourse.datamethods.org/u/ESMD)\
**Post date:** [November 20, 2022, 10:14pm UTC](https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226/8 "2022-11-20T22:14:00Z")

</div>

What’s that line from The Godfather? “Just when I thought I was out…they pull me back in.” I’m trying to get a mountain of charts done this afternoon, Pavlos, but you keep writing interesting things and I feel I must respond… 😉

I totally agree re the strong evidentiary value of some case reports- when they are well-reported, and in the right clinical situation, they can be incredibly compelling. Examples include scenarios where 1) the disease in question has a strongly predictable adverse outcome in the absence of treatment (e.g., death is inevitable within a very short period of time- with rabies or very aggressive cancers, for example), or 2) when there is evidence of “positive dechallnge” and/or “positive rechallenge.”

If I saw a patient survive for 1 year a condition that was previously known to be fatal in 100% of patients within 3 months, I too would jump for joy- such a finding would likely be compelling enough to change practice. But of course, these types of scenarios are not the norm in medicine, except perhaps in the context of aggressive infectious diseases and cancers (and maybe certain degenerative neurologic diseases). When death in a short time is a clinical _certainty_, there is very little or no downside to trying new therapies outside the confines of RCTs. But as soon as we move to analyzing case reports involving diseases with any wider “spectrum” of prognosis than certain death in a very short time, the potential downsides of acting on anecdotal evidence manifest themselves (unless positive dechallenge/rechallenge can be demonstrated). If this were not true, then why do we even bother to do clinical trials in oncology _at all_? Why don’t we simply start applying new therapies based on detailed understanding of tumour biology alone and skip the trials?

---

<div class="post-metadata">

**Author:** ![Pavlos\_Msaouel](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/pavlos_msaouel/32/739_2.png) [@Pavlos\_Msaouel](https://discourse.datamethods.org/u/Pavlos_Msaouel)\
**Post date:** [November 20, 2022, 10:50pm UTC](https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226/9 "2022-11-20T22:50:36Z")

</div>

> [@ESMD](#):
>
> If this were not true, then why do we even bother to do clinical trials in oncology _at all_? Why don’t we simply start applying new therapies based on detailed understanding of tumour biology alone and skip the trials?

Good point. As you mentioned above, we can do trials even in one patient challenging them with a certain therapy to learn more about how their disease and constitution react to it. In general, trials and other experiments / interventions tend to be more informative for my applications than observational data. The advantage of RCTs over other experimental evidence is a brilliant inferential framework meticulously developed over the last century to provide plausible and statistically robust estimates of relative treatment effect uncertainty from RCTs.

A practical example I encounter regularly with my patients is when and whether to offer nephrectomy in metastatic clear cell renal cell carcinoma (the most common kidney cancer)? The only contemporary RCT addressing this question [was published in the New England Journal of Medicine](https://www.nejm.org/doi/10.1056/NEJMoa1803675) and is notoriously uninformative: it used a non-inferiority design (with all its [limitations](https://pubmed.ncbi.nlm.nih.gov/26607238/)) and had a slew of internal validity issues, including the fact that many patients did not follow the treatment assigned in their cohort (see Figure 1), which is the reason why the per-protocol analysis was never presented (it would have been inconclusive as evidenced by the weird semi-per protocol analyses called PP1 and PP2 buried in the supplemental results). The per-protocol analysis was not shown even though in the [same journal a year prior](https://www.nejm.org/doi/10.1056/NEJMra151006) it was recommended to show both intention-to-treat and per-protocol analyses for non-inferiority trials and that concerns should be raised if the two analyses are inconsistent with each other.

I do think the authors should be commented for attempting such an RCT. Reading it you can feel the blood, sweat, and tears needed to get it to the finish line. And publishing it in a high impact journal can be considered a reasonable motivation for others to continue designing and conducting ambitious RCTs. But from a practical standpoint, there is little actionable information yielded from this RCT. Thankfully there are [many other sources of information](https://www.nejm.org/doi/full/10.1056/NEJMe1806331) that do allow us to make informed recommendations for our patients regarding nephrectomy.

This happens often in various situations and probably not only in oncology. Those not familiar with the nuances of this RCT will simply take it as “level I evidence” or something like that, which can bias recommendations. But this is the value of knowing the subject matter well. Impossible to do for everything and I have the utmost respect for physicians dealing on a regular basis with multiple diverse conditions. The practical value of guidelines, often based on evidential hierarchies, is that they provide an often reasonable roadmap for generalists.

---

<div class="post-metadata">

**Author:** ![R\_cubed](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/r_cubed/32/1518_2.png) [@R\_cubed](https://discourse.datamethods.org/u/R_cubed)\
**Post date:** [November 20, 2022, 11:23pm UTC](https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226/10 "2022-11-20T23:23:31Z")

</div>

> [@ESMD](#):
>
> The priority placed on RCT evidence as the basis for drug approval by regulatory agencies …

Regulatory authorities have a more nuanced approach to this issue than simplistic hierarchies would lead you to believe.

This is a question of logic and mathematics that the EBM literature simply gets wrong.

For controlled trials (which RCTs are a proper subset):  
The CONSORT Statement – explanation and elaboration ([pdf](https://www.consort-statement.org/Media/Default/Downloads/CONSORT%202010%20Explanation%20and%20Elaboration%20Document-BMJ.pdf)), pg 9 box 2. states regarding the adaptive allocation method known as minimization:

> Blockquote  
> Nevertheless, in general, trials that use minimization **are considered methodologically equivalent to randomized trials, even when a random element is not incorporated.**

From Doug Altman’s 2005 article:  
[https://www.bmj.com/content/330/7495/843](https://www.bmj.com/content/330/7495/843)

> Blockquote  
> Minimisation is a valid alternative to ordinary randomisation and has the advantage, **especially in small trials** , that there will be only minor differences between groups in those variables used in the allocation process.

I could point to similar statements by the FDA and EMA.

My point: randomization is (very) **useful** but not **necessary** for causal inference.

A stronger case for skepticism of observational reports is that the allocation mechanism is not controlled by the experimenter. While true, this is something that can be overcome, and should be judged on a case by case basis, not decided against a priori simply by a design feature that becomes irrelevant after the data are in.

[https://projecteuclid.org/journals/statistical-science/volume-17/issue-3/Covariance-Adjustment-in-Randomized-Experiments-and-Observational-Studies/10.1214/ss/1042727942.full](https://projecteuclid.org/journals/statistical-science/volume-17/issue-3/Covariance-Adjustment-in-Randomized-Experiments-and-Observational-Studies/10.1214/ss/1042727942.full)

---

<div class="post-metadata">

**Author:** ![ESMD](https://discourse.datamethods.org/letter_avatar_proxy/v4/letter/e/9e8a1a/32.png) [@ESMD](https://discourse.datamethods.org/u/ESMD)\
**Post date:** [November 20, 2022, 11:34pm UTC](https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226/11 "2022-11-20T23:34:44Z")

</div>

I’ll be the first person to admit that I’m no stats expert. But I’m pretty familiar with the type of evidence that regulatory agencies expect in making decisions around drug approval and withdrawal- I worked for one for 10 years (!) The job provided a considerable education in how observational evidence is handled by regulatory authorities. Is there a long list hidden somewhere of drugs that were approved without evidence from RCTs?

---

<div class="post-metadata">

**Author:** ![R\_cubed](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/r_cubed/32/1518_2.png) [@R\_cubed](https://discourse.datamethods.org/u/R_cubed)\
**Post date:** [November 20, 2022, 11:42pm UTC](https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226/12 "2022-11-20T23:42:34Z")

</div>

I’m not specifically talking about drugs, but the logic of causal inference for _any_ intervention. Randomization simply isn’t **necessary** , and this is clear from decision theory.

A case study:

> **[Immobilization in External Rotation Versus Internal Rotation After Primary...](https://pubmed.ncbi.nlm.nih.gov/26116355/)**
>
> Immobilization in external rotation is not significantly more effective in reducing the recurrence rate after primary anterior shoulder dislocation than immobilization in internal rotation. Additionally, this review suggests that there is minimal...

This meta-analysis of 6 small sample studies of immobilization in external rotation vs internal rotation for initial dislocation of the shoulder “failed to find a significant difference.”

The CI is 0.42-1.14 with p=0.15. The data are compatible with a pretty large improvement or a very small increase in risk of recurrence.

Do we really need a study with 200+ subjects in each arm (for a reliable RCT), or can this be studied using techniques that would increase power via matching on covariates?

---

<div class="post-metadata">

**Author:** ![ESMD](https://discourse.datamethods.org/letter_avatar_proxy/v4/letter/e/9e8a1a/32.png) [@ESMD](https://discourse.datamethods.org/u/ESMD)\
**Post date:** [November 20, 2022, 11:49pm UTC](https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226/13 "2022-11-20T23:49:34Z")

</div>

Fortunately, I don’t think that someone like me (a family doctor) would ever presume to make treatment recommendations for a patient with metastatic cancer 🙂 So you don’t need to worry about the potential consequence if I were to misinterpret the results of that particular NEJM trial. Nor would I expect that the results of a trial with the limitations you have outlined would necessarily be put into practice by the specialist community (the group in the best position to identify such limitations).

Nobody is arguing that the results of RCTs should automatically be put into practice without deep scrutiny of their methods and examination for potentially fatal flaws, ideally by those with the deepest subject matter expertise. Only if there is agreement by those with both subject matter _and_ methodologic expertise (and sometimes the Venn diagram doesn’t entirely overlap) should a strong recommendation make its way into a Clinical Practice Guideline. In other words, it’s important not to conflate the idea of an efficacy “hierarchy” with the idea that RCTs are “infallible”- as is true with any study design, RCT results always require deep critical analysis.

---

<div class="post-metadata">

**Author:** ![ChristopherTong](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/christophertong/32/2938_2.png) [@ChristopherTong](https://discourse.datamethods.org/u/ChristopherTong)\
**Post date:** [November 21, 2022, 12:41am UTC](https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226/14 "2022-11-21T00:41:22Z")

</div>

Perhaps some additional background reading for this topic could be S. Piantadosi, _Clinical Trials_ 3/e, chapter 2, especially Sec. 2.2.4 “Clinical and Statistical Reasoning Converge in Research” and Sec. 2.3.4, “Experiments Can Be Misunderstood” but also the rest of the chapter.

> **[Clinical Trials: A Methodologic Perspective, 3rd Edition](https://www.wiley.com/en-us/Clinical+Trials%3A+A+Methodologic+Perspective%2C+3rd+Edition-p-9781118959206)**
>
> Presents elements of clinical trial methods that are essential in planning, designing, conducting, analyzing, and interpreting clinical trials with the goal of improving the evidence derived from these important studies This Third Edition builds on...

---

<div class="post-metadata">

**Author:** ![Pavlos\_Msaouel](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/pavlos_msaouel/32/739_2.png) [@Pavlos\_Msaouel](https://discourse.datamethods.org/u/Pavlos_Msaouel)\
**Post date:** [November 21, 2022, 2:07am UTC](https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226/15 "2022-11-21T02:07:48Z")

</div>

Excellent and relevant book on this topic indeed. Section 2.2.4 is also pertinent to the original [random sampling vs random allocation thread](https://discourse.datamethods.org/t/random-sampling-versus-random-allocation-randomization-implications-for-p-value-interpretation/6177).

---

<div class="post-metadata">

**Author:** ![ChristopherTong](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/christophertong/32/2938_2.png) [@ChristopherTong](https://discourse.datamethods.org/u/ChristopherTong)\
**Post date:** [November 21, 2022, 5:59am UTC](https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226/16 "2022-11-21T05:59:03Z")

</div>

Agreed, I should have thought of it ealier!

---

<div class="post-metadata">

**Author:** ![f2harrell](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/f2harrell/32/165_2.png) [@f2harrell](https://discourse.datamethods.org/u/f2harrell)\
**Post date:** [November 21, 2022, 12:21pm UTC](https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226/17 "2022-11-21T12:21:51Z")

</div>

> [@ESMD](#):
>
> In short, the _over-riding_ priority of physicians is to avoid exposing patients to dud treatments.

Wonderful post and discussion following it. On this one statement I’d offer that finding treatments that are truly effective must be at least equal in importance to a physician. A “dud” can be compensated for by switching treatments, but if there are no effective treatments to switch to the patient is out of luck.  
And in oncology I see too much excitement about costly, minimally effective treatments.

---

<div class="post-metadata">

**Author:** ![Pavlos\_Msaouel](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/pavlos_msaouel/32/739_2.png) [@Pavlos\_Msaouel](https://discourse.datamethods.org/u/Pavlos_Msaouel)\
**Post date:** [November 21, 2022, 3:15pm UTC](https://discourse.datamethods.org/t/information-evidence-and-statistics-for-critical-research-appraisal/6226/18 "2022-11-21T15:15:18Z")

</div>

> [@f2harrell](#):
>
> And in oncology I see too much excitement about costly, minimally effective treatments.

Worth digging deeper into this statement: the concept of cost relates to the trade-offs used to make _decisions_. It may help to distinguish between descriptions, inferences, and decisions, e.g., as we describe using oncology examples [here](https://www.mdpi.com/2072-6694/13/11/2741). Often in this forum I focus on statistical inferences, highlighting the role of biological and other mechanisms that could generate the observed data. However, in the statistical methodology literature our team’s contributions are, at least superficially, closest to the Bayesian decision-theoretic framework that converts inferences into decisions.

Along these lines, we recently published methodology to personalize treatment decisions in oncology based on covariate-specific utility functions from RCT datasets [here](https://rss.onlinelibrary.wiley.com/doi/10.1111/rssc.12582). In a conceptually similar manner, we also recently developed a [Bayesian treatment selection approach](https://onlinelibrary.wiley.com/doi/10.1111/biom.13738) that capitalizes on the advantages of random allocation and concurrent control.
