Royal Statistical Society Discussion on Randomization vs Model Based Inference July 1, 2026

David, this is an excellent question. The core issue is that fruit type does not function as a simple baseline adjustment within a shared treatment-response system. Apples and oranges are governed by different causal structures, each with its own treatment effect and its own covariate vector, with distinct predictors and weights. When mixed together in one trial these form what I call the covariate mass at E3.

This was my concern with Professor Senn’s otherwise masterful presentation and, more generally, with discussions of covariates and adjustment in RCTs. They typically treat the enrollment gate as fixed rather than as part of the causal structure. Yet, as the infographic illustrates, the gate determines whether there is one causal system or many. If the gate selects many, the gate, in combination with the population sampled, determines which causal systems are combined and this causal system mixture is not likely to be reproduced thereby creating the cause mixture paradox.

The gate therefore determines which layer the covariates inhabit and for that reason cannot be left out of the discussion.

1 Like

I conceded that those considerations need to be “half of the discussion” but no further. Adjusting for generalized risk factors such as age and comorbidities helps a great deal with no consideration of pathways.

1 Like

If you are speaking of 3 randomized experiments, I see a lot of inefficiencies in that design.

No. I’m speaking of 3 balanced, or in more recent language, minimized designs, as Gossett argued for in is last paper in Biometrika.

Student, (1938). Comparison Between Balanced and Random Arrangements of Field Plots Comparison Between Balanced and Random Arrangements of Field Plots | JSTOR

In that same issue, Neyman and Pearson elaborated on Gossett’s points, as he had passed before completing the paper.

Note on Some Points in “Student’s” Paper on “Comparison Between Balanced and Random Arrangements of Field Plots” Note on Some Points in "Student's" Paper on "Comparison Between Balanced and Random Arrangements of Field Plots" | JSTOR

Much of the language in the paper involves technical details related to integrating the experimental plan with farming techniques, so it isn’t always easy to see what is going on. i have had to read it a few times in order to really comprehend the dispute.

The entire debate between Gossett and Fisher about bias was Fisher’s insistence on comparing 1 random sample to 1 balanced sample. Gossett argued that maybe 1 balanced sample might be biased in estimating the mean (ie. the grand mean in a series of trials in a set of different farms), due to local factors, but a number of small samples would not be biased for the true purpose of the experiment – to test the relevant treatment effect. Gossett called this difference in means actual error of the experiment.

Gossett fit a normal model when it wasn’t precisely correct, as model variance was an overestimate actual variance. He termed this the calculated error of the experiment. Balanced designs result in data sets that have thinner tails than would be expected under the normal model. This was also mentioned in Kallus’s RSS paper on balanced designs and why he used a bootstrap procedure to evaluate his balanced Kernel Allocation method.

He argues that balanced experiments has more power when it is needed most – large, economically relevant effects. A more charitable interpretation was Gossett’s use of the normal model for a balanced experiment analogous to using a small, but unspecified interval null hypothesis.

Gossett’s error was not a major one in terms of making decisions. But I see that is a moot point when we can resample and compute valid variance estimates regardless of the difference in means.

Gossett’s repeated experimentation with small samples was a simultaneous test of both the hypothesis in question and the measurement process itself. Excessive variance indicated your knowledge of the relevant factors was incomplete. Scientific creativity would be needed to solve the problem, but at least the statistics indicated there was a problem.

Injecting a source of randomness when other sources are not controlled will make it extremely difficult, if not impossible, to figure out what is going on. This is how I interpret the sepsis trials that have been mentioned by Dr. Lynn.

1 Like

I would love to see @Stephen comment on that. He may identify major flaws or unnecessary complexity in such balanced designs.

4 Likes

Senn has made the following comments on balanced designs. There are a number of algorithms to do it, but I’m arguing for a more general claim that experimenter imposed balance deserves more consideration.

  1. Senn’s comment on Altman’s Allocation by Minimization article in BJM in 2005
  2. A sadly unproductive dialog with economist and Gossett scholar Stephen Ziliak over at the Lancet in 2010: Zilliak’s article:
    Ziliak, S. T. (2010). The Validus Medicus and a new gold standard. The Lancet, 376(9738), 324-325. (link)
    Senn, S. (2010). Significant errors. The Lancet, 376(9750), 1390-1391. (pdf)
    Ziliak, S. T. (2010). Significant errors–Author’s reply. The Lancet, 376(9750), 1391.

Ziliak’s reply follow’s Senn’s rebuttal in the second PDF I linked to.

Senn elaborated on the use of alternatives to randomization in this Statistics In Medicine article:

Senn, S., Anisimov, V. V., & Fedorov, V. V. (2010). Comparisons of minimization and Atkinson’s algorithm. Statistics in Medicine, 29(7‐8), 721-730.

A more insightful commentary by Ziliak can be found on Gelman’s blog from 2014 Post 1, Post 2,

There have been new balanced allocation mechanisms designed since that article. Some of the problems with balanced designs involve having to dichotomize some variables. Taves goes into details on prospectively combating selection bias in a trial using minimization in this 2017 paper:’

Taves, D. R. (2017). Flexible Minimization: Synergistic Solution for Selection Bias. In Randomization, Masking, and Allocation Concealment (pp. 229-241). Chapman and Hall/CRC.
(pdf)

He also briefly discusses the Berger-Exner test for detecting selection bias in these types of designs.

There is a page where a section on the FDA website that linked to this article on “information adaptive” designs. Don’t “information adaptive” designs move us progressively closer to Bayesian Optimal designs, which explicitly do not require randomization?

There remains significant merit in Senn’s insights into experimental design. But any robust protocol should incorporate what has also been learned by engineers and computer scientists in designing fault tolerant systems.

The current evaluation methods rely substantially on information provided by self-interested actors, with no external validation. Feynmann indicated failure to independently replicate experiments was a feature of what he called Cargo Cult Science. These replications should not take a large sample to conduct. Randomization encourages this desire to avoid replication attempts, and does not protect against “unscrupulous actors” as the ADVOCATE trial I linked to above, demonstrates.

Bayesians who have followed the math have always been suspicious of randomization. Savage had this to say on the conflict between personal probability and randomization:

Savage, L. J. (1961). The foundations of statistics reconsidered. In Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability, Volume 1: Contributions to the Theory of Statistics (Vol. 4, pp. 575-587). University of California Press.

The theory of personal probability must be explored with circumspection and
imagination. For example, applying the theory naively one quickly comes to
the conclusion that randomization is without value for statistics. This conclu-
sion does not sound right; and it is not right. Closer examination of the road to
this untenable conclusion does lead to new insights into the role and limita-
tions of randomization but does by no means deprive randomization of its
important function in statistics.

The relatively new discipline of cryptography sheds light on Savage’s question. Randomization is admissible when you control the allocation mechanism.

Randomization will not serve the purpose of compelling a rational group of scientists into accepting the specification that the treatment and control are exchangeable on all factors aside from treatment, as Rubin and others have argued in a number of papers. This wouldn’t be known until after the 1980’s, when cryptography became a concern of civilian business and the public and not merely limited to military and intelligence interests. .

Worth reading:
Bernardo, J. M. (1996). The concept of exchangeability and its applications. Far East Journal of Mathematical Sciences, 4, 111-122. (pdf)

From the introduction:

The general concept of exchangeability allows the more flexible modelling of most experimental setups. The representation theorems for exchangeable sequences of random variables establish that any coherent analysis of the information thus modelled requires the specification of a joint probability distribution on all the parameters involved, hence forcing a Bayesian approach. The concept of partial exchangeability provides a further refinement, by permitting appropriate modelling of related experimental setups, leading to coherent information integration by means of so-called hierarchical models.

Wouldn’t a series of well controlled, ie. 3 from the sponsor, then 3 from independent replications, provide the regulator with ample information at much less cost than what is done now, with more credibility? I think so.

1 Like

That would actually create a complete mess.

You provided some excellent references of which I was not aware, and did a great job in summarizing. However I did not believe that the small minority of Bayesians who don’t believe in randomization should carry much weight, and I think the literature still contains many misunderstandings about randomization and still also puts far too much emphasis on “generalizability”. In a nutshell, randomization makes it unnecessary to get much of the data model “right”, allows for blinding, and requires fully prospective data collection and patient management under protocol. It allows you to virtually rule out alternate explanations. Bad actors can still ruin things but without randomization the number of opportunities is increased by an order of magnitude.

Adaptive trials definitely need randomization. They just don’t need a constant probability of treatment assignment.

Minimization and other covariate balancing techniques are nice for (1) avoiding disasters and (2) improving variance of treatment effects by keeping covariates orthogonal to treatment but these advantages are minor in the grand scheme of things.

2 Likes

Professor, I agree with your emphasis on replication and scientific integrity, but I would frame my concern somewhat differently.

My primary criticism is not that measurements are too noisy or that randomization, even when combined with covariate adjustment, is fundamentally inadequate for achieving balance (although I will show when that IS true). Rather, the problem I raised is that the causal target is often ill-defined before randomization ever occurs.

In the 1980s, critical care syndrome RCTs deviated from Hills methodology and acquired enrollment gates that are disease- and mechanism-agnostic. Participants are selected by prognostic thresholds (e.g., Sepsis-3, Berlin ARDS, AKI) rather than by a disease or causal mechanism. The trial question became essentially centrally controlled. The consensus gate produces a synthetic data generating process (SDGP) which is unique to the gate and population sampled. Randomization then estimates the average treatment effect within that unique constructed disease mixture (the SDGP).

The key point is that the covariate probability distribution (E2) exists within each underlying disease, whereas the disease-mixture distribution (E3) sits one layer above it. E3 absorbs the covariate structure from all of the constituent E2 distributions. Consequently, covariate adjustment cannot recover a coherent biological estimand once heterogeneous diseases have been combined at the enrollment gate (although Professor Harrell makes an excellent point about near universally applicable adjustments, like age).

Still, regardless of adjustment, the resulting estimate may be internally valid for that particular mixture, yet fail to transport or reproduce the treatment effect polarity when the disease mixture changes with the next trial despite the same setting and gate.That is the cause-mixture paradox.

Here we see the math is not independent of the social structure of the discipline, it is integrated with it. (Which is why it is quite courageous to engage my critique in the way you have.) When this social to trial structure integration occurs this is pathological science on its face. This is true because a social structure is not self corrective so that when the social structure develops an inevitable pathology, then the trials become pathological.

Without the pathology of social integration with trial structure, independent groups would challenge whether the enrollment gate represents a coherent biological system and in the alternative whether the mixture can be reproduced in the population during transport. These would be investigated in the interest of assuring the integrity of scientific self-correction rather than seeking to confirm a numerical treatment effect responsive to social expediency.

That is structured failue mode detection which is a fundamentally different form of error detection then replicating the trial. For example, in corticosteroid trials of community-acquired pneumonia and sepsis, the 39 year effort (first massive trial in 1987 n 382. 19 centers) has largely been an unsuccessful attempt to consistently reproduce a positive numerical treatment effect when those few that were generated and reversed may simply have reflected the chance occurrence of a favorable steroid-responsive disease mixture as predicted by the math of the cause mixture paradox.

My proposal is therefore to make the causal structure, especially the enrollment gate, an explicit object of scientific scrutiny free of social control. My critique of the Society presentation was that it was incomplete, especially in light of the massive failure of RCT based ventilator guideline transport early in the COVID epidemic. We have to address the real world failures of RCT based EBM as it is presently applied rather than continue to debate the adjustment, minimization, etc. So my comment was focused on the “gate to covariate adjustment dependency” and the correction of ongoing related RCT pathologies which @Stephen did not address.

Until we address that, randomization, replication, and increasingly sophisticated statistical methods all remain vulnerable to the same structural design error.

I have some questions for the practising applied statisticians in this forum:

What proportion of students graduating with a PhD in statistics today would read the text of Dr.Senn’s talk (the focus of the original post in this thread) and be able to say that they understand it deeply after a thorough reading ?

My hazy impression (as someone untrained in statistics, who therefore grasped little of his talk) is that Dr.Senn speaks a different language than most statisticians. I was unable to put my finger on the reason for this impression, until fairly recently. It wasn’t until I read some of Dennis Lendrem’s posts that I became aware that Senn’s training obviously focused heavily on “Design of Experiments.” Since Senn’s style of explanation seems somewhat unique (?), I wonder whether statistical training with this particular focus is somewhat rare (?)

Training in DOE seems to focus on learning principled ways to answer scientific questions. The approach is pre-planned and step-wise and seems to be much more routine in the physical sciences (?engineering, physics) than in the biological sciences (?) Scientists say: “We will plan a series of experiments, with the ultimate goal of answering this specific question:…If our first experiment shows X, then we will be able to rule out Y. In this case, our next experiment, will be designed as follows…” (?)

What proportion of graduate programs in statistics include Design of Experiments as core course requirements (?) And, if PhD applied statistics students are not required to take courses in DOE, what alternative knowledge base do they fall back on when designing clinical trials (?) Are there important, specific knowledge gaps that might exist without dedicated training in DOE (?)

Apologies if these questions are terribly naive. Physicians, almost certainly will never receive training in DOE (except, perhaps, for the very rare research clinician (?)). Therefore, if applied statisticians speak the language of DOE when collaborating with physicians in planning clinical trials, how can they possibly communicate effectively with each other (?) Note that I’m mostly sympathizing with applied statisticians here, who might struggle mightily to gain co-operation from clinical colleagues regarding certain study design features…

3 Likes

DOE is almost never taught in biostatistics programs in the past 25 years. It is to our great detriment. DOE applies to every field and is ironically being used the most by internet marketers and web site developers nowadays. If DOE were taught to more clinical trialists much would be gained. I learned what I know by doing a consulting project with Corning Electronics where they used various factorial, fractional factorial, and central composite designs for making better capacitors.

You should look at Stephen’s Statistical Issues in Drug Development which is the bible for statistics for drug development and clinical trials in general and will expose you to many facets of the world’s leading pharmaceutical statistician.

Stephen has never told me this but I’ll bet he and I share a role model: Sherlock Holmes. Stephen excels at uncovering hidden subtleties that change the interpretation of statistical results, and at uncovering statistical mysteries, finding alternate explanations for many phenomena.

Even when DOE principles can’t be fully implemented in certain human studies, it would be a highly worthwhile exercise to design a perfect clinical trial (especially one that uncovers mechanisms as well as improving patient outcomes) then see how close to that design you can implement.

6 Likes

Thanks. These are just random thoughts and the questions are mostly rhetorical.

Your response begs the question: Why is DOE not taught, if it is so valuable for deep understanding of clinical trial design and analysis? This doesn’t seem to make sense (?)

This is my strong impression. But, at the risk of sounding morbid, what happens to all this “corporate memory” when the (?few) biostatisticians trained in DOE are no longer around ? Does much of Dr.Senn’s frustration over the persistence of unsound statistical practices stem from the fact that so few biostatisticians/clinicians today have the training in DOE that’s needed to deeply internalize the problems with these practices?

I would imagine that DOE is an essential feature of the training of researchers in other fields e.g., experimental physics, industrial engineering, agricultural science (?) But, presumably, it’s the subject matter experts in these fields who are, themselves, learning the applied statistics they need to apply in their daily work (?)

If the biostatisticians being trained today are not being taught DOE, what type of training fills the gap to ensure sound trial design/analysis (?) If they don’t have training in DOE to fall back on, how will biostatisticians dissuade naive clinical researchers who insist on unsound trial design/analysis practices? Would RNCTs, NNT, responder analysis, and dichotomization have gained footholds in clinical medicine if more biostatisticians had understood DOE well enough to push back against clinicians with unreasonable aspirations?

It seems like there might be two distinct “streams” of training for biostatisticians that can improve RCT design and analysis: 1) causal inference epidemiological techniques, as applied to the design of RCTs (e.g., emphasizing concepts like non-identifiability of individual causal effects in parallel group trials); and/or 2) training in DOE. Has training in 1) supplanted training in 2) (?) Would students trained only in 1) be at risk of making serious design/analysis mistakes (and vice versa) (?)

It feels like a lot of current bad practices (examples mentioned in the preceding paragraph) might have slipped into routine use during a period in time (?late 1980s through to present day) when there were too few biostatisticians with training in either DOE OR causal inference epidemiology to push back, in a co-ordinated manner, against their emergence.

2 Likes

DOE wouldn’t have solved all of those messes but it would have made a huge positive impact. I too am worried about the next generation. Over the last 10 or so years biostatisticians have embraced high-dimensional data and causal inference without understanding DOE that would have explained limitations of both. One of the greatest beauties of DOE is that sometimes you can design a definitive experiment with minimal sample size, thus saving a LOT of time and money. For example there are occasions where you can do a 2^5 factorial design with 5 replications at just one of the combinations, for a total of 37 observations, to understand the importance and interplay of 5 independent variables.

Perhaps more important are the lessons from DOE related to designing to cancel artefacts, including using randomization in studies where you don’t think you needed it such as mass spectrometry for biomarker development. There’s also a lot of other measurement issues related to DOE.

3 Likes

A quarter century ago DOE was an elective in both statistics and industrial engineering at Purdue, where I studied (they did not have a biostats program back then). DOE was/is not systematically taught to physicists, despite Youden’s vocal advocacy for it. There are deep reasons for that (which are way off topic for this thread), but some physicists doing applied work (eg, semiconductors, nuclear weapons) are conversant in DOE.

Returning to biostatistics, DOE shows up in every job ad for nonclinical biostatistics that I’ve seen in the last 14 months. However most academic biostats programs are training for clinical, not nonclinical, biostatistics.

6 Likes

Interesting…and odd…I reiterate: If DOE can provide such useful insights, why is it not being taught to clinical biostatisticians? Given Stephen Senn’s widely recognized expertise, the reason can’t be that the topic is somehow considered ill-suited for application to the field. Is it because learning the topic thoroughly would require too much dedicated study, crowding out curriculum time that is (today) allotted for other topics (e.g., causal inference epidemiology?). Or is it because so few practising statisticians are expert enough in the topic to teach it? Or is it because the topic is so hard to understand that demand for the courses was low (and supply therefore dried up)?

3 Likes

DOE is easy to teach and its methods are easily retained by students. There’s not much math. My hunch is the DOE was crowded out by more data analysis-heavy courses related to survival analysis, longitudinal data, and especially in the past 25 years, high-dimensional methods for the latest and greatest biomedical tech, starting with gene microarrays, then on to other “omics” data.

3 Likes

This is profoundly insightful question and I would add …why is it not taught to trialists or clinicians?

I discovered how difficult it was to learn DOE when doing failure mode analysis of critical care RCT. I had to fall back on this Bradford Hill study and my clinical knowledge of TB. Reading this study within the context of its date (1948) is absolutely fascinating. It’s like looking at Michelangelo’s David.

https://doi.org/10.1136/bmj.2.4582.769

This is where I recommend starting. Then follow the history. One will quickly see the emergence of a new species RCT which bypass (perhaps eclipse depending on your perspective) Hill’s teachings.

Teaching casual structure of Hills model is inconsistent with the present view of a single gold standard entity (“The” RCT). Perhaps this is a reason DOE education quickly converts to the complexity of @Stephen ‘s focus on the mathematics of estimators, drop outs, intention to treat, etc. All important but when presented without a basic DOE causal structural and transportable foundation the work is incomplete.