Royal Statistical Society Discussion on Randomization vs Model Based Inference July 1, 2026

@Stephen announced on LinkedIn he is presenting his paper on randomization vs model based inference before the Royal Statistical Society on July 1, 2026. I thought participants here would be interested, as it will also be presented online.

Event info and a link to the preprint can be found here:

Next Discussion Meeting
‘Randomisation in Clinical Trials: Is a Pragmatic Compromise between Randomisation- and
Model-Based Inference Possible?’ by Stephen Senn

Broadway House, Westminster, London and online
Broadway House
Tothill St
Westminster
London
SW1H 9NQ>
Wednesday, 1 July 2026
Time 3pm to 5pm (UK time)
Post-meeting reception 5pm to 6pm

Paper: ‘Randomisation in Clinical Trials: Is a Pragmatic Compromise between Randomisation- and Model-Based Inference Possible?’
Author: Stephen Senn, Consultant Statistician University of Sheffield, Medical University of Vienna and University of St. Andrews

A link to the paper can be found at:

7 Likes

Thanks so much for posting this, Robert! I just finished a first reading, and very much liked both the fresh argument on change-from-baseline and the rough handling of Howson & Urbach, whose book struck me as flippant and dishonest in its treatment of Popper.

Senn’s emphasis on DoE always makes his reasoning challenging for me to follow, but I trust this is a healthy challenge like a high-fiber diet.

The distinction between the ‘LM’ and ‘MV’ approaches (p.6) remains unclear to me. I’m thinking the latter would be characterized by the presence of latent variables providing for information flows that improve efficiency? Would anyone have a plain example of contrasting LM and MV approaches to the same problem?

3 Likes

This is indeed a brilliant article worth multiple rereads. For the linear-Gaussian case that @Stephen wisely focuses on to avoid the whole non-collapsibility hoopla, I think we can distinguish (dichotomize :wink: ) between the design-based versus the model-based approaches, whereby model-based approaches use a superpopulation to provide aleatory probability interpretations.

In Senn’s terms, the randomization distribution (RD) approach is purely design-based but less practical. The multivariate (MV) and linear model (LM) approaches both effectively invoke a superpopulation. But the purely LM view treats the baseline prognostic covariates as fixed and does not justify omitting them. The only random object is the outcome. Conversely, the MV view treats also the baseline prognostic covariates as random and jointly distributed with the outcome Y. This is the old Galton–(Karl) Pearson correlation worldview where everything is jointly normal (as recounted in the Aldrich 2005 paper cited by Senn) and the marginal of a multivariate normal remains a perfectly well-specified normal. Thus, in the MV view, conditioning on a subset of random baseline prognostic covariates is legitimate. That is what gives conceptual permission not to measure everything. LM is the worldview under which omitting an observed prognostic covariate is outright misspecification (the very discomfort that drives Senn to MV).

As an aside, RD is an example of frequentism-as-model (FAM) where the repetition (frequency) is not derived from a superpopulation but re-randomization (a real and performable operation). Thus, RD is the one framework where the frequentist probability nearly exists qua design. As also noted here, particularly by @f2harrell, one may instead dispense with the superpopulation clause by putting epistemic probabilities on the parameters via the Epistemic-Probability-as-Model (EPaM) view.

4 Likes

Many thanks, Pavlos! I was a bit worried I’d have to read Aldrich, and it seems in fact I do; Fisher’s idea of ancillarity of X would seem to be at the core of LM-vs-MV. Your Hennig reference looks extremely interesting, too, and is also queued up.

2 Likes

Since @Stephen is presenting his paper tomorrow, this looks like a good thread to post his interview on the Scott Berry podcast discussing statistics and clinical trials.

5 Likes

There is a wealth of knowledge shared in this podcast. I highly recommend it to anyone interested in clinical trial design.

5 Likes

The discussion was a bit of a let down to be honest, since no one disagreed with him :frowning: . Excellent paper nonetheless.

3 Likes

Well since you asked for it, I will provide a question.

Senn advocates adjustment for prognostic covariates, my question is, prognostic for what? It seems that the concept of a prognostic covariate cannot be separated from the enrollment gate. The answer depends on the population constructed by the gate. If the gate selects a coherent causal system, the prognostic model has a clear biological referent. If the gate constructs a cause-agnostic disease mixture, then the prognostic model pertains to that potentially non-reproducible mixture rather than to any individual disease represented within it.

So the first question the biology trained reader may have is; “I see the math works but is it biologically meaningful, indeed, is it a blind assumption-based mathematical technique when performed without connecting the math to the gate. Does the joint distribution being modeled correspond to a biological data-generating process for which the covariates are prognostic or are they prognostic for an artifactual mixture distribution induced by a cause-agnostic enrollment gate (S=1)”.

In that case, the regression model estimates a property of the disease mixture-defined estimand (E₃), not of any individual disease-specific estimand (E₂).

2 Likes

A fair rebuttal to Senn’s paper deserves much more detail and citations than I can provide here. But a quick sketch of my thoughts, that builds upon the observations of @llynn are as follows:

  1. Just a few days before Senn’s talk, the New England Journal of Medicine retracted a Phase 3 study which the author notes, is extraordinary for NEJM. NEJM Retracts Avacopan - by Mike Putman

This is not just a story about avacopan. It is a story about how the normal systems of sponsor oversight, contract research organization (CRO) data handling, peer review, regulatory approval, and post-publication scrutiny failed to detect a shocking case of data manipulation.

Throughout a number of papers, Senn makes the argument that randomization protects the trial from “unscrupulous actors.” More detail can be provided in his paper Fisher’s Game with the Devil. The arguments in the paper are correct if you are either the agent, or trust the agent in control of the allocation process.

The growing number of case studies on scientific fraud should provide evidence that this faith in randomization is misplaced. Would randomization be credible if the Devil, in Senn’s paper, had control of the allocation process?

The problem [of deceptive research reports] is larger than one might imagine and could be as high as 20% of publications.

He makes useful points regarding randomization (from the perspective of the agent in control of the allocation mechanism running an honest experiment), but as in all of his writings on this topic, it makes a critical assumption that I no longer believe holds: the agent in control of the allocation process that using randomization is either trustworthy (will not attempt to deceive), or strategies for cheating are too costly to implement, and are easily detected. This makes it problematic for modern science that must rely upon the reports of others.

  1. If we wish to discount the lack of protection from actors with the intent to mislead, that brings us to Lawrence’s complaint – randomization in the context of clinical trials in sepsis, have not produced much in useful results.

The fundamental problem in this context is that the measurement process, the definition of the condition under study depends on other hidden assumptions (that may be wrong) in order for the assumption of groups created via a randomization process to be exchangeable for statistical purposes.

Principled Bayesians used to complain about randomization and randomized decision rules. As was pointed out in another thread, Jaynes wrote::

Of course, Fisher’s randomized planting methods – which we think to be not actually wrong, but hopelessly inefficient at information handling – were not reproduced by Jeffreys, nor would he wish to.

It is not disputed that you can control more prognostic covariates in a balanced, controlled trial at a smaller sample size than you can with randomization. If you budget for a large trial, you can perform some error detection by running 3 smaller controlled, but balanced trials, and then combine them via meta-analysis at the end.

You can actually test if your 3 samples are reasonably homogeneous using one of the tests of distributional equality. The most general is the Kolmogorov-Smirnov. Low p values indicate heterogeneity and information loss, and the effective sample size can be adjusted downward from the relevant experiment.

Of course, there are going to be logistical challenges and various economic and clinical considerations to balance. I’m not sure how to handle safety monitoring, for example. I’d be inclined to let each experiment assess that independently, but that might not be economical nor ethical, and centralized safety monitoring might be better. I don’t think you would need to present much evidence to me that centralized safety monitoring along with decentralized efficacy monitoring, is preferable. This would require complex simulations, though.

But it is now time to do the hard thinking and try to do better than randomization.

Related thread:

3 Likes

The Avacopan study retraction shows that unscrupulous people can be involved in RCTs and do a lot of harm. But major scandals involving approved drugs, stemming from the actions of a small number of people with nefarious motives (i.e., corporate “coverups”), have, historically, come to light only very infrequently- I’d estimate approximately every 5-15 years. Even though they are undoubtedly heinous because they can affect a lot of people, such scandals are distinctly UNcommon, given the huge number of people involved in RCT oversight.

RCTs can be defective in their design, conduct, or analysis, just like any other study design. Authors of all types of studies can have conflicts of interest (intellectual, reputational, or financial). Of all study designs, RCTs tend to have the most eyeballs (cheers to David :clinking_glasses:) focused on their conduct and analysis. As such, it’s reasonable to assume that a higher proportion of bad (or incompetent) actors will be deterred or ferreted out before a drug ever gets approved.

From this point on, I’m talking about bad actors (i.e., those with overtly nefarious motives). If you’re suggesting that we should routinely disbelieve the results of RCTs (i.e., “if we found bad actors once, what’s to say they’re not involved in MOST trials”?), then what, exactly, is your proposal for a more trustworthy process for assessing therapeutic efficacy ? Any suggestion that routine disbelief in RCT results should cause us to place more faith in observational efficacy comparisons than we have done historically, would be a non sequitur. In other words, the ever-present spectre of egregious corporate malfeasance in no way supports an argument that study designs other than RCTs should be the preferred method for assessing therapeutic efficacy.

If we take pharmaceutical companies out of the picture, who, exactly, do you propose should finance drug development instead? Academic clinicians with tenure on the line? They would NEVER be conflicted, after all…? Taxpayers, through their funding of national research bodies? Do you really think that these are credible options for financing the development and testing of all the drugs we need to treat our patients (?)

Let’s think this through, with informed consideration of how drug development and approval actually works:

  • Many important basic science discoveries are made by nationally-funded research organizations (at least in the U.S.) or universities. These organizations are subsidized by taxpayers;
  • Pharmaceutical companies exploit these basic science discoveries for the purpose of developing new therapies for diseases;
  • Using their own funding, generated from profits derived from sales of other drugs, drug companies first test potential therapies in the lab, then in non-human animals;
  • With regulatory agency oversight, a small subset of these therapies are then tested in a small number of humans to rule out unacceptable toxicity and determine the appropriate dose;
  • With regulatory agency oversight, a further subset of therapies progress to larger pivotal trials;
  • BOTH pharmaceutical companies AND regulators analyze the results of pivotal trials;
  • Regulators decide whether or not to approve the therapy.

Given that funding for U.S. national research bodies and universities has been decimated (to the point that their ability to conduct basic research is already seriously compromised), do you really think that governments (with taxpayer money) are going to step in to perform all the tasks currently being performed by pharmaceutical companies? So, if pharmaceutical companies don’t take the initiative to explore potential new therapies and underwrite their stepwise, methodical study (including paying for the necessary manpower), who will?

In the real world, where we need to make real decisions, we usually need to operate under the assumption that MOST people are acting in good faith, MOST of the time. The only alternative is to adopt a paralyzingly cynical worldview in which we assume that MOST people are actually acting in BAD faith (e.g., anyone working for pharmaceutical companies). We would literally get nowhere in medicine if we adopted this outlook.

As I see it, the only workable approach to drug regulation is “trust but verify.” Very occasionally, we’ll get burned by bad actors, in spite of all our safeguards- Avacopan is a case in point. So regulators certainly need to remain vigilant. But the “very occasionally” part is key.

3 Likes

Bad actors come in two forms.

The rare… the nefarious, and the more common…the uninformed or incurious, such as those who, despite decades of reversal, diligently recycle pathological RCT design. (Which may be meticulously embellished by “prognostic” covariate adjustment.)

1 Like

Older people, people living with diabetes, those with heart failure will always have worse outcomes for nearly everything than those without. Adjust for things like these please

3 Likes

Whoa:

Welcome back, Christos! Great to see you and your bison cappuccino again!

2 Likes

It’s a lot simpler than this. Adjust for baseline factors strongly suspected or known to predict the outcome. And if you adjust for a couple of factors that turn out to be non-prognostic, no harm done.

A big mistake is to fail to explain easily explainable outcome variation by not conditioning on covariates, which boosts power and makes the model fit better for free.

1 Like

Absolutely and thank you as I have learned this from you.

But my point, which was not addressed in Professor Senn’s paper is that the type of gate to matters when considering adjustment. For the adjustment to be valid the gate must select a biologic data generating process not a synthetic data generating process.

All RCT are not the same species and some end at E2 with covariate vectors and some end at E3 with a covariate mass of unrelated covariate vectors . I have found this quite hard to teach because of the “averaging similarity” of the E2 and E3 layers. So I am resorting to infographics, Here is my latest infographic which I plan to present at an upcoming conference. I hope there will be criticism/comments.

3 Likes

Why not simply include a fruit_type factor (taking values orange or apple) among the predictors?

1 Like

Perhaps @llynn will correct me if I’m wrong, but his complaint seems to be: how do you discover what variables to condition on when the measurement process is so noisy?

My complaint is that randomization can work if other parts of the model have very low probabilities of being in error. This isn’t the case in sepsis research, as Lawrence has overwhelmingly demonstrated

In the context of sepsis research: the RCT methodology didn’t aid his community of medical scientists from discovering a crucial flaw in their assumptions. This lead to 30+ years of wasted research resources.

This wasn’t clear in Fisher’s time, but any protocol that is deemed “scientific” needs to be:

  1. decentralized ie. no central authority that determines “truth”. Consensus determines truth.
  2. resistant to misleading reports, whether through honest error, or through intentional manipulation.

All of these properties are implied in Feynman’s classic Cal Tech speech, where he introduced the notion of Cargo Cult Science. The essence of science, according to Feynman is:

  1. Don’t fool others, have scientific integrity.
  2. Don’t fool yourself, and never forget you are very easy to fool.

There are a number of methodological guidelines he provides, that statistical recommendations (ie. randomization) are in conflict with – especially the need for repetition in order to validate a claim.

Other kinds of errors are more characteristic of poor science. When I was at Cornell. I often talked to the people in the psychology department. One of the students told me she wanted to do an experiment that went something like this—I don’t remember it in detail, but it had been found by others that under certain circumstances, X, rats did something, A. She was curious as to whether, if she changed the circumstances to Y, they would still do, A. So her proposal was to do the experiment under circumstances Y and see if they still did A.

I explained to her that it was necessary first to repeat in her laboratory the experiment of the other person—to do it under condition X to see if she could also get result A—and then change to Y and see if A changed. Then she would know that the real difference was the thing she thought she had under control.

She was very delighted with this new idea, and went to her professor. And his reply was, no, you cannot do that, because the experiment has already been done and you would be wasting time. This was in about 1935 or so, and it seems to have been the general policy then to not try to repeat psychological experiments, but only to change the conditions and see what happens.

These intuitive notions of Feynman regarding honest scientific process were formalized in the computer science literature, starting with the discussion by:

Lamport, L.; Shostak, R.; Pease, M. (1982). “The Byzantine Generals Problem” (PDF). ACM Transactions on Programming Languages and Systems. 4 (3): 382–401.

No communications engineer would today implement a protocol that does not at least have some form of error detection or correction if possible. In critical systems that are decentralized, resistance to parts of the system to arbitrary errors in other parts is known as being tolerant to
Byzantine Fault.

Byzantine fault tolerance is a family of problems, with the answer contingent upon how much computational power you grant to the agents submitting misleading signals/reports.

A system is Byzantine Fault Tolerant if > 2/3 of the signals/reports are reliable.

(Agents in this context need not be human beings; they could be sensors or computer processors).

The question I’m asking myself currently: how can this finding be adapted to the clinical trial context, without requiring huge sample sizes?

I believe this total cost of experimentation (initial claim + validation) can done with sample sizes of < 2n (with n being the maximum size of an initial randomized trial) with some hard thinking, as E.T. Jaynes recommended. But that will require rethinking the role of randomization among a community of scientists.

Who could object to a process that provides Byzantine fault resistance for less than the cost of 2 RCTs?

2 Likes

There’s a lot of value in that way of thinking; I just wanted to point out that even silly factors that are still true baselines (can’t be caused by treatment, etc.) can explain variance in Y and increase power and precision.

1 Like

Isn’t Senn’s point rather that if you want to have any hope of protecting against nefarious actors randomisation is necessary (but not necessarily sufficient), because without randomisation no concurrent control, and without concurrent control no blinding.

1 Like

That is confusing controlled experiments with observational studies. I’m discussing explicitly balanced allocation strategies along the lines of Gossett, Donald Taves, and Simon and Pocock.

I’m arguing it would be preferable to do 3 independent, balanced designs under the same protocol of at most 1/3 the maximal random sample size, and then combine the results using meta-analytic techniques (ie. likelihood functions).

To validate, have another party or set of parties run another set of 3 balanced experiments. If you do all of this on a sequential basis, you are likely to need much less than 2n (two full sized randomized trials), with a drastically lower chance of being mislead.

In the language of Seidenfeld and Kadane, randomization might be useful (but not essential) for “experiments to learn”, but I fail to see how they are useful for “experiments to prove”, especially if we are looking for Byzantine Fault Tolerance. The required sample sizes for randomization to be valid are too large, and prevent redundancy in passing messages (ie. experimental repetition).

See also:
Seidenfeld, T. (1990). Randomization in a Bayesian perspective. Journal of statistical planning and inference.

The Bayesian rule described in this source adds more randomness as the sample size gets larger.

Atkinson AC. (2014). Selecting a Biased-Coin Design, Statistical Science, Statist. Sci. 29(1), 144-163