A fair rebuttal to Senn’s paper deserves much more detail and citations than I can provide here. But a quick sketch of my thoughts, that builds upon the observations of @llynn are as follows:
- Just a few days before Senn’s talk, the New England Journal of Medicine retracted a Phase 3 study which the author notes, is extraordinary for NEJM. NEJM Retracts Avacopan - by Mike Putman
This is not just a story about avacopan. It is a story about how the normal systems of sponsor oversight, contract research organization (CRO) data handling, peer review, regulatory approval, and post-publication scrutiny failed to detect a shocking case of data manipulation.
Throughout a number of papers, Senn makes the argument that randomization protects the trial from “unscrupulous actors.” More detail can be provided in his paper Fisher’s Game with the Devil. The arguments in the paper are correct if you are either the agent, or trust the agent in control of the allocation process.
The growing number of case studies on scientific fraud should provide evidence that this faith in randomization is misplaced. Would randomization be credible if the Devil, in Senn’s paper, had control of the allocation process?
The problem [of deceptive research reports] is larger than one might imagine and could be as high as 20% of publications.
He makes useful points regarding randomization (from the perspective of the agent in control of the allocation mechanism running an honest experiment), but as in all of his writings on this topic, it makes a critical assumption that I no longer believe holds: the agent in control of the allocation process that using randomization is either trustworthy (will not attempt to deceive), or strategies for cheating are too costly to implement, and are easily detected. This makes it problematic for modern science that must rely upon the reports of others.
- If we wish to discount the lack of protection from actors with the intent to mislead, that brings us to Lawrence’s complaint – randomization in the context of clinical trials in sepsis, have not produced much in useful results.
The fundamental problem in this context is that the measurement process, the definition of the condition under study depends on other hidden assumptions (that may be wrong) in order for the assumption of groups created via a randomization process to be exchangeable for statistical purposes.
Principled Bayesians used to complain about randomization and randomized decision rules. As was pointed out in another thread, Jaynes wrote::
Of course, Fisher’s randomized planting methods – which we think to be not actually wrong, but hopelessly inefficient at information handling – were not reproduced by Jeffreys, nor would he wish to.
It is not disputed that you can control more prognostic covariates in a balanced, controlled trial at a smaller sample size than you can with randomization. If you budget for a large trial, you can perform some error detection by running 3 smaller controlled, but balanced trials, and then combine them via meta-analysis at the end.
You can actually test if your 3 samples are reasonably homogeneous using one of the tests of distributional equality. The most general is the Kolmogorov-Smirnov. Low p values indicate heterogeneity and information loss, and the effective sample size can be adjusted downward from the relevant experiment.
Of course, there are going to be logistical challenges and various economic and clinical considerations to balance. I’m not sure how to handle safety monitoring, for example. I’d be inclined to let each experiment assess that independently, but that might not be economical nor ethical, and centralized safety monitoring might be better. I don’t think you would need to present much evidence to me that centralized safety monitoring along with decentralized efficacy monitoring, is preferable. This would require complex simulations, though.
But it is now time to do the hard thinking and try to do better than randomization.
Related thread: