I was thinking specifically of balanced allocation experimental designs such as minimization. The entire literature on this issue is rather frustrating to read, especially since the development of resampling methods.
AFAICT, the entire dispute since Fisher and Gossett, was the validity of using a model such as Student’s t distribution, to examine the data from a controlled, but non-random experiment. The biggest complaint (which isn’t obvious) is that the data from such experiments has thinner tails than what would be expected from a normal distribution.
The following discusses the Gossett vs Fisher design of experiments debate from the perspective of Gossett. It is more of a historical narrative than a mathematical statistics paper, but it does touch upon the debate. The figure 1 (page 33 in the PDF) shows power curves for balanced vs. randomized designs.
For large (and economically worthwhile effects), balanced experiments analyzed by the t-test had greater power to detect effects. Power to detect effects declined as the effect size gets smaller.
Ziliak, S. T. (2014). Balanced versus randomized field experiments in economics: why WS Gosset aka “Student” matters. Review of Behavioral Economics, 1(1-2), 167-208. (PDF)
Correct me if I’m wrong, but to me, this debate is moot because we can simply resample from the data, and compute a valid empirical null distribution, as Nathan Kallus discussed in his paper. Whether you want confidence distributions or likelihood functions, resample techniques permit computation of statistical summaries that satisfy a wide audience.
The reduction in sample size for a single experiment that wishes to account for N prognostic factors permits the possibility of actually repeating the protocol as a form of cross validation and builds in redundancy and at least error detection, as the information theorists and communication engineers would think about it.
Essentially, I would budget for a valid randomized design, but conduct 3 smaller, balanced experiments, and then synthesize them using meta-analytic techniques.
This might not be possible for rare disease, but balanced designs still maximize the available information to the experimenter.
All the powerful tools that have been developed – bias and sensitivity analysis as advocated by @Sander, the use of permutation techniques to detect bias, as described by Rosenbaum in this old article, are now available with the vast increases in computing power.
Rosenbaum, Paul R. (1989) “On Permutation Tests for Hidden Biases in Observational Studies: An Application of Holley’s Inequality to the Savage Lattice.” Ann. Statist. 17 (2) 643 - 653, June, 1989. On Permutation Tests for Hidden Biases in Observational Studies: An Application of Holley's Inequality to the Savage Lattice link
Further reading:
Saint-Mont calculates the sample sizes needed for randomization to balance factors (binary through continuous) with varying degrees of frequency. The more common a prognostic factor is, the larger the sample needs to be in order for randomization to account for it. Groups formed by randomization aren’t really exchangeable until hundreds of observations have been collected.