Perhaps @llynn will correct me if I’m wrong, but his complaint seems to be: how do you discover what variables to condition on when the measurement process is so noisy?
My complaint is that randomization can work if other parts of the model have very low probabilities of being in error. This isn’t the case in sepsis research, as Lawrence has overwhelmingly demonstrated
In the context of sepsis research: the RCT methodology didn’t aid his community of medical scientists from discovering a crucial flaw in their assumptions. This lead to 30+ years of wasted research resources.
This wasn’t clear in Fisher’s time, but any protocol that is deemed “scientific” needs to be:
- decentralized ie. no central authority that determines “truth”. Consensus determines truth.
- resistant to misleading reports, whether through honest error, or through intentional manipulation.
All of these properties are implied in Feynman’s classic Cal Tech speech, where he introduced the notion of Cargo Cult Science. The essence of science, according to Feynman is:
- Don’t fool others, have scientific integrity.
- Don’t fool yourself, and never forget you are very easy to fool.
There are a number of methodological guidelines he provides, that statistical recommendations (ie. randomization) are in conflict with – especially the need for repetition in order to validate a claim.
Other kinds of errors are more characteristic of poor science. When I was at Cornell. I often talked to the people in the psychology department. One of the students told me she wanted to do an experiment that went something like this—I don’t remember it in detail, but it had been found by others that under certain circumstances, X, rats did something, A. She was curious as to whether, if she changed the circumstances to Y, they would still do, A. So her proposal was to do the experiment under circumstances Y and see if they still did A.
I explained to her that it was necessary first to repeat in her laboratory the experiment of the other person—to do it under condition X to see if she could also get result A—and then change to Y and see if A changed. Then she would know that the real difference was the thing she thought she had under control.
She was very delighted with this new idea, and went to her professor. And his reply was, no, you cannot do that, because the experiment has already been done and you would be wasting time. This was in about 1935 or so, and it seems to have been the general policy then to not try to repeat psychological experiments, but only to change the conditions and see what happens.
These intuitive notions of Feynman regarding honest scientific process were formalized in the computer science literature, starting with the discussion by:
Lamport, L.; Shostak, R.; Pease, M. (1982). “The Byzantine Generals Problem” (PDF). ACM Transactions on Programming Languages and Systems. 4 (3): 382–401.
No communications engineer would today implement a protocol that does not at least have some form of error detection or correction if possible. In critical systems that are decentralized, resistance to parts of the system to arbitrary errors in other parts is known as being tolerant to
Byzantine Fault.
Byzantine fault tolerance is a family of problems, with the answer contingent upon how much computational power you grant to the agents submitting misleading signals/reports.
A system is Byzantine Fault Tolerant if > 2/3 of the signals/reports are reliable.
(Agents in this context need not be human beings; they could be sensors or computer processors).
The question I’m asking myself currently: how can this finding be adapted to the clinical trial context, without requiring huge sample sizes?
I believe this total cost of experimentation (initial claim + validation) can done with sample sizes of < 2n (with n being the maximum size of an initial randomized trial) with some hard thinking, as E.T. Jaynes recommended. But that will require rethinking the role of randomization among a community of scientists.
Who could object to a process that provides Byzantine fault resistance for less than the cost of 2 RCTs?