My Response to the Challenge from Iván Díaz
This is What Happened
- Researcher R poses an unconditional (on the future) question about response Y
- Deaths before Y is measured prevent the question from being answered
- Causal formalism asks the researcher to change the question to a conditional one with unverifiable assumptions and a reduced sample size
- R should have been asked to change the question to a different unconditional one that is consonant with available data
- The new unconditional question meshes with available data and intent-to-treat and requires no causal formalism
My Argument
Formalizing a physical or thought process can expose logic and assumptions. It can lead to better understanding and to process improvement. For formalization to be helpful, I submit that the following conditions must hold.
- If the process is intended to involve data, the formalization should be concordant with the data generating process.
- The original question must be cohesive and answerable.
- The formalization must answer the question in a way that is useful for time zero decision-making.
- Assumptions made by the process must be visible.
- The assumptions must be verifiable by data even if the sample size needed to check assumptions is quite large.
When formalization requires changing the question, requires making assumptions that could never be verified, or is discordant with the likely data generating process, it is not very helpful.
For straight-up RCTs with no interruptions to the outcome flow, where the analysis includes all randomized patients, I have asserted that no causal formalism is needed. This is because the design excludes all but two explanations for a treatment difference: bad luck (data randomness) and the experimental manipulation. Data randomness is addressed by having probability distributions for statistics (in the frequentist world) or parameters (with Bayes). I also stated in the linked blog article that the best place for formal causal inference may be for RCTs where secondary non-intent-to-treat questions are asked when there is interruption to the outcome flow, resulting in comparisons of non-random samples from the original RCT cohort. Examples of such settings include estimating the treatment benefit were everyone assigned to the treatment actually 100% adhering to it, comparison of longitudinal trends when there are dropouts, and the example to be discussed below
Iván Díaz posed a good example of a non-straight-up RCT in which the target outcome Y = weight at month m may not be observed because some of the patients may die between randomization (time zero) and month m. The clinical question is to estimate the effect of an antidepressant (vs. control) on weight Y at month m. The naive approach of comparing Y among m-month survivors is usually wrong. Using Iván’s notation with assignment to antidepressant denoted by A=1 and to control by A=0, this naive comparison when stated as a difference in non-baseline-weight-adjusted means (not a great idea) is E(Y|A=1,S=1) - E(Y|A=0, S=1) where S=1 means survival until at least time m. The selection process involved here is based on a different outcome (death) that is assessed post time zero. The post-time-zero exclusion represents non-random sampling which creates a bias of unknown magnitude. The researcher might be extremely lucky and the bias=0 but because the bias is not restricted in any way, the contrast is a biased one and is frowned upon in the RCT world.
Going back to my 5-item list, item 1 is only somewhat satisfied in the example. The data generating process needs to include a huge amount of variation in Y that can be explained by baseline weight, and probably needs to account for baseline height. More important is whether item 2 is satisfied, i.e., is the researcher’s question answerable? Her question was “what is the effect of treatment on weight”. In the presence of death, a truncating, terminating event, this question cannot possibly be answered without unverifiable assumptions (more about that later). An informal poll conducted in 2021 found that researchers have great difficulty interpreting an outcome interrupted by death, and far less difficulty interpreting an outcome that includes death. Thus, at its heart the original question is defective and could be replaced by the following questions that deal only with observables:
- Do patients having baseline weight b who are on the antidepressant tend to have a higher probability of being alive and having Y < y than control patients with baseline weight b?
- Do antidepressant patients have higher chance of being dead or having weight > y than control patients? (the complement of 1)
- If death as considered to be the worst outcome what is the difference in median death-penalized weight?
- Best: Make a 3-dimensional plot showing continuously varying baseline weight b, month m weight threshold y, and P(Y > y or death | A=1, b) - P(Y > y or death, A=0, b).
These questions can all be answered with a multi-state transition model that treats Y as ordinal and has death as an absorbing state. Once a true prospective question is posed by the researcher, we see that the causal formalization does not help to answer it. The RCT once again becomes “straight up.”
Why have I brought up all these other points? Because in drug development, sponsors have had great difficulty operationalizing causal formalizations of intercurrent events, and have more frequently adopted compound endpoints as a solution. They have seen the real advantages of reliance only on observables.
Díaz claims that I would never want to use the potential outcomes framework. He is correct. It doesn’t matter to me what great statisticians’ names are attached to this framework.
Getting back to formalization, consider a randomized crossover study. Unlike the classic parallel-group RCT design we have been discussing, in a crossover study one actually observes Y(0) and Y(1), the outcomes under control (A=0) and treatment (A=1). That means that data are available for computing the variance of the within-patient difference no matter what the correlation between the A=1 response and the A=0 response. Now return to the parallel-group study in which Y(0) and Y(1) can never both be observed on the same patient. Causal inference provides a way to estimate the sample average Y(1)-Y(0) were that to be interesting without accounting for the baseline distribution. The problem is that to be able to operationalize this causal estimand to get causal estimates we must also be able to get uncertainties about the treatment differences. As AP Dawid has elegantly shown, a cohesive analysis that involves “what might have happened had a patient received the treatment they didn’t receive” requires one to know or well-estimate the variance of Y(1)-Y(0). The joint distribution between a patient’s two potential outcomes can never be estimated or checked, even in principle, and even for infinite sample sizes.
So although the marginal causal contrast has a point estimate available, it has no uncertainty estimator, unlike crossover designs. The unobservable within-patient contrast would need to be known to compute a variance. Neyman derived bounds for the variance of the contrast, but conservative bounds reduce the effective sample size; patients are too precious to be wasted in an RCT.
Regarding Iván’s gratuitous point 12, saying that a patient is poorly served when a rigorous causal argument is absent, when the causal argument requires the patient to play the “if I live” game when patients already have enough to consider at baseline, is problematic.
Regarding Iván’s point 16, the assumption that treatment changes no individual’s survival at m is not obviously plausible. And this uses incomplete conditioning since it makes no use of the timing or cause of death. It may be desirable to simply condition on survival until month m, but if there is a mixture of causes of death, anything can happen. One treatment may have more of a certain cause of death than the other, and even if the total number of deaths is the same, the mixture of patients having weight measured may be different across treatment arms in ways that make the conditional estimand impossible to interpret.
On point 17, a “well-defined average causal effect among observed survivors” is of little help to the time zero decision maker. So I see that this is a theorem, but question its utility and the plausibility of its assumptions. For item 20 these “testable implications” do not have much relevance unless the power of the survival comparison is extremely high. With few total deaths this cannot happen. And I am still not clear on whether there is a hidden problem of “different kinds of patients dying on depressant than dying on control therapy” even if overall survival is the same.
The well-defined causal effect among survivors turns the study into a landmark analysis where observation starts (and ends) at time m.
Overall I can see that causal formalism sheds positive light on the “death getting in the way of weight assessment” issue, but find the research question ill-defined, assumptions implausible, and uncertainty calculations especially tenuous. I don’t find the notion of intercurrent events very clear, and would rather just talk about events. A multi-state transition model with death as an absorbing state would cut through a lot of complexity and would involve only observables. There is a trend in drug development with sponsors moving towards ordinal outcomes and multi-state models and away from complexities of intercurrent events.
Notes
- A minor quibble with Iván’s point 15: “population” should be replaced with “RCT cohort” (or approximately “sample”).
- Causal inference formalists who have attempted to take this further into the realm of individual patient decision-making have created a true mess.
- The antidepressant study of weight gain would have been better designed as a longitudinal study with frequent weight measurements. The more frequent the measurements the less damage done by dropouts and death and the more likely a missing-at-random assumption can be true.
- General thoughts about RCT estimands and the role of statistical models in RCTs are here.