# Reference Collection to push back against "Common Statistical Myths"

**URL:** <https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787>\
**Category:** data analysis\
**Tags:** teaching, journal\
**Created:** [June 27, 2019, 12:51pm UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787 "2019-06-27T12:51:23Z")\
**Posts on this page:** 20\
**Page:** 2

<div class="post-metadata">

**Author:** ![pakeezahs](https://discourse.datamethods.org/letter_avatar_proxy/v4/letter/p/ebca7d/32.png) [@pakeezahs](https://discourse.datamethods.org/u/pakeezahs)\
**Post date:** [July 17, 2019, 1:55pm UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/22 "2019-07-17T13:55:46Z")

</div>

Hello,

Not sure if I can edit the entry directly but I found this to be helpful:

Gary King on [Why Propensity Scores Should Not Be Used for Matching](https://www.youtube.com/watch?v=rBv39pK1iEs)

---

<div class="post-metadata">

**Author:** ![ADAlthousePhD](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/adalthousephd/32/124_2.png) [@ADAlthousePhD](https://discourse.datamethods.org/u/ADAlthousePhD)\
**Post date:** [July 17, 2019, 2:07pm UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/23 "2019-07-17T14:07:35Z")

</div>

> [@pakeezahs](#):
>
> Not sure if I can edit the entry directly but I found this to be helpful:

@f2harrell can comment but there may be a restriction that prevents one from editing unless you have contributed to the forum before, or posted a certain number of times (a bot/quality control issue, I think).

I’ll add this to the wiki, though. Thanks!

---

<div class="post-metadata">

**Author:** ![Matt\_Williams](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/matt_williams/32/829_2.png) [@Matt\_Williams](https://discourse.datamethods.org/u/Matt_Williams)\
**Post date:** [August 22, 2019, 11:42pm UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/24 "2019-08-22T23:42:32Z")

</div>

Fantastic thread. I can’t edit at the moment because I’m a new user, but under **TOPIC: Misunderstood “Normality” Assumptions** this paper might be relevant:

Williams, M. N., Grajales, C. A. G., & Kurkiewicz, D. (2013). Assumptions of multiple regression: Correcting two misconceptions. Practical Assessment, Research & Evaluation, 18(11). [http://www.pareonline.net/getvn.asp?v=18&n=11](http://www.pareonline.net/getvn.asp?v=18&n=11)

(Excuse the self-promotion!)

---

<div class="post-metadata">

**Author:** ![RonanConroy](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/ronanconroy/32/242_2.png) [@RonanConroy](https://discourse.datamethods.org/u/RonanConroy)\
**Post date:** [August 24, 2019, 5:52pm UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/25 "2019-08-24T17:52:29Z")

</div>

What an incredibly useful post. There must be a ‘bravo’ emoji, but I am for sure far too old to know where to find it.

---

<div class="post-metadata">

**Author:** ![ADAlthousePhD](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/adalthousephd/32/124_2.png) [@ADAlthousePhD](https://discourse.datamethods.org/u/ADAlthousePhD)\
**Post date:** [August 26, 2019, 12:42pm UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/26 "2019-08-26T12:42:19Z")

</div>

Thanks Ronan - please feel free to add your own suggestions or references!

---

<div class="post-metadata">

**Author:** ![SameeraDaniels](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/sameeradaniels/32/496_2.png) [@SameeraDaniels](https://discourse.datamethods.org/u/SameeraDaniels)\
**Post date:** [August 27, 2019, 3:25pm UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/27 "2019-08-27T15:25:09Z")

</div>

This thread should be highlighted on Twitter and other social media platforms. Perhaps on Facebook Psychological Methods Discussion Page.

> **[Psychological Methods Discussion Group](https://www.facebook.com/groups/853552931365745/?multi_permalinks=1730398057014557&notif_id=1525483414622303&notif_t=group_highlights&ref=notif)**
>
> "For here we are not afraid to follow truth wherever it may lead, nor to tolerate any errors so long as reason is left free to combat it"
> Thomas Jefferson, 1789
> 
> In this spirit, the Psychological...

---

<div class="post-metadata">

**Author:** ![natea](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/natea/32/864_2.png) [@natea](https://discourse.datamethods.org/u/natea)\
**Post date:** [September 2, 2019, 9:52pm UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/28 "2019-09-02T21:52:39Z")

</div>

Since I haven’t posted before, I don’t think I’m able to directly edit the wiki, but I wanted to provide a nice pair of references that might be valuable additions to the propensity score matching section resources!

Brooks JM, Ohsfeldt RL. Squeezing the balloon: propensity scores and unmeasured covariate balance. Health services research. 2013 Aug;48(4):1487-507.

> <https://www.ncbi.nlm.nih.gov/pubmed/23216471/>
>
> To assess the covariate balancing properties of propensity score-based algorithms in which covariates affecting treatment choice are both measured and unmeasured.A simulation model of treatment choice and outcome.Simulation.Eight simulation scenarios varied with the values placed on measured and unmeasured covariates and the strength of the relationships between the measured and unmeasured covariates. The balance of both measured and unmeasured covariates was compared across patients either grouped or reweighted by propensity scores methods.Propensity score algorithms require unmeasured covariate variation that is unrelated to measured covariates, and they exacerbate the imbalance in this variation between treated and untreated patients relative to the full unweighted sample.The balance of measured covariates between treated and untreated patients has opposite implications for unmeasured covariates in randomized and observational studies. Measured covariate balance between treated and untreated patients in randomized studies reinforces the notion that all covariates are balanced. In contrast, forced balance of measured covariates using propensity score methods in observational studies exacerbates the imbalance in the independent portion of the variation in the unmeasured covariates, which can be likened to squeezing a balloon. If the unmeasured covariates affecting treatment choice are confounders, propensity score methods can exacerbate the bias in treatment effect estimates.

and:

Ali MS, Groenwold RH, Klungel OH. Propensity score methods and unobserved covariate imbalance: comments on “squeezing the balloon”. Health services research. 2014 Jun;49(3):1074-82.  
[https://onlinelibrary.wiley.com/doi/abs/10.1111/1475-6773.12152](https://onlinelibrary.wiley.com/doi/abs/10.1111/1475-6773.12152)

---

<div class="post-metadata">

**Author:** ![baxpr](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/baxpr/32/1172_2.png) [@baxpr](https://discourse.datamethods.org/u/baxpr)\
**Post date:** [September 12, 2019, 7:21pm UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/29 "2019-09-12T19:21:34Z")

</div>

I stumbled on this article about issues with categorization (responder analysis): [Responder analyses and the assessment of a clinically relevant treatment effect | Trials | Full Text](https://trialsjournal.biomedcentral.com/articles/10.1186/1745-6215-8-31)

> Ideally, a clinical trial should be able to demonstrate not only a statistically significant improvement in the primary efficacy endpoint, but also that the magnitude of the effect is clinically relevant. One proposed approach to address this question is a responder analysis, in which a continuous primary efficacy measure is dichotomized into “responders” and “non-responders.” In this paper we discuss various weaknesses with this approach, including a potentially large cost in statistical efficiency, as well as its failure to achieve its main goal. We propose an approach in which the assessments of statistical significance and clinical relevance are separated.

---

<div class="post-metadata">

**Author:** ![lbautista](https://discourse.datamethods.org/letter_avatar_proxy/v4/letter/l/ea5d25/32.png) [@lbautista](https://discourse.datamethods.org/u/lbautista)\
**Post date:** [September 13, 2019, 3:29pm UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/30 "2019-09-13T15:29:52Z")

</div>

> [@f2harrell](#):
>
> I have to strongly disagree with that. Much has been written about this. Briefly, you have to covariate adjust in RCTs to make the most out of the data, i.e., to get the best power and precision

I agree with Dr. Harrel. But I think an argument for adjustment could be made, in addition to the known gain in precision in the estimate of effect. Randomizing treatments is not a full proof method. Even if you randomize a million patients, there is no guarantee the potential outcomes will be the same in treated and “untreated”. It would be very unlikely if the potential outcomes are not very similar, but unlikely/rare things do happen. They are bound to happen due to the very nature of randomization. If we put aside issues of variable selection and how variables will be modelled, adjusting will provide evidence about the exchangeability of the treatment groups beyond the evidence provided in the traditional Table 1 comparing prognostic factors in treated and untreated. Even if each prognostic factor in Table 1 is balanced, this does no imply combinations of multiple prognostic factors are also balanced. In other words, prognostic factors and treatment may not be associated in a crude analysis (presented in Table 1), but may be associated in a multivariate analysis (never presented). To avoid conscious or unconscious manipulation of the analysis, we could decide on what variables we would adjust for pre-facto, as part of the study protocol. Actually, what we report in Table 1 is a list of the variables we believe we should adjust for. These variables could be selected using the same substantive-based approaches we use in observational studies. There doesn’t seem to be a methodological reason for adjusted effect estimates from RCT to be more biased than crude estimates (again, assuming modeling assumptions are correct). In most cases, particularly in mid-size and small trials, the validity of the estimate of the effect of the treatment will be enhanced, and credibility of the RCT findings would increase, if crude and adjusted estimates are consistent.

---

<div class="post-metadata">

**Author:** ![f2harrell](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/f2harrell/32/165_2.png) [@f2harrell](https://discourse.datamethods.org/u/f2harrell)\
**Post date:** [September 14, 2019, 11:40am UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/31 "2019-09-14T11:40:18Z")

</div>

This is described in the “Table one” topic where it is shown that even if you don’t bother to measure any covariates the inference is sound (though not efficient). So I can’t say I agree with this angle on the problem.

---

<div class="post-metadata">

**Author:** ![f2harrell](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/f2harrell/32/165_2.png) [@f2harrell](https://discourse.datamethods.org/u/f2harrell)\
**Post date:** [September 14, 2019, 11:41am UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/32 "2019-09-14T11:41:33Z")

</div>

> [@baxpr](#):
>
> stumbled on this article about issues with categorization (responder analysis):

I’ll add that to the separate responder analysis “loser x4” topic. Great paper.

---

<div class="post-metadata">

**Author:** ![lbautista](https://discourse.datamethods.org/letter_avatar_proxy/v4/letter/l/ea5d25/32.png) [@lbautista](https://discourse.datamethods.org/u/lbautista)\
**Post date:** [September 14, 2019, 1:19pm UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/33 "2019-09-14T13:19:18Z")

</div>

> [@f2harrell](#):
>
> This is described in the “Table one” topic where it is shown that even if you don’t bother to measure any covariates the inference is sound (though not efficient). So I can’t say I agree with this angle on the problem.

I do not argue the non-ajusted estimates are biased. I argue that in “small” and “moderate” size the exchangeability of treatment arms may be compromised and that small differences in several prognostic factors could lead to significant bias in the estimate of effect. This can not be appreciated in univariate comparisons of the distribution of prognostic factors across treatment groups, which is what is presented in Table 1. Therefore, if I see small differences in several prognostic factors or if I see a large difference in a single prognostic factor, I would present crude and adjusted estimates, and would give more weight to the adjusted one, for the purpose of inferences, if they are different. I also argue that even in the case of “large” trials, adjusting would not introduce bias. This is a direct consequence of the independence between treatment assigned and potential outcome that results from randomization. Therefore, if adjusted and crude estimates differ in a large trial, I’d be inclined to believe something was wrong with the model used for the adjustment. Briefly, there is nothing wrong with adjusting for prognostic factors in a RCT, either from the perspective of precision or bias, unless the model used for the adjustment is misspecified.

---

<div class="post-metadata">

**Author:** ![PerPersvensson](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/perpersvensson/32/1198_2.png) [@PerPersvensson](https://discourse.datamethods.org/u/PerPersvensson)\
**Post date:** [November 25, 2019, 4:24pm UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/36 "2019-11-25T16:24:45Z")

</div>

Great post  
I wonder if the first topic could be broadened to also apply to observational “table one’s” such as descriptives of baseline data in different exposure groups in a cohort study ? The STROBE criteria argue against significance testing.

> **[Strengthening the Reporting of Observational Studies in Epidemiology...](https://journals.plos.org/plosmedicine/article?id=10.1371%2Fjournal.pmed.0040297)**
>
> In this explanatory and elaboration document Mattias Egger and colleagues provide the meaning and rationale of each checklist item on the STROBE Statement.

  
Would be interesting to hear your thought also on observational studies  
Thanks

---

<div class="post-metadata">

**Author:** ![tho\_ols](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/tho_ols/32/1173_2.png) [@tho\_ols](https://discourse.datamethods.org/u/tho_ols)\
**Post date:** [November 28, 2019, 11:18am UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/37 "2019-11-28T11:18:30Z")

</div>

Added topic on significance testing in pilot studies with some useful references, feel free to expand.

---

<div class="post-metadata">

**Author:** ![SteveSchwartz](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/steveschwartz/32/1420_2.png) [@SteveSchwartz](https://discourse.datamethods.org/u/SteveSchwartz)\
**Post date:** [March 26, 2020, 8:41pm UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/38 "2020-03-26T20:41:12Z")

</div>

I feel that this issue of not calculating and presenting p-values in Table 1 extends to observational studies, for multiple reasons. That said, I am not aware of any published papers that have made this argument.

---

<div class="post-metadata">

**Author:** ![mgrafit](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/mgrafit/32/1568_2.png) [@mgrafit](https://discourse.datamethods.org/u/mgrafit)\
**Post date:** [May 28, 2020, 9:59am UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/39 "2020-05-28T09:59:51Z")

</div>

Thanks for initiating this list. I discovered 3 articles I didn’t know about before that.  
Please see below some news references, by topic.

**TOPIC: Analyzing “Change” Measures in RCT’s**

- Archie J.P. Mathematic coupling of data – A common source of error. Annals of Surgery, 1980, 193: 296-303
- Yanez N.D. et al. The effects of measurement error in response variables and tests of association of explanatory variables in change models. SiM, 1998, 17: 2597-2606.
- Senn S. Change from baseline and analysis of covariance revisited. SiM, 2006, 25: 4334-4344.
- Tu Y-K., et al. Revisiting the relation between change and initial value: A review and evaluation. SiM, 2007. [https://doi.org/10.1002/sim.2538](https://doi.org/10.1002/sim.2538)
- Braun J., et al. Accounting for baseline differences and measurement error in the analysis of change over time. SiM, 2013. [https://doi.org/10.1002/sim.5910](https://doi.org/10.1002/sim.5910)
- Tu Y-K. Testing the relation between percentage change and baseline value. ScientificReports, 2016. [https://doi.org/10.1038/srep23247](https://doi.org/10.1038/srep23247)
- Clifton et al. Comparing different ways of calculating sample size for two independent means: A worked example. CCT, 2019. [https://doi.org/10.1016/j.conctc.2018.100309](https://doi.org/10.1016/j.conctc.2018.100309)

**TOPIC: Stepwise Variable Selection (Don’t Do It!)**

- Heinze G. et al. Variable selection – A review and recommendations for the practicing statistician. BiomJ, 2017. [https://doi.org/10.1002/bimj.201700067](https://doi.org/10.1002/bimj.201700067)
- Ahamadi M. et al. Operating characteristics of stepwise covariate selection in pharmacometric modeling. JPKPD, 2019. [https://doi.org/10.1007/s10928-019-09635-6](https://doi.org/10.1007/s10928-019-09635-6)

**TOPIC: Inappropriately Splitting Continuous Variables Into Categorical Ones**

- Weinberg C.R. How bad is categorization? Epidemiology, 1995, 6:345-346.
- Senn S. Disappointing dichotomies. PharmStat, 2003. [https://doi.org/10.1002/pst.090](https://doi.org/10.1002/pst.090)
- Chen H. et al. Biased odds ratios from dichotomization of age. SiM, 2007. [https://doi.org/10.1002/sim.2737](https://doi.org/10.1002/sim.2737)
- VanWalraven C. et al. Leave ‘em alone – why continuous variables should be analyzed as such. Neuroepidemiology 2008, [https://doi.org/10.1159/000126908](https://doi.org/10.1159/000126908)

Hope it will be useful to you all.

---

<div class="post-metadata">

**Author:** ![verbeekmc](https://discourse.datamethods.org/letter_avatar_proxy/v4/letter/v/919ad9/32.png) [@verbeekmc](https://discourse.datamethods.org/u/verbeekmc)\
**Post date:** [December 3, 2020, 9:43am UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/40 "2020-12-03T09:43:51Z")

</div>

Related to power and the measurement of change scores vs. group differences, does anyone know references about the number of measurements (for instance, adding an ‘in-between measurement’) to increase power? Does it matter if you are only going to investigate group differences or is just a pre-measurement enough?

---

<div class="post-metadata">

**Author:** ![f2harrell](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/f2harrell/32/165_2.png) [@f2harrell](https://discourse.datamethods.org/u/f2harrell)\
**Post date:** [December 3, 2020, 1:17pm UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/41 "2020-12-03T13:17:01Z")

</div>

Since this does not fit in with the ‘myths’ topic please start a new topic with appropriate primary and secondary topic and tag choices. Then I’ll remove this one.

---

<div class="post-metadata">

**Author:** ![EpiLearneR](https://discourse.datamethods.org/letter_avatar_proxy/v4/letter/e/f9ae1b/32.png) [@EpiLearneR](https://discourse.datamethods.org/u/EpiLearneR)\
**Post date:** [December 14, 2020, 2:45am UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/42 "2020-12-14T02:45:39Z")

</div>

TOPICS suggestion: It would be useful to have a reference collection on P value and confidence interval myths

---

<div class="post-metadata">

**Author:** ![R\_cubed](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/r_cubed/32/1518_2.png) [@R\_cubed](https://discourse.datamethods.org/u/R_cubed)\
**Post date:** [December 14, 2020, 2:47pm UTC](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787/43 "2020-12-14T14:47:24Z")

</div>

@EpiLearneR You might find this open access paper valuable:

> **[Statistical tests, P values, confidence intervals, and power: a guide to...](https://link.springer.com/article/10.1007/s10654-016-0149-3)**
>
> Misinterpretation and abuse of statistical tests, confidence intervals, and statistical power have been decried for decades, yet remain rampant. A key problem is that there are no interpretations of these concepts that are at once simple, intuitive,...

Search for @Sander (Sander Greenland) here, and you will find a lot of excellent papers on these misinterpretations as well as corrective measures. If you read them slowly and are prepared to look up the mathematics you do not know, they will teach you a lot.

Here is a good link to some of the things he has written on this:

> [@Predicting survival of cancer patients using the horoscope: astrology, causal inference, seasonality, frequentist & Bayesian approach](https://discourse.datamethods.org/t/predicting-survival-of-cancer-patients-using-the-horoscope-astrology-causal-inference-seasonality-frequentist-bayesian-approach/3647/39):
>
> Answering that gets into a big topic area… Various transforms of the observed P-value p such as 1−p, 1/p, and log(1/p) have been used as measures of evidence against statistical models or hypotheses (which are parts of models). Part of the interpretation problem is that those statistical models get confused with the theories that predict them. Correct understanding however requires keeping them distinct, because the measures refer only to the statistical model used to derive the p-value; purely…

[Previous page](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787.md?page=1)

[Next page](https://discourse.datamethods.org/t/reference-collection-to-push-back-against-common-statistical-myths/1787.md?page=3)
