# How to interpret “confidence intervals” in observational studies

**URL:** <https://discourse.datamethods.org/t/how-to-interpret-confidence-intervals-in-observational-studies/28318>\
**Category:** general\
**Created:** [August 11, 2025, 9:49pm UTC](https://discourse.datamethods.org/t/how-to-interpret-confidence-intervals-in-observational-studies/28318 "2025-08-11T21:49:23Z")\
**Posts on this page:** 1\
**Showing post:** 39

<div class="post-metadata">

**Author:** ![R\_cubed](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/r_cubed/32/1518_2.png) [@R\_cubed](https://discourse.datamethods.org/u/R_cubed)\
**Post date:** [August 16, 2025, 12:25pm UTC](https://discourse.datamethods.org/t/how-to-interpret-confidence-intervals-in-observational-studies/28318/39 "2025-08-16T12:25:18Z")

</div>

> [@f2harrell](#):
>
> **Do note that @ChristopherTong‘s wonderful paper is mostly related to a frequentist perspective.** Bayesian models can take model uncertainty into account, when practitioners are diligent to add the extra parameters needed.

This was going to be my criticism. I’ll only add that your Bayesian solution will not be satisfactory to the dogmatic frequentists because it ends up being “subjective”.

Taken literally, the recommendations in that paper cited by Christopher Tong would require scientists to so severely discount **any** source of data that he or she did not personally collect, it might as well be treated as having no credibility. That would include meta-analysis of RCT’s, since meta-analyses of RCTs are also merely “observational”.

**Berlin JA, Golub RM.** _Meta-analysis as Evidence: Building a Better Pyramid._ _JAMA._ 2014;312(6):603–606 [link](https://jamanetwork.com/journals/jama/fullarticle/1895230)

> _JAMA_ considers meta-analysis to represent an observational design, with measures that should be interpreted as associations rather than causal effects.

I thought Efron’s old paper “Why isn’t everyone Bayesian” settled the dispute on foundations, where commentators agreed that the subjectivists had the best arguments, but practice dictated use of other procedures that were close approximations.

**Efron, B. (1986).** Why isn’t everyone a Bayesian?. _The American Statistician_, _40_(1), 1-5. [link](https://www.tandfonline.com/doi/abs/10.1080/00031305.1986.10475342)

If the frequentists are going to complain about the Bayesian prior, why should the Bayesian meekly accept that a study **that claims** to use randomization is inherently more credible than any other type of formal model for the data collection procedure, when other claims from that same agent are treated with doubt, if not strong skepticism?

The Bayesians in recent times (with an important exception) have been too charitable in conditioning on the reports of others that claim to use randomization. Rethinking this naive acceptance of reports that claim to use randomization can lead to an impasse – how does a community of scientists collect data and conduct experiments to decide questions of fact that leads to a convergence of opinion, which seems to be the explicit goal of scientific inquiry?

> [@Definition of statistics](https://discourse.datamethods.org/t/definition-of-statistics/6191/10):
>
> Blockquote “Bayesian decision making” it’s not very common in med research, as far as I can see. And it is also not very commonly meant in intro statistics books The fact that a decision theory perspective is mostly absent from medical research and intro stat books I see as an important cause of poor research conduct and misinterpretation of statistical methods. I have not seen any rebuttal to the claim that the maximization of utility, when applied to the scientific research domain is simp…

The field of mechanism design provides some answers, but none of this appears to be on the radar of statisticians. It would unify the frequentist and Bayesian perspectives on the design of experiments by providing a common language to discuss goals and methods using game theoretic concepts, which have already been used in theoretical statistics.

> **[Mechanism Design - Leonid Hurwicz](https://leonidhurwicz.org/mechanism-design/)**
>
> \[et\_pb\_section fb\_built=”1″ admin\_label=”section” \_builder\_version=”3.22″ bb\_built=”1″ \_i=”0″ \_address=”0″\]\[et\_pb\_row admin\_label=”row” \_builder\_version=”3.25″ background\_size=”initial” background\_position=”top\_left” background\_repeat=”repeat” \_i=”0″...

> The key idea of mechanism design is identifying goals first and then attempting to design a system that achieves those goals. In other words, at the beginning of the process, the goals are given, and the ideal mechanism is the unknown. This contrasts with “positive” or predictive economics, which studies the actual or likely outcomes of a given system. In that case, the system is the given, and the outcomes are the unknowns.

The closest paper that comes to approaching this Frequentist/Bayesian dispute from a mechanism design point of view (without realizing it) is a paper from 2014 by Atkinson, that discusses biased coin designs, which I mentioned in this thread:

> [@Design of Experiments in Economics vs Medicine: a Decision Theory POV](https://discourse.datamethods.org/t/design-of-experiments-in-economics-vs-medicine-a-decision-theory-pov/6567):
>
> This post is inspired by the @f2harrell reference to the economist Lars P Syll’s post [The Limited Value of Randomization](https://heterodox.economicblogs.org/lars-p-syll/2023/syll-value-randomization). It is a good entry point into the criticisms of so-called “evidence based medicine” heuristics more generally. Syll is often linked to by @Sander_Greenland on Twitter for his skepticism of applied stats and mathematics in the realm of social science (especially econometrics). First, it should be mentioned that different subject areas have different challenges, which need t…

**Atkinson AC. (2014).** Selecting a Biased-Coin Design, Statistical Science, Statist. Sci. 29(1), 144-163 [link](https://projecteuclid.org/journals/statistical-science/volume-29/issue-1/Selecting-a-Biased-Coin-Design/10.1214/13-STS449.full)

> Biased-coin designs are used in clinical trials to allocate treatments with some randomness while maintaining approximately equal allocation. More recent rules are compared with Efron’s [Biometrika 58 (1971) 403–417] biased-coin rule and extended to allow balance over covariates. The main properties are loss of information, due to imbalance, and selection bias. Theoretical results, mostly large sample, are assembled and assessed by small-sample simulations. The properties of the rules fall into three clear categories. A Bayesian rule is shown to have appealing properties; at the cost of slight imbalance, bias is virtually eliminated for large samples.

Recommended Reading

> **[The Randomized In Randomized Controlled Trials Is Pure Superstition; Bad Magic](https://www.wmbriggs.com/post/47450/)**
>
> Why You Need To Read This My dear readers, a complex subject today, presented in the guise of a book review. We are increasingly beset by lunatic psychotic sociopathic rulers wielding The Science l…

Briggs is a bit extreme in his critique, but his fundamental point is valid, and his argument that “randomization” is a religious ritual that “blesses” a data set can explain the rise of “non-comparative randomized trials” discussed in this other thread:

> [@Randomized non-comparative trials: an oxymoron?](https://discourse.datamethods.org/t/randomized-non-comparative-trials-an-oxymoron/20863):
>
> Randomized non-comparative trials (RNCTs) are becoming increasingly more popular, [particularly in oncology](https://pubmed.ncbi.nlm.nih.gov/38402886/), and being published in prominent clinical journals. The idea is to randomize between two or more treatment arms and then not compare them with each other but instead compare each arm with historical controls or prespecified values. Essentially, RNCTs act as single-arm trials for each treatment group. Convenience sampling is used, similarly to standard comparative randomized trials, so the…

The following should be studied together, as it presents a coherent way to do probabilistic bias analysis that @Sander has advised for decades, based on Bayesian Trust modelling (which is used in information security contexts):

[![](https://discourse.datamethods.org/uploads/default/original/2X/0/000b85ecd4381b06ec1666d09ee8c97056cf3b58.jpeg "Greenland seminar, April 19, 2023") ](https://www.youtube.com/watch?v=N7-yn5dd7Hg&t=735s)

**Josang, A., Hayward, R., & Pope, S. (2006).** Trust network analysis with subjective logic. In _Conference Proceedings of the Twenty-Ninth Australasian Computer Science Conference (ACSW 2006)_ (pp. 85-94). Australian Computer Society. [link](https://eprints.qut.edu.au/10146/)

Of course this 2005 paper by @Sander_Greenland remains relevant for this thread. Note he also mentions the importance of bias analysis in the context of randomized studies.

**Greenland, S. (2005).** Multiple-bias modelling for analysis of observational data. _Journal of the Royal Statistical Society Series A: Statistics in Society_, _168_(2), 267-306. [PDF](http://www.medicine.mcgill.ca/epidemiology/Joseph/courses/common/greenland2005.pdf)

This paper by Philip Stark is also worth re-examination (updated for the actual published citation):

**Stark, P. B. (2022).** Pay no attention to the model behind the curtain. Pure and Applied Geophysics, 179(11), 4121-4145. [link](https://link.springer.com/article/10.1007/s00024-022-03137-2)

> [@Recommended Reading: Pay No Atttention to the Model Behind the Curtain by Philip B. Stark](https://discourse.datamethods.org/t/recommended-reading-pay-no-atttention-to-the-model-behind-the-curtain-by-philip-b-stark/4216):
>
> I was looking at some threads over at Andrew Gelman’s blog and found an important paper recommended by @Sander in the comments to a discussion on modelling: [https://statmodeling.stat.columbia.edu/2020/12/04/discussion-of-uncertainties-in-the-coronavirus-mask-study-leads-us-to-think-about-some-issues/#comment-1604678](https://statmodeling.stat.columbia.edu/2020/12/04/discussion-of-uncertainties-in-the-coronavirus-mask-study-leads-us-to-think-about-some-issues/#comment-1604678) Link to actual paper: Stark, P (2016) Pay No Attention to the Model Behind the Curtain preprint ([pdf](https://www.stat.berkeley.edu/~stark/Preprints/eucCurtain15.pdf)) It’s a tough paper regardless of your statistical philosophy, as he raises i…

---

_[View the full topic](https://discourse.datamethods.org/t/how-to-interpret-confidence-intervals-in-observational-studies/28318)._
