# What Should a Newly Applied Statistician Know Beyond the Standard Ciriculum

**URL:** <https://discourse.datamethods.org/t/what-should-a-newly-applied-statistician-know-beyond-the-standard-ciriculum/28855>\
**Category:** education\
**Created:** [September 10, 2026, 2:35pm UTC](https://discourse.datamethods.org/t/what-should-a-newly-applied-statistician-know-beyond-the-standard-ciriculum/28855 "2026-09-10T14:35:05Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![DylanArmbruster](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/dylanarmbruster/32/5319_2.png) [@DylanArmbruster](https://discourse.datamethods.org/u/DylanArmbruster)\
**Post date:** [September 10, 2026, 2:35pm UTC](https://discourse.datamethods.org/t/what-should-a-newly-applied-statistician-know-beyond-the-standard-ciriculum/28855/1 "2026-09-10T14:35:05Z")

</div>

Hello,

I’m interested in the tools, methods, and practical skills that people think applied statisticians should know today—particularly those that may not be emphasized in a typical university statistics curriculum.

As a newly practicing applied statistician, the following come to mind:

- Quarto and reproducible reporting

- LLM’s: how to use them responsibly, and perhaps some basic theory

- Optimization

- Nonlinear regression

- Categorical data analysis (e.g., Agresti)

- Version control with Git and GitHub

- LaTeX

- Data privacy, confidentiality, and governance

- Survey design

- Reproducible analytical workflows

- Data engineering

- Causal inference

- Measurement-error modeling

- Communicating data and statistical results to nontechnical audiences

What else would you add? Are there particular tools, topics, or professional skills that you have found especially valuable in applied statistical work?

I’ve found the last one to be the most lacking. Seems like a lot of people just learn on their own. But practice in an academic setting is really not similar to a professional work setting. The story and where you start is different.

---

<div class="post-metadata">

**Author:** ![f2harrell](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/f2harrell/32/165_2.png) [@f2harrell](https://discourse.datamethods.org/u/f2harrell)\
**Post date:** [September 10, 2026, 3:35pm UTC](https://discourse.datamethods.org/t/what-should-a-newly-applied-statistician-know-beyond-the-standard-ciriculum/28855/2 "2026-09-10T15:35:45Z")

</div>

A good question, and a tough one. The most important things to know, in descending order of importance, are in my current view:

1. [Statistical principles](https://fharrell.com/post/principles)
2. Experimental and study design (including sample survey design) and understanding biases in observational research
3. Measurement including why not to use [change scores](https://hbiostat.org/bbr/change)
4. [Preservation of information in data](https://hbiostat.org/bbr/info)
5. Reproducible analysis and reporting including pre-specification
6. Effective sample size and what limits that places on analysis; overfitting
7. Statistical modeling [subsumes all statistical tests](https://fharrell.com/post/cpm)

For general statistical analysis topics see [BBR](https://hbiostat.org/bbr).

I would not teach \\LaTeX but rather Quarto/markdown and [Typst](https://typst.app).

Formal causal inference is far less important than experimental and study design and understanding the vast difference between prospective studies and [routinely collected data](https://www.bmj.com/content/bmj/393/bmj-2025-087812.full.pdf).

Categorical data analysis is a special case of statistical models.

---

<div class="post-metadata">

**Author:** ![DylanArmbruster](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/dylanarmbruster/32/5319_2.png) [@DylanArmbruster](https://discourse.datamethods.org/u/DylanArmbruster)\
**Post date:** [September 10, 2026, 4:43pm UTC](https://discourse.datamethods.org/t/what-should-a-newly-applied-statistician-know-beyond-the-standard-ciriculum/28855/3 "2026-09-10T16:43:32Z")

</div>

Hello Frank,

Thank you for these. I noticed your list didn’t include anything on communication. Why is that? Have you had no issues with it? And I don’t just mean how to create and interpret plots. I mean, how to converse with non-statisticians. For me, it seems like a really good skill, yet neither my undergrad or my grad program had anything for it!

I think the closest thing we had was a Statistical Consulting class.

---

<div class="post-metadata">

**Author:** ![f2harrell](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/f2harrell/32/165_2.png) [@f2harrell](https://discourse.datamethods.org/u/f2harrell)\
**Post date:** [September 11, 2026, 12:12pm UTC](https://discourse.datamethods.org/t/what-should-a-newly-applied-statistician-know-beyond-the-standard-ciriculum/28855/4 "2026-09-11T12:12:31Z")

</div>

A clear omission on my part - thanks for pointing it out. Communication and clear, intuitive graphics that don’t give any optical illusions are both very important.

---

<div class="post-metadata">

**Author:** ![trumanfrancis](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/trumanfrancis/32/47_2.png) [@trumanfrancis](https://discourse.datamethods.org/u/trumanfrancis)\
**Post date:** [September 11, 2026, 3:34pm UTC](https://discourse.datamethods.org/t/what-should-a-newly-applied-statistician-know-beyond-the-standard-ciriculum/28855/5 "2026-09-11T15:34:49Z")

</div>

Where do you think decision analysis or decision making would apply? Ultimately when I read a study that is what I am interested in, does the statistician simply provide results?

---

<div class="post-metadata">

**Author:** ![f2harrell](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/f2harrell/32/165_2.png) [@f2harrell](https://discourse.datamethods.org/u/f2harrell)\
**Post date:** [September 11, 2026, 4:47pm UTC](https://discourse.datamethods.org/t/what-should-a-newly-applied-statistician-know-beyond-the-standard-ciriculum/28855/6 "2026-09-11T16:47:15Z")

</div>

Oh man - more super important topics that I forget to list! I think decision theory is super important and **very** seldom taught to statisticians. Only by understanding expected utilities and how they depend on the distribution of risk estimates can someone really understand the irrelevance of \alpha and p-values.

Then there is another all-important learning goal: [Bayesian thinking](https://fharrell.com/post/bthink).

---

<div class="post-metadata">

**Author:** ![samw235711](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/samw235711/32/4255_2.png) [@samw235711](https://discourse.datamethods.org/u/samw235711)\
**Post date:** [September 13, 2026, 2:22pm UTC](https://discourse.datamethods.org/t/what-should-a-newly-applied-statistician-know-beyond-the-standard-ciriculum/28855/7 "2026-09-13T14:22:16Z")

</div>

Re LLMs: they are estimators, completely in the statistician’s wheelhouse. I would say that the average statistician understands LLMs very well, because the average statistician understands estimators (not just how to estimate them, but also how they behave (inference)). However, the current LLM terminology makes them seem like some new species. We see the standard things like estimation, extrapolation, prediction, etc anthropomorphized as learning, hallucinating, reasoning, etc.

Further, decision theory interacts very closely with the reinforcement learning used to then take an LLM and build a human-in-the-loop chatbot (in fact it is an application of decision theory).

I talk about all this in the paper Treatment, evidence, imitation, chat.

I hope that statisticians such as yourself @DylanArmbruster engage fully here, for the sake of the healthcare system, at least, and probably many other parts of society, too.

---

<div class="post-metadata">

**Author:** ![f2harrell](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/f2harrell/32/165_2.png) [@f2harrell](https://discourse.datamethods.org/u/f2harrell)\
**Post date:** [September 13, 2026, 3:23pm UTC](https://discourse.datamethods.org/t/what-should-a-newly-applied-statistician-know-beyond-the-standard-ciriculum/28855/8 "2026-09-13T15:23:17Z")

</div>

I put some thoughts [here](https://hbiostat.org/LLM) where I argue that the best use of LLMs for statisticians is to make them better statisticians, for example by having LLM clarifying assumptions being made and writing simulation code to check the performance of the proposed method. We should be running simulations **far** more often, and LLMs are exceptionally helpful for this purpose. (As an aside, statisticians are really good with simulation use feature selection far less often because they see how unreliable it is. They also learn how irrelevant the CLT is).

---

<div class="post-metadata">

**Author:** ![peterk](https://discourse.datamethods.org/letter_avatar_proxy/v4/letter/p/7993a0/32.png) [@peterk](https://discourse.datamethods.org/u/peterk)\
**Post date:** [October 3, 2026, 6:56pm UTC](https://discourse.datamethods.org/t/what-should-a-newly-applied-statistician-know-beyond-the-standard-ciriculum/28855/9 "2026-10-03T18:56:00Z")

</div>

it depends where you reside i feel, ie academia versus industry versus government. At the regulator you’ll need to know the drug development process, all phases, and the various guidelines. If you want to freelance on the other hand then you’ll need to be familiar with a broad collection of methods and diseases, animal studies, pk/pd, and you need to be a capable programmer too. Yet a statistician in industry does not need to write code, however they need to understand/appreciate the commercial side. There are inevitable clashes with these different perspectives, eg if a regulatory guideline stipulates change from baseline then youre just stuck with it. Academics are freed up. They produce mostly phase II studies and given the smaller sample sizes and freedom they can contemplate Bayesian methods etc. If youre not in academia you can get by not knowing a single thing about Bayes (unless youre in eg rare diseases, pediatrics). Communication is not that important i feel. Likely you’ll be communicating with other technical people, eg at the regulator. Communication is something for marketing to worry about, they would say. Based on feedback from reviewers when submitting to medical journals i would say: learn mediation analysis, they all want it, also estimands

---

<div class="post-metadata">

**Author:** ![llynn](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/llynn/32/2827_2.png) [@llynn](https://discourse.datamethods.org/u/llynn)\
**Post date:** [October 4, 2026, 1:21pm UTC](https://discourse.datamethods.org/t/what-should-a-newly-applied-statistician-know-beyond-the-standard-ciriculum/28855/10 "2026-10-04T13:21:42Z")

</div>

Excellent list. An important addition for enhancing the patient safety of RCT derived guidelines is teaching the statistician gate to guideline (G2G) analysis.

The trialist may not have this knowledge because it comprises an integration of both biological (clinical) and statistical (mathematical) domains. G2G teaching should include methods for disambiguation (formal explication) of the gate and of the guideline target population to identify when and to whom the transport is safely “guideline applicable”.

I call this “G2G analysis” for simplification and clinical integration. The CI community has their preferred terms. Regardless, this addition assures that the statistician learns to enhance the patient safety of her work by determining (with clinical consultation and by interrogation of the clinicians) whether the G2G is primarily “population matching based” (Cause agnostic) or “causal (target) mechanism based” (Bradford Hill type design).

---

<div class="post-metadata">

**Author:** ![f2harrell](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/f2harrell/32/165_2.png) [@f2harrell](https://discourse.datamethods.org/u/f2harrell)\
**Post date:** [October 4, 2026, 3:36pm UTC](https://discourse.datamethods.org/t/what-should-a-newly-applied-statistician-know-beyond-the-standard-ciriculum/28855/11 "2026-10-04T15:36:46Z")

</div>

> [@peterk](#):
>
> given the smaller sample sizes and freedom they can contemplate Bayesian methods

Well said except for that part. Bayesian methods have major advantages in all sample size settings and should most often be used without basing priors on external data, i.e., not usually used in a way that boosts the effective sample size.
