I’m interested in the tools, methods, and practical skills that people think applied statisticians should know today—particularly those that may not be emphasized in a typical university statistics curriculum.
As a newly practicing applied statistician, the following come to mind:
Quarto and reproducible reporting
LLM’s: how to use them responsibly, and perhaps some basic theory
Optimization
Nonlinear regression
Categorical data analysis (e.g., Agresti)
Version control with Git and GitHub
LaTeX
Data privacy, confidentiality, and governance
Survey design
Reproducible analytical workflows
Data engineering
Causal inference
Measurement-error modeling
Communicating data and statistical results to nontechnical audiences
What else would you add? Are there particular tools, topics, or professional skills that you have found especially valuable in applied statistical work?
I’ve found the last one to be the most lacking. Seems like a lot of people just learn on their own. But practice in an academic setting is really not similar to a professional work setting. The story and where you start is different.
I would not teach \\LaTeX but rather Quarto/markdown and Typst.
Formal causal inference is far less important than experimental and study design and understanding the vast difference between prospective studies and routinely collected data.
Categorical data analysis is a special case of statistical models.
Thank you for these. I noticed your list didn’t include anything on communication. Why is that? Have you had no issues with it? And I don’t just mean how to create and interpret plots. I mean, how to converse with non-statisticians. For me, it seems like a really good skill, yet neither my undergrad or my grad program had anything for it!
I think the closest thing we had was a Statistical Consulting class.
A clear omission on my part - thanks for pointing it out. Communication and clear, intuitive graphics that don’t give any optical illusions are both very important.
Where do you think decision analysis or decision making would apply? Ultimately when I read a study that is what I am interested in, does the statistician simply provide results?
Oh man - more super important topics that I forget to list! I think decision theory is super important and very seldom taught to statisticians. Only by understanding expected utilities and how they depend on the distribution of risk estimates can someone really understand the irrelevance of \alpha and p-values.
Then there is another all-important learning goal: Bayesian thinking.
Re LLMs: they are estimators, completely in the statistician’s wheelhouse. I would say that the average statistician understands LLMs very well, because the average statistician understands estimators (not just how to estimate them, but also how they behave (inference)). However, the current LLM terminology makes them seem like some new species. We see the standard things like estimation, extrapolation, prediction, etc anthropomorphized as learning, hallucinating, reasoning, etc.
Further, decision theory interacts very closely with the reinforcement learning used to then take an LLM and build a human-in-the-loop chatbot (in fact it is an application of decision theory).
I talk about all this in the paper Treatment, evidence, imitation, chat.
I hope that statisticians such as yourself @DylanArmbruster engage fully here, for the sake of the healthcare system, at least, and probably many other parts of society, too.
I put some thoughts here where I argue that the best use of LLMs for statisticians is to make them better statisticians, for example by having LLM clarifying assumptions being made and writing simulation code to check the performance of the proposed method. We should be running simulations far more often, and LLMs are exceptionally helpful for this purpose. (As an aside, statisticians are really good with simulation use feature selection far less often because they see how unreliable it is. They also learn how irrelevant the CLT is).
it depends where you reside i feel, ie academia versus industry versus government. At the regulator you’ll need to know the drug development process, all phases, and the various guidelines. If you want to freelance on the other hand then you’ll need to be familiar with a broad collection of methods and diseases, animal studies, pk/pd, and you need to be a capable programmer too. Yet a statistician in industry does not need to write code, however they need to understand/appreciate the commercial side. There are inevitable clashes with these different perspectives, eg if a regulatory guideline stipulates change from baseline then youre just stuck with it. Academics are freed up. They produce mostly phase II studies and given the smaller sample sizes and freedom they can contemplate Bayesian methods etc. If youre not in academia you can get by not knowing a single thing about Bayes (unless youre in eg rare diseases, pediatrics). Communication is not that important i feel. Likely you’ll be communicating with other technical people, eg at the regulator. Communication is something for marketing to worry about, they would say. Based on feedback from reviewers when submitting to medical journals i would say: learn mediation analysis, they all want it, also estimands