So much has been said about p-values and null hypothesis significance testing (NHST) that one blog post probably won't change anybody's opinion.
But I want to recommend the wonderful paper "the null ritual" [1] by Gerd Gigerenzer et al. It shows that a precise understanding of what a p-value means is extremely rare even among statistics lecturers, and, more interestingly, that there have always been fundamentally different understandings even when p-values were invented, e.g. between Fisher and Pearson.
Beyond that, there was a special issue recently in the american statistician discussing at length the issues of p-values in general and .05 in particular [2].
Personally, I of course feel that null hypothesis testing is often stupid and harmful, because null effects never exist in reality, effects without magnitude are useless, and because the "statelessness" of NHST creates or exacerbates problems such as publication bias and lack of power.
Thanks for that. I'll give [1] a read. I'm familiar with [2], and cited one of those papers in the blog.
About the stupid or harmful nature of null hypothesis testing in general, what do you recommend instead for decision making and for summarization of uncertainty? In the scenario of large (yet fast moving) organizations where most people will have little stats background.
Thanks, I hope you find Gigerenzer useful. The paper is a bit academic, but he also wrote a couple of nice popular science books on the (mis-)perception of numbers and statistics, those might be useful in a business environment.
For real-world applications outside engineering and academia, I would rely heavily on confidence intervals and/or confidence bands. For example, the packages from easystats [1] in R have quite a few very useful visualization functions, which make it very easy to interpret results of statistical tests. You can even get a textual precise description, but then again, that's intended for papers and not a wider audience.
Apart from that, I would mainly echo recommendations from people like Andrew Gelman, John Tukey, Edward Tufte etc.: Visuals are extremely useful and contain a lot of data. Use e.g. scatterplots with jittered points to show raw data and the goodness of fit. People will intuitively make more of it than of a single p-value.
Totally agree about visualization, and that those authors are great advocates for it. Confidence intervals are definitely much more informative and intuitive than p-values.
Would the policy be "look at our confidence intervals later and then decide what to do"? One remaining issue is how to have consistent decision criteria, and to convey it ahead of time. Imagine a context with 10-50 teams at a company that run experiments, where the teams are implicitly incentivized to find ways to report their experiments as successful. Quantified criteria can be helpful in minimizing that bad incentive.
Thanks for this, looking forward to read it! I would also recommend in turn the paper by Greenland et al (2016) called "Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations.". It's a great read.
I recently had a chat with a colleague who has "abandoned p-values for bayes factors" in their research, on how, in principle, there's nothing stopping you from having a "bayesian" p-value (i.e. the definition of the p-value at its most general can easily accommodate bayesian inference, priors, posteriors, etc). The counter-retort was more or less "no it can't, educate yourself, bayes factors are better" and didn't want to hear about it. It made me sad.
p-values are an incredibly insightful device. But because most people (ab)use it in the same way most people abuse normality assumptions or ordinal scales as continuous, it's gotten a bad rep and means something entirely different to most now by default.
But I want to recommend the wonderful paper "the null ritual" [1] by Gerd Gigerenzer et al. It shows that a precise understanding of what a p-value means is extremely rare even among statistics lecturers, and, more interestingly, that there have always been fundamentally different understandings even when p-values were invented, e.g. between Fisher and Pearson.
Beyond that, there was a special issue recently in the american statistician discussing at length the issues of p-values in general and .05 in particular [2].
Personally, I of course feel that null hypothesis testing is often stupid and harmful, because null effects never exist in reality, effects without magnitude are useless, and because the "statelessness" of NHST creates or exacerbates problems such as publication bias and lack of power.
[1] http://library.mpib-berlin.mpg.de/ft/gg/GG_Null_2004.pdf
[2] https://www.tandfonline.com/toc/utas20/73/sup1