Jonathan Shuster
I wish to commend Professor Gang Xie for his courageous editorial, “A Practicing Statistician’s Plea,” that appeared recently in Amstat News. I concur with his plea and hope his advice is taken seriously by our profession. To paraphrase his directives, we should ideally make no assumptions in our analyses, and when we must, disclose any we make in the statistical considerations we put into our co-authorships. We should also use point and interval estimates of effect size and downplay the P-values. I would like to add the following three important caveats to this advice.
- Avoid diagnostic testing for assumptions, especially if you would consider “changing horses in midstream” if you reject the assumptions. Whether or not you change horses, you would need to account for the impact of the diagnostic tests upon your point and interval estimates you ultimately obtain. Whether these tests passed or failed, you cannot legitimately ignore the fact that such testing was done. This is an unsolved dilemma. Jonathan Shuster, in his 2005 Statistics in Medicine article, “Diagnostics for Assumptions in Moderate to Large Simple Clinical Trials: Do They Really Help?” presents a compelling case of the dangers involved in diagnostic testing for assumptions. Failure to reject a null hypothesis on assumptions is an inconclusive result, not proof that they are correct.
- Although point and interval estimates (or their Bayes counterparts) should be the driving force behind most analysis, civil law cases should clearly be based on hypothesis testing, not upon interval estimation. The verdict boils down to a yes/no decision on a case. A simple example might be the yet unproven allegation that driverless cars have more fatalities per year owned than other cars. The standard legal “presumption of innocence” is equivalent to a null hypothesis that driverless cars have the same (or lower) fatality rates per year owned as contemporary cars with strictly active drivers. The endpoint in a class-action civil lawsuit might be fatal accidents per year owned. The jury must be presented with the results of this hypothesis test, with a P-value for a verdict determined by precedents (often below 0.05 one-sided) rejecting the null hypothesis. If a decision is made against driverless cars, then estimation might be used to determine damages. But the verdict rests strictly upon hypothesis testing. It should be noted that, often, there is no statistical parameter, and yet hypothesis testing can work. One example appears in Shuster and Mark Handler’s 2020 European Polygraphy article, “Trying an Accused Serial Sexual Harasser for Libel in a US Civil Court.”
- When designing large clinical trials of public health importance, we need to bear the above in mind and conduct “large simple randomized trials.” Susan Ellenberg and Mary Foulkes provide excellent justification for this when dealing with the treatment of AIDS in their 1994 Statistics in Medicine article titled “The Utility of Large, Simple Trials in the Evaluation of AIDS Treatment Strategies.”

Leave a Reply