Showing posts with label ANOVA. Show all posts
Showing posts with label ANOVA. Show all posts

Tuesday, October 11, 2016

Does Misuse of Statistical Significance Give Us a Higher Likelihood of Immortality?

This past week, a letter in Nature was published by the Department of Genetics at Albert Einstein College of Medicine in New York. The letter’s title was pithy and shocking “Evidence for a limit to human lifespan.” GASP! “But this can’t be true,” some may say, “with advances in medicine, we’re only continuing to lengthen the human lifespan!” First off, that’s what I think is a bit of a teleological argument – assuming just because humans have certain intelligences about the world, our purpose is to expand life or live forever. There’s some sort of moral purpose that is implied in that answer. Anyways, that’s not the point, but it sets off an interesting point of debate

Maybe we, humans, aren’t supposed to live forever. We forget that we are but specks on the lithograph of time; we’re babies! Only 200,000 years has the species Homo sapiens existed. On a 4.5 billion-year old planet. Many things have lived before us, and many things will live after us. Perhaps this sounds a bit nihilistic, but I was happy to hear these reports from these geneticists. Humans are single-handedly taking the gauntlet to ourselves and our planet, making sure life, in any form, is going to have a hard time existing on a planet where the temperature is rising almost 2 degrees Fahrenheit each year. So yes, I was overjoyed, ecstatic, and relieved to hear we may be shortening the execution of our planet Earth -- until I heard these geneticists used statistics to come to their results. Oh, great, another potential mishap with statistics by those not formally trained in statistics, I thought.

The methods that the study employed were actually pretty simple. Most people who took a basic stats class could probably understand everything up until the cubic smoothing splines. The authors plotted the maximum reported age at death (MRAD) for 534 people over the age of 110 or what they call “supercentenarians” gathered from the International Database on Longevity. These supercentarians came from Japan, France, US and the UK. Two linear regressions were done for subsets of data between 1968-1994 and 1995-2006. The scientists found an increase in MRAD of 0.15 per year before 1995 (with r=0.68 and p=0.0007) and a decrease in MRAD of 0.28 years from 1995-2006. There was a r=-0.35 and a p-value of p=0.27 for the decreasing MRAD. The scientists didn’t expand on the statistical shortcomings of their decreasing MRAD points. The scientists conducted the same procedure with other data points, and found a similar overall trend – that is, a significant increase until breakpoint and an insignificant decrease after breakpoint -- but did not discuss the weak correlation or significance.

The initial criticisms of the work have been predictable, to say the least. Why didn’t the scientists discuss their dismal p-values? Moreover, some say that seeing a significant increase in MRAD then observing a decrease, even if it was insignificant, is still something. This is, however, an amateur cop out, in my opinion. There are so many more reasons to be critical of the work than just the terrible p-values, and it starts with study design.

The main figure in Dong et. al. and the bane of my existence these last four days...

First, it’s really not clear why the scientists decided to use the arbitrary breakpoint between 1994 and 1995. Clarification of this would be helpful, if not crucial, to understand, as the entire crux of the paper’s argument leans on this breakpoint and its subsequent data analysis (Figure 2a). It appears that the breakpoint was chosen arbitrarily to support their initial hypothetical claim that humans have reached a plateau in age advancement, and the linear regression decrease model in MRAD is used in a rhetorical sense, rather than in a strict statistical-sense of proving their ad-hoc null-hypothesis statistical test (NHST). Is it even right to use NHST here? I’d argue no. If their alternative hypothesis was that r=0, it’s increasingly hard to test effect sizes closest to zero – you probably need to have a larger sample size, and with this data, data that exists on the margins of collectability, that’s hard to do.

The authors do make an implied comment on the power of their statistical tests, noting that they probably don’t have a large enough sample size to make case for a more robust statistical model. To alleviate this, they applied a post-hoc sample expansion to include several RADs as the highest point, including the MRAD and the second through fifth RAD; this appears to be more or less a sneaky way of doing away with outliers to me. They concluded that the average RADs had not increased since 1968, and that “all series showed the same pattern as the MRAD.” But instead of applying an ANOVA to test their hypothesis about average RADs not affecting the means, they never reference any supplemental material. It appears that they just eyed it. Another suggestion would be to compare the slopes of the “plateau” regions by doing some sort of t-test.


In the end, the statistical procedures taken to prove their own point would have any data scientist’s head spinning; probably enough so that people just wouldn’t take a hard look at it for too long. This is dangerous, but we can’t say Nature hasn’t committed the crime before: many times, the journal has published funky stats simply because the title of the study was provocative (as demonstrated by our class this semester). Nonetheless, you have to tip your cap at the scientists who wanted to make a headline, they surely did it, then perhaps bonk them on the head with your cap and tell them to do better stats before they make such grand claims.

Sunday, April 24, 2016

John W.Tukey - Playing in Everyone's Backyard



John W. Tukey, a chemist, mathematician, educator, consultant, researcher, data analyst, and statistician, was born in New Bedford, MA in 1915, the only child of Adah M. Tasker and Ralph H. Tukey. It was evident from an early age that Tukey had great potential. His parents discovered that he had learned to read at age three after he informed them about a bridge closure that he had read about in the newspaper. Tukey’s parents were both trained as high school teachers, and his mother educated him at home until college, when he attended Brown University and received both his bachelors and masters degrees in chemistry in 1936 and 1937, respectively. Tukey then entered Princeton University with the intention of studying chemistry, but instead earned his Ph.D. in mathematics in just two years. His thesis entitled “On Denumerability in Topology” was submitted in 1939 under the supervision of the distinguished mathematician Solomon Lefschetz. After receiving his doctorate, Tukey became an instructor of mathematics at Princeton. Shortly thereafter, however, the second World War began and changed the trajectory of Tukey’s career.
Tukey served as a consultant then as the assistant director of the Fire Control Research Office (FCRO), one of the departments Princeton organized to help support the goals of the National Defense Research Committee during U.S. involvement in WWII. The FCRO studied the mechanics of artillery weapons in order to improve their precision and efficacy. Here, Tukey worked closely with physicists and engineers as well as mathematicians, including an engineer with a Ph.D. in physiology who focused on statistics named Charles Winsor. Tukey later explained that, “It was Charlie and the experience of working on the analysis of real data that converted me to statistics.” By the end of the war, Tukey described himself as a statistician, a field that he viewed not as an offshoot of mathematics but as a scientific discipline equivalent to chemistry, biology, or physics. Tukey maintained his faculty position at Princeton, helping to establish and direct its first statistics department. Following the war, he also began work at Bell Telephone Laboratories as a member of the technical staff, where he contributed to research in numerous diverse areas over the course of his career, eventually serving as the Associate Executive Director of Research Information Sciences.
Harvard professor Frederick Mosteller is credited with saying that Tukey “probably made more original contributions to statistics than anyone else since World War II,” an assertion that seems credible when looking at a list of Tukey’s accomplishments. To the field of statistics, Tukey contributed a method of determining confidence intervals used in ANOVAs, the fast fourier transform, and a series of methods of data depiction that includes the stem-and-leaf diagram and box-and-whisker plots, among many other works. He also contributed substantially to paradigms of statistical analysis, promoting the utilization of and distinction between “exploratory data analysis” and statistical hypothesis-driven “confirmatory data analysis,” and he was a vocal advocate for statistical rigor in experimental design.
John Tukey’s accomplishments were recognized with numerous accolades including the National Medal of Science awarded by President Nixon in 1973 and the IEEE Medal of Honor in 1982 for his research in the development of the fast Fourier transform (FFT) algorithm. Aside from his work in academia and at Bell Laboratories, he made an impact in many other areas, for, as Tukey put it, “The best thing about being a statistician is that you get to play in everyone’s backyard.” The NY Times’ obituary of Tukey credits him with coining the term “bit” in reference to computer data and the word “software” in 1958. He also served in a number of public capacities, including chairing a committee investigating ozone depletion in the 70s, serving as a U.S. delegate to nuclear discontinuance conferences, and serving on the President’s Science Advisory Committee.
John Tukey retired in 1985 at the age of 70, and he passed away in 2000. He aspired to be a scientific generalist, and he achieved this, leaving a broad legacy of advancement in science and technology.


Another good reference:
Brillinger, D.R. (2002). John W. Tukey: His life and professional contributions. The Annals of Statistics. 30(6): 1535-1575.

Thursday, April 14, 2016

One Test Does Not Fit All





In 2007, Toscano, et al. published the article “Differential glycosylation of TH1, TH2 and TH-17 effector cells selectively regulates susceptibility to cell death” in Nature Immunology. This study reported that some T helper cell subsets (Th1 and Th17) were susceptible to galectin-1-mediated anti-inflammatory regulation and others were resistant (Th2) due to differing surface glycosylation patterns. The article contains eight multi-pane figures, one table, and six supplementary figures/tables and utilizes at least nine different experimental procedures, and the authors report that their statistical testing consisted entirely of Mann-Whitney U-tests.

The Mann-Whitney U-test is a non-parametric analysis that tests the null hypothesis that values from two groups derive from the same population. It is only appropriate for statistical comparisons of two groups in which all values are independent, and technically, it is best applied when the values do not conform to a normal distribution. Even ignoring the last technicality, there are very few instances in this paper in which the Mann-Whitney U-test was correctly applied. The most fundamental statistical errors are outlined below.


More than two groups compared

The authors performed experiments comparing properties of three different groups of T cells, a design in which a one-way ANOVA would have been an appropriate test, but they only indicate significant comparisons between two of the three groups, suggesting that they either ignored one group in statistical testing or that they performed multiple Mann-Whitney tests within each three-group experiment(Fig. 1, 2, 3b-e, 4b, 6b, 6d, Sup3, Sup5b, Sup6).


Comparisons of groups with more than one explanatory variable

Several experiments compare a variable in these three groups of cells over time, with increasing dose, or under three different treatments, conditions that require a two-way ANOVA(Fig. 1b-e, 3d, 3e, 4, 5a, 5b, 6c, 7c, Sup4). The images below depict perhaps the most egregious examples of this.


Non-independent samples
Each experiment with human cells was conducted using a sample from a single human donor split into three groups. Clearly then, the cells in each group are not independent and require a statistical test that accounts for repeated measures to be appropriately analyzed. Similarly, the authors frequently state that their reported data represent the mean of several experiment replicates. In these cases, each replicate would need to be considered paired for the purposes of statistical analyses, and the Mann-Whitney test is once again inappropriate (all figures).