Showing posts with label The rest of statistics. Show all posts
Showing posts with label The rest of statistics. Show all posts

Tuesday, April 12, 2016

The Trees for the Forest


            I picked up my first pipette in 2009 in a virology lab that focused on a single domain of a single protein of a single virus. The majority of our assays were in vitro, and even cell culture was used sparingly. We essentially studied a single cog of a machine outside the context of that machine, and this was standard practice in the field.
            Today I am sifting through massive piles of RNA sequencing, proteomics, and metabolomics data in an effort to see the machine more clearly. As our high throughput and analytical methods have progressed, scientific research has gained the ability to investigate not only the intricacies of biology, but also the opportunity to examine the bigger picture. With that opportunity, however, come significant challenges.  We can spend months chasing down artifacts or trying to piece together seemingly contradictory data with little success. Statistics is our only guide in this process, and many of us have only rudimentary training in the nuances of analysis necessary to navigate these mountains of data.
            The days of “your favorite protein” may be coming to an end, replaced by a more complex and wholistic approach at the bench. Amin et. al. describe this new approach in an excellent article about systems biology and the rise of big data. They summarize the “omics” approach in the figure below, and argue that research must move away from any one aspect of biology and begin to integrate all approaches into the investigation of a scientific question.

            With this more complex approach comes a need for more advanced statistics in graduate and even undergraduate education. Big data has become an integral part of our scientific lives and increasingly an important aspect of our personal lives as well. In the midst of this contested presidential election, popular interest in big data is incredibly high, and interested onlookers would do well to familiarize themselves with the art and science of statistics.

Meta-Analysis of Meta-Analyses

Reading through the chapters of Motulsky’s Intuitive Biostatistics, my attention was caught by the section on meta-analyses in Chapter 43. Given the problem of reproducibility in scientific research, the meta-analysis seems to be one of the great tools for addressing the problem of reproducibility; confirming or refuting prevailing scientific theories and advancing humanity’s body of scientific knowledge.

Looking at the “Assumptions of Meta-Analysis”, I was intrigued by the idea that meta-analyses fall into two general categories, each based on a different assumption. Either

(A)    All subjects are sampled from one large population; each scientific study is estimating the same effect. Measurement error comes from random selection of subjects, or
(B)    Each study population is unique, and the differences in population and random selection of subjects both contribute to the error.

Sound familiar? To me, these two assumptions seemed derived from the two philosophies of scientific realism we discussed at the beginning of class. One being that the truth (the large population) exists somewhere, and that science’s job is to uncover it, the other being that unobservable truth is irrelevant, and that utility of knowledge is paramount as is relates to advancement of medicine or technology. The idea that a population’s true response to a therapeutic exists is reflected in assumption (A), whereas (B) describes the anti-realist philosophy that there is no “global” population response, only individual subset responses as described in each sub-study of the meta-analysis.


The value of the meta-analysis under the framework of (B), then, would be to predict the efficacy of a given therapeutic in the next population of patients, given all those who have been tested before. Motulsky goes on to described this second model as the more commonly used one underlying most meta-analyses. The anti-realist philosophy is commonly associated with being applied and utilitarian, though I’m wondering if there’s a fundamental application of the anti-realism paradigm. Specifically, what does the idea that each sub-population in a meta-analysis is inherently disconnected from the others mean for drawing scientific conclusions from a meta-analysis? Is there a connection to the scientific philosophies of confirmation and falsification inherent in the above assumptions? Do the answers to these questions even affect the conclusions we can draw from meta-analyses, or are they irrelevant exercises in navel-gazing?

A first look at meta-analyses

I’m going to confess that I had never given meta-analysis much thought before reading this section of Motulsky. I can also confess that I’m really, really glad this has never applied to my research so far.
The subject itself is conflicting. Motulsky introduces meta-analysis as a way to combine evidence from multiple studies, usually clinical trials that are used to determine the effectiveness of some therapeutic. At first read, it doesn’t sound like the worst idea—the fact is, not every study has the resources to follow and collect the thousands of patient samples necessary to determine efficacy of a treatment. Pooling many, well designed, smaller trials offers a solution for researchers interested in meta-analysis. It also offers a giant problem.
A quick google search of “meta-analysis and research bias” yields over one million results. It seems from reading this chapter, that there’s a good reason for this. Meta-analysis lends itself to publication bias, and not necessarily because people bloodthirsty, competitive, publish-or-perish, nightmare monsters. Not that this doesn’t happen, but honestly, performing a non-biased seems incredibly difficult. There’s a huge table of challenges in Motulsky’s chapter on it (seriously, check it out: p. 412 table 43.1). Some of the challenges involved in performing a meta-analysis include needing to seek out ALL relevant data, including unpublished studies and studies published in other languages.
One survey published in The BMJ states that of the 31 most recent met-analyses they examined only 9 included participant data from unpublished studies. Many of the studies included in this survey didn’t list any limitations to their analysis, which leads the authors of this particular survey to strongly caution reviewers when reading meta-analyses.
It seems like there are some ways to detect bias in your meta-analyses by the generation of funnel plots, where small sample bias can be seen as asymmetry on an axis:



As opposed to: 


But even this isn’t uncontentious.

At this point in my researching of meta-analysis, I’m genuinely glad that I have no interest in clinical efficacy of therapeutics. Honestly, I think I would stick to well written narrative review articles.  



Saturday, April 9, 2016

Nonparametric statistics in nonhuman primate research shouldn’t be taboo

My dissertation research has focused on defining the immune response of nonhuman primates infected with simian malaria parasites. One of the biggest challenges in nonhuman primate research is small sample sizes due to the cost of performing research utilizing these models. To put it into perspective, one monkey can range anywhere from $2,000 - $8,000 depending on the species and specific experiments that will be performed, and each animal costs $8 - $10 a day to house and feed. These costs add up quickly so researchers are required to limit the number of animals used, and in most cases, a monkey study consisting of anywhere from 3-7 animals is considered “well-powered” by the NHP research community. However as we have learned throughout the course and based on my experience with the heterogeneous responses of outbred models like NHPs, this sample size is typically not sufficient to rigorously and fairly test most scientific and statistical hypotheses. Further, most of the time the data does not meet the assumptions of a most parametric statistical tests, but most papers will use these test to gain significance, or in other words “p-hack”. This introduces bias into the nonhuman primate literature and provides a reason why many people are skeptical of NHP research. To fix this problem, there needs to be appropriate funding available for this NHP research to properly power and fairly assess the question being evaluated, or appropriate nonparametric statistical test should be used. However, this becomes difficult with small sample sizes as many nonparametric test require at least 5 subject to obtain a p-value of less than 0.05.



I have experienced the burn of an underpowered experiment that requires analysis by a nonparametric statistic first hand. Whenever I conducted one of the first experiments for my PhD, I had a result that was clearly significant based on the “bloody-obvious” test (see image to left). Prior to the experiment, I predicted the phenotype based on the literature and did a power calculation, which stated that I needed 5 animals to fairly test my hypothesis. My statistical hypothesis was that the mean parasite burden during a primary infection was different than the mean parasite burden during a relapse infection. Unfortunately during the experiment, one of my animals succumbed to the infection prior to having a relapse, which brought my sample size down to 4 animals. Whenever I performed a Wilcoxon matched pairs test (which I argue was the appropriate statistical test in this situation), I did not have enough data points to fairly test the hypothesis and got a P value of 0.0625 despite the phenotype; I should point out that this wasn't graphed correctly and should actually be connected by a line to imply that the analysis was paired. Whenever I presented this data, there was a huge debate on whether I had run the statistical test incorrectly and many thought, including PIs, I shouldn’t use nonparametric statistics even though this is clearly the appropriate test to perform. In the end, I succeeded in arguing my point and was able to report the phenotype in a publication even though we didn’t have significance because due to the death of one animal the study lacked the power needed to assess the data by nonparametric statistical methods. 
              Overall, I think that NHP researchers should embrace nonparametric statistics even though it may require more resources to generate significant data, but the benefit of producing reliable data that draws reproducible conclusions is key. Overall, the extra resources are well worth it even if it means that one has to do less experiments or hire one less technician. I think that “less is more” whenever it comes to science, particularly in the realm of NHP research. 

Tuesday, April 5, 2016

Statistical Agreement - Duh!

           I recently experienced a situation in which I was faced with what I believed to be a statistical impossibility of sorts (Granted I had yet to embark on TJ's biostats course at the time). This conundrum was exposed when attempting to establish equity of quantitation between two entirely dichotomous experimental methods aimed to provide similar readout metrics. Explicitly, the change in measurement from an isotope release assay to that of a single cell based analysis.  While ample statistical tests could be run internal to each individual experiment, cross comparisons between methods attempting to glean “significance” were determined to be illegitimate in theory due to the contrasting experimental designs.
            Briefly, cell mediated lympholysis (CML) has traditionally been enumerated by loading peripheral blood mononuclear (PBMC) target cell populations with the radioactive isotope 51Cr that diffuses into live cells and is released upon the induction of apoptosis and subsequent lysis as mediated by alloreactive cytotoxic lymphocyte (CTL) effector cells. The goal of this study was to develop an isotope-free approach to the quantification of CML in the context of transplantation immune monitoring, by moving from 51Cr release as the metric of target cell lysis to determining the total loss of healthy target cells as measured by flow cytometry. To further complicate matters, the combination of isotope decay and natural day-to-day variation of both the donor and recipient PBMC, along with minor histocompatibility antigen discrepancies prevents the normalization of individual experiments. While these differences may seem negligible to a third party - years of alloreactive CML assays run with chromium said otherwise leading to an adoption of the method to simply provide a binary "responsive" or "unresponsive" result ignoring magnitude all together. This aptly illustrates the reason for my ambition to establish a more modern approach.    
             In order to provide the journal reviewers with a visual representation of correlation I resorted to contacting my bosses go to biostatistician to try to make sense of how this could be done (Don't worry TJ I was still at MGH at the time). After discussing the matter with the biostatistician I was absolutely dumbfounded that I had not thought of graphing the correlation as a simple agreement between the two methods. Anyways, to accurately graph the data correlation we plotted the result from each method for a single experiment with a "perfect fit" line containing a slope of 1.0 to show gross deviation and relative magnitudes of agreement.   


Simpson's Paradox

Statistics to many is a non-intuitive, mathematical jungle fought through by hours of rigorous study and contemplation. Even to those that are well versed, there are still aspects of statistics that seem to defy logic. One phenomenon that illustrates this paradigm well is Simpson’s paradox, which is defined as a set of data where a trend observed in each of the individual groups disappears or reverses when the groups are combined. This paradox has been shown in data collected from baseball, clinical research, and sociology studies.

For example, let us consider a group of five individuals: Josh, Michael, Jessica, Faith and Marcus. Each of these individuals decided after watching the Nathan’s Famous hotdog eating competition that their new passion in life was to become a competitive eater. Over a ten-year period, each of the individuals competed in five competitions and enlisted the help of a statistician to analyze their progress. As you can see in the graphs below, through rigorous training each of the individuals managed to steadily increase the amount of hotdogs they could consume in a ten-minute period over the ten-year time period they were followed.
From this data, one may conclude that overall individuals will tend to increase the amount of hotdogs they can consume in a ten-minute period. This seems logical considering there is a strong upwards trend in each of the graphs above. When all of the data is grouped together, however, the opposite trend appears. As shown in the graph below, the grouped data from the five individuals clearly shows that the overall trend between hotdogs eaten in ten-minutes and age is negative.
 
This is Simpson’s paradox in action. The trend observed in each of the individual data sets reversed when these data sets were combined.
Personally, I believe this example, and many data sets in which Simpson’s paradox is observed, should raise some questions. Do these data belong in a grouped analysis? Are there confounding variables that can explain the trend in the grouped data relative to the individual data? Were the proper experimental and statistical protocols followed?
In this example, one could argue that the data do not belong in a grouped analysis since each individual began competitive eating at a different age. If the data were plotted with the x-axis as years after beginning competitive eating, the overall trend matches the individual data as seen below.
Overall, if data you are analyzing displays Simpson’s paradox, it may be wise to step back and analyze your data collection and analysis. Although the data may just display a paradoxical trend, it is quite possible that the trend arose from flawed collection or analysis.