Showing posts with label #badstats. Show all posts
Showing posts with label #badstats. Show all posts

Friday, May 6, 2016

Grad School Rankings...What do they mean??

In light of the recent post regarding ranking competitiveness of the UAE, I started turning over the idea of rankings in my head.  We all value rankings, whether we admit it or not.  For one, rankings help us make decisions.  I'm sure a number of us peeked at the US News and World Report ranking biomedical research programs before selecting Emory as our home.  From personal experience, I found one of the prime emphases in grant writing class to be citing the number of F31 grants awarded to Emory GDBBS students (apparently we are currently ranked program number 2 in the US, and at one point last year, we were in first place). Oddly enough, despite our high ranking in terms of NRSAs, we are only ranked number 30 in terms of biological science graduate programs, according to the US News and World Report website. How could this be?  And what does statistics have to say about this?

How Emory Stacks Up...maybe.


Indeed, some rankings are based purely on objective, raw numbers, such as the NRSA statistic. The student either recieved F31 funding, or they did not. Others, such as the US News and World report rankings, and the UAE competitiveness rankings, are based on an amalgamation of a number of factors, including some that are subjective.  I did a bit of digging to figure out what the US News and World Report numbers are based on. Perusing their website left me with more questions than answers.  Any information given on how the rankings were determined is murky at best.  The most data I could find on how they rank describes their methodology for their undergraduate ranking system (the original). As one article from their own website investigated this topic correctly points out, "The host of intangibles that makes up the [college] experience can't be measured by a series of data points."  Factors such as reputation, selectivity and student retention are cited as some of the data points that determine ranking.  In terms of quantification, however, I did not find any clear answer.  The website cites a "Carnegie Classification" as the main method for determining rank. 
The Carnegie Classification was originally published in 1973, and subsequently updated in 1976, 1987, 1994, 2000, 2005, 2010, and 2015 to reflect changes among colleges and universities. This framework has been widely used in the study of higher education, both as a way to represent and control for institutional differences, and also in the design of research studies to ensure adequate representation of sampled institutions, students, or faculty.

As you can probably gather, the quantification involved in the Carnegie Classification is never really defined and left me wondering if there is some sort of conspiracy underlying these rankings.  Reading about these rankings has left me feeling what I imagine a nonscientist feels like when they try to understand scientific data from reading pop culture articles.  Still, we continue to value these rankings, even if we have no idea what they really mean.  Yet again, we revisit the theme that there is a strong need, not only for statistical literacy, but for statistical transparency. Statistical analysis needs to be clearly laid out so that the layperson can fully appreciate the true value of a ranking.

I also wonder about how quantitative and qualitative factors that are apparently used in the Carnegie Classification are combined together to determine one final ranking value.  As we have learned in class, continuous and categorical variables simply do not mix.  Yet here and most everywhere, we see them being combined. Maybe it is beyond my level of statistical comprehension, but I wonder if there is a way to correctly combine the two?

Tuesday, April 19, 2016

Main effects vs. interaction

In order to find an example of bad statistical analysis used in the field of behavioral neuroscience, I turned to an article published in Nature Neuroscience (2011) by Nieuwenhuis, Forstmann & Wagenmakers. This article published the most prevalent incorrect statistical analyses observed in Science, Nature, Nature Neuroscience, Neuron and The Journal of Neuroscience. Here, I will provide an example of one of the more prevalent statistical errors in published papers - one that involves comparing p-values  between two groups to show that one group that received treatment is significantly different, whereas the other is not. 

For example, "Optogenetic photoinhibition of the locus coeruleus decreased the amplitude of the target-evoked P3 potential in virally transduced animals (P = 0.012), but not in control animals (P = 0.3)."
 

Recently we were taught about how two use a two-way repeated measures ANOVA to assess main effects. In this example, one should have performed this analysis and then looked to see if there was a significant interaction between the two factors here: viral group and photoinhibition status. However, this example of bad statistics, instead, compares p-values representative of main effects, only. Here, P=0.012 represents the asterisk above the bar graphs showing reduced P3 amplitude in virally transduced mice. This asterisk does NOT indicate that the difference is a significant interaction. In other words, it does not indicate that P3 amplitude is significantly reduced by photo inhibition in the virally transduced mice, and not the control mice.


P-hacking, and proud of it

I did not pick up this paper in the hopes of using as a BadStats example, which makes the improper statistics all the more disappointing. The behavioral output used by Gaskin et al. (2010) was an investigation ratio (IR) between exploration times of repeated and novel objects. This metric made their abuse of statistical design personal, because I use the exact same behavioral output in my experiments. The premise of the study was sound, and was developed in response to an earlier, poorly designed study.  Albasser et al. (2009) claimed to find a relationship between the amount of time a rat explores an object and later object memory. These results were pooled from subjects in different experimental conditions, and therefore the correlation they ran encompassed multiple independent variables. Gaskin et al. (2010) strove to experimentally test the relationship between exploration duration and memory strength.

The experimental setup was fairly sound. In total, the team performed 4 experiments to assess any correlation between exploration and memory, and then manipulated exploration time to determine if a causal role existed. Problems came in most notably in experiments 3 and 4, in which the authors compared the memory of control and dorsal hippocampus-lesioned rats from a novel object preference task. Importantly, the only difference between experiments 3 and 4 were the retention intervals between the study session (in which rats viewed two identical copies of an object) and the multiple test sessions (in which rats viewed a copy of the object from the study phase and a novel object). Experiment 3 used a 2-hour delay for the first test, then 24 hour intervals for the next 4 days, and experiment 4 separated all tests by 35 seconds. Not surprisingly, the data from the two experiments were presented in identical fashion:



Really the only way to tell which is which is the labelling of Days (experiment 3) versus Tests (experiment 4). Even with the identical design, their analyses smack of p-hacking. To their credit, they are fully transparent about it.

In experiment 3, analyses of IRs included:
1)      One-sample t-tests for both groups at every time point to determine if IRs were above chance (0.5) – found significant differences
2)      Independent sample t-tests to compare IRs between control and lesioned groups at every time point – found significant differences
3)      One-way RM ANOVA (test days as repeated measure) run separately for control and lesion groups – they did report a significant F-value in the control group, but claim it “was only due to a significant difference between the IRs obtained during Days 2 and 3”, and therefore dismiss it
Alright. Got that?  Here were the analyses for experiment 4:
1)      One-sample t-tests for both groups at every time point to determine if IRs were above chance (0.5) – found significant differences
2)      Independent sample t-tests to compare IRs between control and lesioned groups at the first time point – found significant differences
3)      One-way RM ANOVA on both groups – showing no effect of test session on either group.
4)      Two-way mixed ANOVA by group and test session – found a trend toward group effect and a significant effect of test session.
a.       Post hoc comparisons of IRs between groups – found significant differences

By the end, I was confused.  Why run a two-way ANOVA in experiment 4 and not 3?  Why also run a one-way ANOVA in experiment 4?  Why run separate one-way ANOVAs on the two groups that end up being compared by t-tests? Why run independent t-tests and an ANOVA for the same data in experiment 3?

The authors set up the paper by saying that differences in IR between groups are not behaviorally relevant - only differences from chance.  That being the case, the independent sample t-tests are improper. A mixed factor two-way ANOVA does address changes over time and differences between groups, and their use of a two-way ANOVA in experiment 4 was correct.  Adding on one-way ANOVAs to the same data, however, was not.

There is no report of the raw data, so I cannot reevaluate their results.  However, the varied analyses and redundant tests (with varied appropriateness) make me feel like I was sitting in on a graduate student running every test they could think of to drive that p-value down.  I commend the authors for reporting them all, but transparency does not make biased, exploratory analyses good practice.

References

Limits of Detection (and Credibility)

Methods of Effective Conjugation of Antigens to Nanoparticles as Non-Inflammatory Vaccine Carriers

Nanoparticles are promising carriers for recombinant vaccine antigens, although the physical properties that make some particles good vaccines and others not are still unknown. In this work[1], the authors conjugate model ovalbumin (OVA) antigen to polystyrene nanoparticles for use as a vaccine delivery system. They immunize mice and examine the in vivo responses in order to elucidate an immunological mechanism, however their poor statistical analysis of the resulting data leaves little in the way of conclusions.

Before I get to that, another issue I had with their data was an odd regression curve used to determine concentration from absorbance (Figure 3). In the adjacent figure, the authors used what appear to be serial dilutions of nanoparticles to generate a regression between optical density (OD) at 248 nm and the number of nanoparticles per mL. While their R2 value is especially high, as would be expected for a serial dilution, there are no error bars on the points or indications of the number of replicates performed. More seriously, however, is the regression equation they use. In drawing a correlation between the number of nanoparticles in units of 1013/mL and the optical density, both numbers with one significant figure, it seemed odd that their regression equation would have an intercept with 4 significant figures, leading me to question the value of their equation and its corresponding R2 value. Additionally, the fact that serial dilutions of a solution would have linearly decreasing absorbance values is not necessarily novel or informative, and this figure could have been put in the supplemental information.
More statistically egregious, however, is their table of multiplexed, cytokine-bead array results, expressed as pg/mL (Tables 3 and 4):

Looking only at the first column, one wonders what the difference between 1.4 ± 1.9, or 0 ± 0.21 and “Not detectable” is. Given the absurd concentrations seen in naïve serum and serum from mice injected with NPs, it’s clear that the data is very widely skewed, in which case actually seeing the data points in a graph would be much more helpful. It is also suspicious that these values were the only data in the paper displayed in tabular form, whereas all others were either histograms or bar graphs of some sort. No mention of the number of mice used for this analysis was found in the figure caption, or the results or methods sections either. Not only are their conclusions on immunological mechanism weak, but their poor statistical analysis calls the significance of other data they present into question as well.


[1] Xiang SD, Wilson K, Day S, Fuchsberger M, Plebanski M. Methods of effective conjugation of antigens to nanoparticles as non-inflammatory vaccine carriers. Methods 2013;60:232-41.

Monday, April 18, 2016

Paired vs. Unpaired.


Let’s say you’ve invented a weight loss pill and you’re ready to show the world what that this pill is the real deal.  You start to think up some experiments that will showcase your work with human patient trials and you’re deciding between unpaired or paired designs.  Which should you choose?  Historically, in these situations, the majority of experimenters would probably lean towards a paired design.  The ability to detect an effect of treatment within an individual over time or in a new context has usually been considered a “more impressive” result.  It makes for a better story when you’re presenting the data.  Luckily, it’s hard to convince an audience that a study uses a paired design when that’s untrue, however sometimes experimenters might miss the fact that their own data is paired and opt for unpaired tests with confusing results.
Figure 1: Unpaired comparison in ER chaperone activity with and without sevoflurane.




In the paper “The Effect of Endoplasmic Reticulum Stress on Neurotoxicity Caused by Inhaled Anesthetics”, Komita et al. investigate a possible calcium-dependent mechanism for neuronal death following exposure to sevoflurane anesthetic.  Sevoflurane exposure causes calcium to be released from the endoplasmic reticulum (ER) of neurons into the cytosol.  The rapid change in calcium dynamics in the cell is thought to cause ER stress which may lead to apoptosis and underlie the neurotoxic effects that have been reported after anesthetic exposure.  To test their hypotheses about the effects of inhaled anesthetics, they chose to expose cultured neurons to sevoflurane while monitoring markers of ER function.  To analyze this data, the authors used immunofluorescence microscopy to compare ER chaperone levels with inert gas and sevoflurane gas.  Given that these experiments are done in cell culture, the authors could have chosen to use a paired design, but instead this test is run as an unpaired design (figure 1).
 


This may seem to be a simple, inconsequential mistake in comparison to the paper’s findings.  The effect is rather large and would have been noticed given either test, but the authors go on to make further mistakes regarding the treatment of their data.  The worst being their improper analysis of the data in Figure 2 as a 1-way ANOVA instead of a 2-way analysis (figure 2).  For the following figure, different genotypes of rats had their ER chaperone response tested in both control and sevoflurane conditions.  The study design is clearly deserving of a 2-way ANOVA since comparisons are being made between control and sevoflurane and all the genotypes.  However, the experimenters instead decide to treat all of these groups as unpaired and unrelated.  It’s unclear whether they ignored the difference in the genotype or treatment in order to justify analysis with one-way ANOVA, but they do it anyway.  It’s improper experimental design here, and leads to confusing results that seemed to be cherry-picked.



Komita M, Jin H, Aoe T. (2013). “The Effect of Endoplasmic Reticulum Stress on Neurotoxicity Caused by Inhaled Anesthetics”. 117(5): 1197-1204.


Friday, April 15, 2016

If you don't provide the details, that doesn't make it okay!



      In this paper, the authors investigate how a protein implicated in schizophrenia, dysbindin, regulates dendritic spine dynamics in hippocampal neurons. The authors decided to look at the connection between dysbindin and dendritic spines because there is a well documented spine pathology in schizophrenia. The authors also investigate the mechanism as to how dysbindin regulates dendritic spine dynamics by looking at Ca2+/calmodulin-dependent protein kinase II (CaMKIIa) and Abl interactor 1 (Abi-1) levels. To investigate these aims, the authors utilized time-lapse imaging of the hippocampal neurons, immunoblots, and immunoprecipitation.
Despite being published in the reputable Journal of Neuroscience, this paper fails to comply with a major responsibility in publishing scientific results: stating how they analyzed their data. Though the authors put their sample size and p values under each figure, they never explicitly state what statistical tests they utilized. There is no mention of the statistical tests in the methods section. Since Journal of Neuroscience does not have a supplementary section, the analysis information cannot be hiding there either. It is crucial that readers know what statistical tests the researchers performed in order to be assured that they performed the correct tests to back up their claims. In Figure 1, the researchers wanted to test the hypothesis that the protein dysbindin regulates dendritic spine dynamics. They observed the dynamics of dendritic protrusions in neurons from WT mice and mice lacking dysbindin (sdy mice) by transfecting with a fluorescent protein construct and imaging a dendritic branch every minute for 30 minutes. They then defined whether an event was a formation, retraction, mushroom to filopodia conversion, filpodia to mushroom conversion, mushroom to stubby spine conversion, or stubby spine to mushroom conversion. Figures 1B,D,F demonstrate this categorization and assert that the dynamics in sdy mice significantly differed from WT mice. However, it is unclear if the authors correctly did a two-way ANOVA to find this significance or if they incorrectly analyzed the data with two separate t-tests.
Figures1,B,D,F

 A similar statistical question is presented with Figure 2B, where the authors quantify protrusion number per µm of mushroom, stubby, and filopodia spines for both WT and Sdy mice. Once again, this data should have been analyzed as a two-way ANOVA but appears as though it may have incorrectly analyzed with separate T-tests. Even in figures that are obviously two-way ANOVAs, such as Figure 2G and Figure 3B,D,F,H, the authors do not go into any detail regarding their multiple comparisons statistics.
Figure 2B

            Aside from the omission of specifics regarding statistical analysis, an aspect of experimental analysis was incorrect in Figure 7. In this figure, the authors sought to test the hypothesis that CaMKIIa activity is lowered in the sdy neurons due to inhibition by activity-dependent modulation of Abi interactor 1 (Abi1). The authors used western blots to determine the amount of Abi1, CaMKIIa, and dysbindin (as a control) present in WT versus sdy mice. Additionally, they did an immunoprecipitation for Abi from the P2 fractions of both genotypes. The authors present the quantification of these western blots as “normalized fold change”. However, the authors fail to define what they are normalized to. It is convention to normalize to actin levels (which are present on the western blots for whole cell lysate and P2 fraction), but the authors do not explicitly say they normalized their Abi1 and CaMKIIa levels to actin. What appears to be the biggest flaw in these sets of experiments is how they normalized their immunoprecipitation data. Typically, immunoprecipitation data is normalized to “input”, meaning that 10% of protein lysate that was used for the IP is usually run on western blot and used to normalize. However, it appears that the data for the immunoprecipitation in sdy mice is just normalized to what was immunoprecipitated in WT mice. This does not control for if there was initially more Abi1 or CaMKIIa in the input for sdy mice. If there is more protein inputted into an immunoprecipitation, then more will immunoprecipitated. Thus, it could be possible that more of the sdy sample was used for the immunoprecipitation than WT sample, which would only make it appear that CaMKIIa was increased. A loading control is NECESSARY to make any conclusion from this immunoprecipitation experiment.
Figure 7


            Though this was an interesting paper in demonstrating how dysbindin regulates spine dynamics, the authors failed to provide crucial information regarding their statistics as well as normalization procedures. Without this information, it is difficult to have full faith in their statistics and subsequent conclusions.

Wednesday, April 13, 2016

Repetitive T-Tests Instead of an ANOVA




Goal of the Paper: The goal of this paper was to determine the change in number and activation status of peripheral blood T-cell subsets during two blood-stage infection models of malaria. One model involved two-short-course infections while the other model used a long-course infection. The investigators were interested to know if T-cell subsets changes within each group over time and if these changes were different between the different infection models at the same time point.

Experimental Design: Three animals were assigned to each group for a total of six animals in the cohort. An initial pre-infection sample was collected from each animal and was used as the baseline value for each macaque. Samples were then taken at specific time points after inoculation for analysis by CBC and flow cytometry. Comparisons were then performed to compare the data from baseline with other points in the infection within infection model and to determine if there were differences between the infection models at a specific time point.

Critiques of the Paper:
1.  Repetitive t-tests to compare within group and between groups whenever the experimental design reflects the need for performing a two-way ANOVA with an appropriate post-hoc analysis.


The most egregious error in this paper is the use of t-tests to compare between group and within group based on this experimental design. Figure 3 is pasted above for reference and confirms this was the approach used by the authors. According to the methods, a Student’s t-test was used to assess if there were differences between groups at different time points, and a paired t-test was used to determine if there were significant changes within group compared to the baseline value. Conceptually, the authors knew that between group analyses did not need a paired analysis and that within group analyses did. However, the approach that was used was incorrect. By performing repetitive t-test, the authors inflated their type-I error well above the established threshold of 0.05, and given the sample size of the study, most of the statistically significant results are likely invalid.
            The experimental design calls for a two-way, repeated measures ANOVA with an appropriate post-hoc to address the objective. The approach that the authors should have taken is to perform an initial two-way, repeated measures ANOVA with group (i.e. infection model) and time point as the two factors. After performing this analysis, the results would inform if there were significant changes based on subject, group, time, and if there was an interaction-effect occurring. If significant, the next step would have been to perform a post-hoc analysis, and in the case of this experiment that appears to be underpowered with only 3 animals per group, only specific planned comparisons should be performed to conserve alpha. Using an unplanned comparison approach would be unwise because it would likely be too underpowered to identify any significant differences, especially if a pairwise analysis was performed for every possible combination.

2.     Figures are poorly designed and do not clearly indicate the relevant information, and the captions are confusing and unclear.

The figures in this analysis clearly indicate that individuals are being followed over time (see figure 3 above). This is appropriate representation of the data, but unfortunately, the other aspects of the figure are lacking. For instance, the arrows on the figure indicate inoculation and drug treatment. One group had a different inoculation and drug treatment regimen than the other, and thus, displaying the data on the same graph is bad data presentation. Further, the repetitive t-tests lead the authors to use a strange convention of denoting the statistical significance between a time point and a baseline value. With an appropriate two-way, repeated measures ANOVA this could have been rectified. Overall, I would likely suggest that the data be graphed separately based on group and a table be generated to show significant differences in outcome variables between groups to make it clearer and more effective for the reviewer/reader.

3.     Biological conclusions should be questioned as treatment could be considered a confounding/third variable.

One of the goals of the study was to determine if there were differences in T-cell responses between groups. Indeed, the two-way ANOVA that I mentioned above would answer this question in the most appropriate statistical manner based on the experimental design. However in that approach, the assumption is that drug intervention and re-inoculation have no effect on the T-cell values. Based on my experience with these drugs and this model, I would say that this is a fair assumption. However, it is worth recognizing this aspect of the design and understanding the appropriate statistics should that assumption not be made. If the drug intervention was added into the current Two-Way ANOVA approach, this would for a three-way ANOVA. As we have learned in the course, it is virtually impossible to interpret the results of a three-way ANOVA because of the complexity of the experimental design and, thus, the null hypothesis. Therefor if treatment was going to be a factor, a linear or nonlinear model would likely be needed to determine the effect between groups, within group over time, and if there was an effect of treatment, and if there were three-way interactions between the different factors.

Overall Conclusion: This paper does not use appropriate statistics, has poor data representation, figure captions, and graphs, and the overall assumptions made should be questioned. Additionally the fact that it is likely severely underpowered, the conclusions drawn are likely erroneous and could largely be false-positives with the specific statistical approach that was taken. Finally, I suspect the authors were “p-hacking” to achieve significance and that is why they went with the t-tests and not the ANOVA analysis that the experiment calls for.