Showing posts with label #p-hacking. Show all posts
Showing posts with label #p-hacking. Show all posts

Tuesday, May 3, 2016

P is for Polygraph


Polygraphs, also referred to as lie detectors, seem to be part of a social experiment that never ends. Though many of us have heard of polygraphs (especially given their widespread presence in old detective shows), it was not until recently that I looked into the experimental design behind the polygraphs, and the highly debated believability of their results.
First, a little background on polygraphs… The polygraph was invented in 1921 by John Augustus Larson, a 21 year-old physiology student at University of California, Berkeley who moonlighted as a police officer. The device, which Larson referred to as the cardio-pneumo psychogram, could measure blood pressure, pulse, respiration, and skin conductivity, all with the premise that these are measurable indicators that change when the subject of the tests is being deceptive.
Like any good experiment, polygraphs use controls – control questions, that is. A series of control questions, usually broad statements meant to inquire into the subject’s past character and truthfulness, are asked by the administer of the test, and then a series of relevant questions concerning the case at hand are asked (a repeated-measures design, if you will). If a subject is being deceptive, or perhaps rather, is being anxious or nervous, the polygraph should record a change in the physiological measures. Someone who is being deceptive concerning the case should theoretically be more anxious when asked to answer relevant questions, while someone who is not being deceptive concerning the case should be more anxious when answering control questions.
Though the polygraph has been regarded as one of the greatest inventions (in fact, the original polygraph device constructed by Larson is housed in the Smithsonian), its validity has been debated since its invention. Larson conducted many tests on his own, compiling a list of cases where his device helped solve murders, thefts, etc. (interestingly, Larson first tested his device on his wife…). However, many scientists today regard the basis of polygraphs as pseudoscience. Former US Attorney General John Ashcroft estimated the false-positive (or type I error, if you will) rate of polygraphs to be 15 percent. Yet, the Journal of General Pyschology published an analysis of 41 criminal cases where polygraph tests were used, concluding an accuracy over 90 percent.
If the test seems to be so accurate, then why the ongoing debate (as with all science, right)? Well, studies are conflicting in their results of accuracy, which could be a result of “p-hacking”, where researchers throw out inconclusive polygraph results to improve the accuracy rate of polygraphs. One also has to question the experimental design of the polygraph, where neither a subject or administer is blinded. One could potentially skew the polygraph results from either perspective by causing undue nervousness or anxiety during questioning, or by being aware of how the measured indicators change and resisting (common spy trick). The best summary of the flaws of the polygraph test come from William Iacono, a psychophysiologist at the University of Minnesota: "A big problem is that it's not really a test of anything,” highlighting the fact that little is really known about how the body behaves when lying, and therefore these physical measurements may not actually be measuring marks of deception.
Needless to say, given all of the debate, I would not want to be on the wrong side of a polygraph in a courtroom.
Sources:
1.     “Telling the Truth About Lie Detectors.” USA Today. http://usatoday30.usatoday.com/news/nation/2002-09-09-lie_x.htm
2.     “The Truth About Lie Detectors.” American Physiological Association. http://www.apa.org/research/action/polygraph.aspx

Tuesday, April 19, 2016

P-hacking, and proud of it

I did not pick up this paper in the hopes of using as a BadStats example, which makes the improper statistics all the more disappointing. The behavioral output used by Gaskin et al. (2010) was an investigation ratio (IR) between exploration times of repeated and novel objects. This metric made their abuse of statistical design personal, because I use the exact same behavioral output in my experiments. The premise of the study was sound, and was developed in response to an earlier, poorly designed study.  Albasser et al. (2009) claimed to find a relationship between the amount of time a rat explores an object and later object memory. These results were pooled from subjects in different experimental conditions, and therefore the correlation they ran encompassed multiple independent variables. Gaskin et al. (2010) strove to experimentally test the relationship between exploration duration and memory strength.

The experimental setup was fairly sound. In total, the team performed 4 experiments to assess any correlation between exploration and memory, and then manipulated exploration time to determine if a causal role existed. Problems came in most notably in experiments 3 and 4, in which the authors compared the memory of control and dorsal hippocampus-lesioned rats from a novel object preference task. Importantly, the only difference between experiments 3 and 4 were the retention intervals between the study session (in which rats viewed two identical copies of an object) and the multiple test sessions (in which rats viewed a copy of the object from the study phase and a novel object). Experiment 3 used a 2-hour delay for the first test, then 24 hour intervals for the next 4 days, and experiment 4 separated all tests by 35 seconds. Not surprisingly, the data from the two experiments were presented in identical fashion:



Really the only way to tell which is which is the labelling of Days (experiment 3) versus Tests (experiment 4). Even with the identical design, their analyses smack of p-hacking. To their credit, they are fully transparent about it.

In experiment 3, analyses of IRs included:
1)      One-sample t-tests for both groups at every time point to determine if IRs were above chance (0.5) – found significant differences
2)      Independent sample t-tests to compare IRs between control and lesioned groups at every time point – found significant differences
3)      One-way RM ANOVA (test days as repeated measure) run separately for control and lesion groups – they did report a significant F-value in the control group, but claim it “was only due to a significant difference between the IRs obtained during Days 2 and 3”, and therefore dismiss it
Alright. Got that?  Here were the analyses for experiment 4:
1)      One-sample t-tests for both groups at every time point to determine if IRs were above chance (0.5) – found significant differences
2)      Independent sample t-tests to compare IRs between control and lesioned groups at the first time point – found significant differences
3)      One-way RM ANOVA on both groups – showing no effect of test session on either group.
4)      Two-way mixed ANOVA by group and test session – found a trend toward group effect and a significant effect of test session.
a.       Post hoc comparisons of IRs between groups – found significant differences

By the end, I was confused.  Why run a two-way ANOVA in experiment 4 and not 3?  Why also run a one-way ANOVA in experiment 4?  Why run separate one-way ANOVAs on the two groups that end up being compared by t-tests? Why run independent t-tests and an ANOVA for the same data in experiment 3?

The authors set up the paper by saying that differences in IR between groups are not behaviorally relevant - only differences from chance.  That being the case, the independent sample t-tests are improper. A mixed factor two-way ANOVA does address changes over time and differences between groups, and their use of a two-way ANOVA in experiment 4 was correct.  Adding on one-way ANOVAs to the same data, however, was not.

There is no report of the raw data, so I cannot reevaluate their results.  However, the varied analyses and redundant tests (with varied appropriateness) make me feel like I was sitting in on a graduate student running every test they could think of to drive that p-value down.  I commend the authors for reporting them all, but transparency does not make biased, exploratory analyses good practice.

References

Wednesday, April 13, 2016

Repetitive T-Tests Instead of an ANOVA




Goal of the Paper: The goal of this paper was to determine the change in number and activation status of peripheral blood T-cell subsets during two blood-stage infection models of malaria. One model involved two-short-course infections while the other model used a long-course infection. The investigators were interested to know if T-cell subsets changes within each group over time and if these changes were different between the different infection models at the same time point.

Experimental Design: Three animals were assigned to each group for a total of six animals in the cohort. An initial pre-infection sample was collected from each animal and was used as the baseline value for each macaque. Samples were then taken at specific time points after inoculation for analysis by CBC and flow cytometry. Comparisons were then performed to compare the data from baseline with other points in the infection within infection model and to determine if there were differences between the infection models at a specific time point.

Critiques of the Paper:
1.  Repetitive t-tests to compare within group and between groups whenever the experimental design reflects the need for performing a two-way ANOVA with an appropriate post-hoc analysis.


The most egregious error in this paper is the use of t-tests to compare between group and within group based on this experimental design. Figure 3 is pasted above for reference and confirms this was the approach used by the authors. According to the methods, a Student’s t-test was used to assess if there were differences between groups at different time points, and a paired t-test was used to determine if there were significant changes within group compared to the baseline value. Conceptually, the authors knew that between group analyses did not need a paired analysis and that within group analyses did. However, the approach that was used was incorrect. By performing repetitive t-test, the authors inflated their type-I error well above the established threshold of 0.05, and given the sample size of the study, most of the statistically significant results are likely invalid.
            The experimental design calls for a two-way, repeated measures ANOVA with an appropriate post-hoc to address the objective. The approach that the authors should have taken is to perform an initial two-way, repeated measures ANOVA with group (i.e. infection model) and time point as the two factors. After performing this analysis, the results would inform if there were significant changes based on subject, group, time, and if there was an interaction-effect occurring. If significant, the next step would have been to perform a post-hoc analysis, and in the case of this experiment that appears to be underpowered with only 3 animals per group, only specific planned comparisons should be performed to conserve alpha. Using an unplanned comparison approach would be unwise because it would likely be too underpowered to identify any significant differences, especially if a pairwise analysis was performed for every possible combination.

2.     Figures are poorly designed and do not clearly indicate the relevant information, and the captions are confusing and unclear.

The figures in this analysis clearly indicate that individuals are being followed over time (see figure 3 above). This is appropriate representation of the data, but unfortunately, the other aspects of the figure are lacking. For instance, the arrows on the figure indicate inoculation and drug treatment. One group had a different inoculation and drug treatment regimen than the other, and thus, displaying the data on the same graph is bad data presentation. Further, the repetitive t-tests lead the authors to use a strange convention of denoting the statistical significance between a time point and a baseline value. With an appropriate two-way, repeated measures ANOVA this could have been rectified. Overall, I would likely suggest that the data be graphed separately based on group and a table be generated to show significant differences in outcome variables between groups to make it clearer and more effective for the reviewer/reader.

3.     Biological conclusions should be questioned as treatment could be considered a confounding/third variable.

One of the goals of the study was to determine if there were differences in T-cell responses between groups. Indeed, the two-way ANOVA that I mentioned above would answer this question in the most appropriate statistical manner based on the experimental design. However in that approach, the assumption is that drug intervention and re-inoculation have no effect on the T-cell values. Based on my experience with these drugs and this model, I would say that this is a fair assumption. However, it is worth recognizing this aspect of the design and understanding the appropriate statistics should that assumption not be made. If the drug intervention was added into the current Two-Way ANOVA approach, this would for a three-way ANOVA. As we have learned in the course, it is virtually impossible to interpret the results of a three-way ANOVA because of the complexity of the experimental design and, thus, the null hypothesis. Therefor if treatment was going to be a factor, a linear or nonlinear model would likely be needed to determine the effect between groups, within group over time, and if there was an effect of treatment, and if there were three-way interactions between the different factors.

Overall Conclusion: This paper does not use appropriate statistics, has poor data representation, figure captions, and graphs, and the overall assumptions made should be questioned. Additionally the fact that it is likely severely underpowered, the conclusions drawn are likely erroneous and could largely be false-positives with the specific statistical approach that was taken. Finally, I suspect the authors were “p-hacking” to achieve significance and that is why they went with the t-tests and not the ANOVA analysis that the experiment calls for.