Showing posts with label repeated measures. Show all posts
Showing posts with label repeated measures. Show all posts

Tuesday, May 3, 2016

P is for Polygraph


Polygraphs, also referred to as lie detectors, seem to be part of a social experiment that never ends. Though many of us have heard of polygraphs (especially given their widespread presence in old detective shows), it was not until recently that I looked into the experimental design behind the polygraphs, and the highly debated believability of their results.
First, a little background on polygraphs… The polygraph was invented in 1921 by John Augustus Larson, a 21 year-old physiology student at University of California, Berkeley who moonlighted as a police officer. The device, which Larson referred to as the cardio-pneumo psychogram, could measure blood pressure, pulse, respiration, and skin conductivity, all with the premise that these are measurable indicators that change when the subject of the tests is being deceptive.
Like any good experiment, polygraphs use controls – control questions, that is. A series of control questions, usually broad statements meant to inquire into the subject’s past character and truthfulness, are asked by the administer of the test, and then a series of relevant questions concerning the case at hand are asked (a repeated-measures design, if you will). If a subject is being deceptive, or perhaps rather, is being anxious or nervous, the polygraph should record a change in the physiological measures. Someone who is being deceptive concerning the case should theoretically be more anxious when asked to answer relevant questions, while someone who is not being deceptive concerning the case should be more anxious when answering control questions.
Though the polygraph has been regarded as one of the greatest inventions (in fact, the original polygraph device constructed by Larson is housed in the Smithsonian), its validity has been debated since its invention. Larson conducted many tests on his own, compiling a list of cases where his device helped solve murders, thefts, etc. (interestingly, Larson first tested his device on his wife…). However, many scientists today regard the basis of polygraphs as pseudoscience. Former US Attorney General John Ashcroft estimated the false-positive (or type I error, if you will) rate of polygraphs to be 15 percent. Yet, the Journal of General Pyschology published an analysis of 41 criminal cases where polygraph tests were used, concluding an accuracy over 90 percent.
If the test seems to be so accurate, then why the ongoing debate (as with all science, right)? Well, studies are conflicting in their results of accuracy, which could be a result of “p-hacking”, where researchers throw out inconclusive polygraph results to improve the accuracy rate of polygraphs. One also has to question the experimental design of the polygraph, where neither a subject or administer is blinded. One could potentially skew the polygraph results from either perspective by causing undue nervousness or anxiety during questioning, or by being aware of how the measured indicators change and resisting (common spy trick). The best summary of the flaws of the polygraph test come from William Iacono, a psychophysiologist at the University of Minnesota: "A big problem is that it's not really a test of anything,” highlighting the fact that little is really known about how the body behaves when lying, and therefore these physical measurements may not actually be measuring marks of deception.
Needless to say, given all of the debate, I would not want to be on the wrong side of a polygraph in a courtroom.
Sources:
1.     “Telling the Truth About Lie Detectors.” USA Today. http://usatoday30.usatoday.com/news/nation/2002-09-09-lie_x.htm
2.     “The Truth About Lie Detectors.” American Physiological Association. http://www.apa.org/research/action/polygraph.aspx

Thursday, April 14, 2016

One Test Does Not Fit All





In 2007, Toscano, et al. published the article “Differential glycosylation of TH1, TH2 and TH-17 effector cells selectively regulates susceptibility to cell death” in Nature Immunology. This study reported that some T helper cell subsets (Th1 and Th17) were susceptible to galectin-1-mediated anti-inflammatory regulation and others were resistant (Th2) due to differing surface glycosylation patterns. The article contains eight multi-pane figures, one table, and six supplementary figures/tables and utilizes at least nine different experimental procedures, and the authors report that their statistical testing consisted entirely of Mann-Whitney U-tests.

The Mann-Whitney U-test is a non-parametric analysis that tests the null hypothesis that values from two groups derive from the same population. It is only appropriate for statistical comparisons of two groups in which all values are independent, and technically, it is best applied when the values do not conform to a normal distribution. Even ignoring the last technicality, there are very few instances in this paper in which the Mann-Whitney U-test was correctly applied. The most fundamental statistical errors are outlined below.


More than two groups compared

The authors performed experiments comparing properties of three different groups of T cells, a design in which a one-way ANOVA would have been an appropriate test, but they only indicate significant comparisons between two of the three groups, suggesting that they either ignored one group in statistical testing or that they performed multiple Mann-Whitney tests within each three-group experiment(Fig. 1, 2, 3b-e, 4b, 6b, 6d, Sup3, Sup5b, Sup6).


Comparisons of groups with more than one explanatory variable

Several experiments compare a variable in these three groups of cells over time, with increasing dose, or under three different treatments, conditions that require a two-way ANOVA(Fig. 1b-e, 3d, 3e, 4, 5a, 5b, 6c, 7c, Sup4). The images below depict perhaps the most egregious examples of this.


Non-independent samples
Each experiment with human cells was conducted using a sample from a single human donor split into three groups. Clearly then, the cells in each group are not independent and require a statistical test that accounts for repeated measures to be appropriately analyzed. Similarly, the authors frequently state that their reported data represent the mean of several experiment replicates. In these cases, each replicate would need to be considered paired for the purposes of statistical analyses, and the Mann-Whitney test is once again inappropriate (all figures).