Showing posts with label Statistical Tests. Show all posts
Showing posts with label Statistical Tests. Show all posts

Saturday, October 15, 2016

Take that Jenny McCarthy! Statistical Tests Show Improvement in Vaccination Completeness

It hasn’t been too long since celebrity Jenny McCarthy let it be known that she is vehemently opposed to our current vaccines. She was one of the first celebrities to support the pseudoscientific view that vaccinations cause autism, and recently she has made it clear that she thinks anyone carrying a virus is deathly “sick,” in her comments made about former co-star Charlie Sheen, who has HIV. Uh, that’s not exactly how the fields of virology and immunology have deduced the process, Jenny.

Nonetheless, public health officials will soldier on because they recognize the benefits of vaccinating people, especially children, and that such vaccination prevents sickness even in the case of contracting the virus. One question I’ve always had as a bench scientist is how is it that public health officials know they’re doing their job efficiently? I see many of my friends going to public health school wanting to help with the education arm of public health issues. How do we know if the methods they use are effective? What is a quantifiable measure for us to obtain a level of effectiveness?

Figure 1. An example of paired participant studies. In the case of the SUNY Upstate article I looked at the comparison group would be a group of similar age and income in an adjacent community compared with a group of interest receiving the intervention
Dually, I’ve enjoyed spending the semester reading about statistical tests we haven’t gone over in class. One of those tests is known as the McNemar’s test. A quick interwebs definition says McNemar’s test is a “statistical test on paired nominal data,” or basically assigning a binomial outcome to paired data (see Figure 1 for paired data example). When I first read about this, I thought of vaccines. A good signal to public health educators that their programs are working are whether populations are vaccinated or not, specifically communities that face traditional barriers to quality healthcare.

In a community health paper published in 2013, public health professionals at SUNY Upstate tested their hypothetical vaccination intervention, which involved partnering with community organizations such as the Salvation Army, allowing patients a Q&A session prior to vaccination, and connect to vaccination specialists through community liasions. The authors of the study paired their subjects based on age and household income across 10 different community sites, separating them by intervention positive or intervention negative status, and measuring proof of influenza vaccination in the presence or absence of the intervention. They wanted to compare if the intervention had successfully raised the vaccination levels across age cohorts and overall. The group then used McNemar’s test to construct their 95% confidence intervals to illustrate the nearly 17% increase (95% CI 15.5-19.5) in influenza vaccination levels (see Figure 2) to compared to state and county level alternative interventions. Impressive! Although the authors don’t report a p-value, with the right null hypothesis, McNemar can calculate one for you. It’s so handy.

Figure 2. The contigency table used to calculate McNemar's test. As shown in Figure 3, McNemar's test relies on reporting those not receiving vaccinations but are enrolled (not explicitly stated in the chart).

One limitation to McNemar’s test is that it’s meant for large groups. However, based on the population scope of public health data, this doesn’t seem to be an issue – in fact it is an advantage for novice public health professionals to know this fact, especially if they’ve never done statistical analysis.   
Figure 3. A screen grab of the McNemar test calculator found on GraphPad. Motulsky recommends readers use this for calculation of confidence intervals and p-values in his book. 
 All this time, I thought public health professionals went off magnitudes of numbers alone, perhaps testing averages across populations in an ANOVA test. As it turns out, they use statistical tests, specifically McNemar’s test when employing tired-and-true case-control designs.

Monday, April 11, 2016

Tiny bird headphones and testing model fit: An example use of the extra sum of squares F test


One of the main avenues of research in our lab is quantitative analysis of behavior. One behavior we are particularly interested in is sensorimotor error correction in songbirds. “Sensorimotor error correction” refers to the process by which sensory feedback (such as the auditory feedback of hearing oneself sing) is used to correct a motor behavior (such as singing).

To induce a sensory “error”, our lab fits birds with sets of miniature headphones. While the bird sings, a microphone in the cage records them, and sound processing software will artificially “shift” the pitch of the song up by a couple semitones. This pitch-shifted version song is then played back to the bird through the headphones, virtually in real time. To compensate for the “error” it hears in the auditory feedback, the bird will start singing at a lower pitch (note that if you artificially shift the pitch down, the bird will shift its pitch up).
Previous work had shown that birds will learn to compensate for pitch shift at a faster rate if the shift was small. For very large pitch shifts, they barely learn at all. It’s important to note here that while each bird has its own individual song, they don’t sing it exactly the same way every time. The pitch of a particular note will vary from rendition to rendition, in a normally distributed manner. A technician in our lab hypothesized that it wasn’t the raw amount of sensory error from the pitch shift that influenced learning rate, it was the overlap between the error and the distribution of pitches the bird typically sang at that mattered most.

To test this, the technician decided to use an extra sum of squares f-test.  He created two models (not discussed in detail here to avoid going too egregiously over the word limit, see the paper for more), one of which included parameters for both error size and overlap with the prior distribution, and one that just included error. Then, he took birds with a range of pitch distributions (young adult birds have more variability, and thus a wider distribution of pitches they sing at, while older birds have a narrower distribution). He then tested those birds with a variety of different error sizes via the headphones pitch shift.
The extra sum of squares F test is a way of comparing the fit of two nested models. “Nested” models are models which are identical, but one has additional parameter(s). The extra sum of squares test asks whether the additional parameters reduce the residual error or not. In the case of my labmate, he wanted to know whether there would be less residual error in the error + prior distribution model than the error-only model.

GraphPad’s help page offers another great example of nesting, which may be more intuitive to most biologists:

If you asked Prism to test whether parameters are different between treatments, then the models are nested. You are comparing a model where Prism finds separate best-fit values for some parameters vs. a model where those parameters are shared among data sets. The second case (sharing) is a simpler version (fewer parameters) than the first case (individual parameters).”

The change in residual sum of squares is divided by the additional degrees of freedom for the extra variables, giving us a mean square. The mean square is then compared to the residual mean square from the full model. An F-test allows us to determine the likelihood of our result, assuming the null is true.

In my labmate’s experiment, the more complex model that included prior distribution overlap significantly reduced the residual error. Check out the full paper here. Here is a longer explanation of extra sum of squares F tests.

Is statistical testing worth it?



In preparation for this assignment, I read an article titled The controversy of significance testing: misconceptions and alternatives. As the title suggests, the article went into some detail regarding the controversy surrounding significance testing. The main controversy was the idea that the P value is often misinterpreted and that other factors such as confidence intervals and effect sizes are ignored. This point was interesting and reminded me of something from the textbook. The idea that just because a result is statistically significant does not mean it is important. While I still believe proper experimental design and statistical analysis are important, I could identify with this critique regarding the misinterpretation of P values and what they really mean. It is frustrating to think that after all the time planning and executing an experiment with a statistically significant P value, that the result was really irrelevant. The book gave the example of a drug that led to a decrease in symptoms with a statistically significant P value. However, the statistically significant result only decreased symptoms by 7%, which was not enough in the broader scheme of things.  Because my research is related to the effects of a certain compound on cancer cell growth, a statistically significant result that would not really benefit patients isn’t ideal.

            In favor of statistical testing is first the fact that this article was written in 1999 and things have since improved with the reliability of statistically significant outcomes. Furthermore, in order to properly run an unbiased experiment, the design must be planned first, improving the integrity of the science conducted. I think the thought that must go in to properly performing research leads to better execution of science, and if more people prepared correctly, it may improve issues with reproducibility in science. Overall, I believe the benefits of statistical testing outweigh the drawbacks. When done correctly, the outcomes of research are more reliable.



Testing your outlier; turtles all the way down




      If you have been doing bench work for any length of time, you have had an experiment that had seemingly beautiful data that easily passes the bloody obvious test but still is not significant. You begin to dig through the individual data points and you find it, that one mouse/well/prep that is wildly off from the others. That little *&$@er. Being a good scientist you don’t want to throw away data, your lab notebook says that you did everything correctly that day, and your controls look good. It is not beyond the pale that the value actually happened and was recorded correctly, biological systems are messy and will spit up on you from time to time. 
      But you really don’t have time/money to repeat it, so you begin the squicky task of seeing if you can justify excluding that value. These are the tests before the test, and could probably stand to be done before all analyses to make sure they conform to your assumptions rather than as a post-hoc measure when something goes wrong. So you begin with a simple Q test, the easiest way to justify an outlier’s removal. So you divide the gap by the range and find that value on the table of Q values. But here you have another set of choices to make depending on your sample size and how sure you want to be of the values outlier status. Do you accept a 90% confidence interval on outlier identification? Or are you more stringent, going for 99%?  Perhaps somewhere in between? Perhaps you just really need this post-doc to be over and consider bumping the range below 90%. 
     Confused, you go to find more options and find a plethora of other outlier tests; Pierce’s criterion, Chauvenet’s, and you panic, realizing that many outlier tests have their own assumptions about the normality and variance of your data. What if you have a system where the variance is expected to go up as the dose does? Worse, how would you even know that your data is actually normal?  Well there are many tests for the latter, each with their own assumptions and methods. You can do it graphically with a qq plot, which may make it easier to explain to your advisor, or you can do it by either frequentist or Bayesian methods, but almost inevitably you will find that there are assumptions underlying each of those, and again you can search for a test to prove your data does or does not fit them. One errant point has consumed your work day learning the nuances of each statistical test to determine only if you could throw it away, nevermind testing your actual question. You sit staring at the fractal decision flowchart in front of you, little lines trailing off into nothingness. All due to that little *&$@er.

Friday, April 8, 2016

Red(shirt) Alert: Stats of expendable crewmen

Statistics-first research planning is a relatively new concept to me, though one that I’m increasingly seeing as important. Over the past few weeks, Dr. Murphy has mentioned several times that a researcher should know exactly which tests they want and use them to guide their experimental design. Scientists are often trained, however, in an experiment first, analysis later manner. Rarely have I seen a lab class include the use of a statistical test to examine data besides linear regression. Instead, statistical testing is often taught in standalone classes where data is contrived. Even working in a non-academic laboratory, the mentality was data now and we’ll ask the staff statistician what to do with it later.

Though, through several stats courses, I’m familiar with most of the standard statistical tests used often in lab work, they are still an afterthought. Combatting years of procedural habits is difficult, but there’s time to get into the habit of statistics-first planning before I really begin in my thesis work. The first step to that is to integrate the use of statistical tests into my daily thoughts and, since many of my musings revolve around one topic in particular, I thought it best to begin there. Ahead, full impulse, Mr. Sulu!


One well-known joke arising from Star Trek: The Original Series is that of the ominous red shirt. (But only TOS red shirts, not any subsequent series) Kirk, Spock, and Bones beam to a planet with a small force of red-shirted security and within a few minutes those escorts are victims of the latest alien/ plant/ disease/ energy/ weather mishap. Just observationally, it seems that wearing a red uniform increases chances of death. But how does that stand up to the rigors of statistics?

In 2013, Matt Barsalou wrote a post for Significance Magazine exploring that idea and several people used his ideas as a springboard for further analysis. Of the 40 on-screen crew deaths (of people wearing uniforms), 24 of them wore red shirts. But Barsalou pointed out that viewing those numbers on their own isn’t necessarily fair considering that there are varied numbers of crewmen of each of the three uniform colors and that multiple departments wear the same colors. One of those departments using red is the security team, which is often included in more dangerous work than, say, the similarly red-shirted engineers. Using Bayes Theorem, he showed that there’s a 64.5% chance that any given red-shirt casualty is on the security team despite security making up a minority of red shirts. Of the remaining non-security red shirts, he calculated that they have an 8.6% chance of being that casualty. Rather than the color of the uniform being the issue, Barsalou concluded, it’s actually the occupation. This makes more sense than a universe-wide aversion to a uniform color.

Jim Frost, an author on the Minitab blog, also confirmed Barsalou in another post where he examined the same data with different tests. When looking at whether red-shirt deaths exceed the overall average of deaths, he found with the chi-square test that red shirts actually have a rate of death that is about average of the three groups, as they die more often but derive from a larger pool of crewmen, whereas gold uniforms die with a higher rate than average due to being such a minority overall. Noting, as Barsalou did, that shirt color isn’t the correct variable to examine, Frost used the two proportions test to look at occupation. With a very low p-value, he found that the rate deaths of red shirts in the security department far exceeded those of other red shirts. Thus, being in security is dangerous, not your shirt color.

Other people on the internet have looked into other Trek-related topics using statistics, including one post using the Mann-Whitney test on Trek film ratings to see if the odd- or even-numbered films are better. Realizing how applicable these statistical tests actually are brought statistics out of the realm of rat stressors and confusion and into the streamlined Federation star ships that I know and love. Spock might agree that statistics-first organization is the logical choice to make, if not the natural choice, and better-designed experiments and better-informed test choices will be the result.

//The video is a song about red shirts performed by the Enterprise Blues Band, a group made of actors from Trek that I saw in person last month. Unless they're all security team, stats says they don't have much to worry about. 



Tuesday, April 5, 2016

Statistical Tests: Don't forget to use your brain

Statistics do not make you smart. Common sense is not to be scoffed at. Canonical designs are often outdated and bad. Statistical tests are not sufficient to determine if your data is acceptable.

Our world changes, our science changes, and we get smarter. Why do we allow ourselves to be crippled by the ignorance of the past? I am a structural biologist, I have worked with x-ray crystallography for several years. Crystallography relies heavily on a set of equations and statistical parameters that determine whether or not our data is “good.” We learn these rules as absolutes when we get started - but every year I learn again and again why these rules are idiotic.

The first rule - where do we throw out data? Anything with a signal:noise below two of course! Except.... Our data collection method involves shooting x-rays at an ordered crystal lattice of our protein and observing the scattering of those x-rays when they interact with electron density clouds around the protein. We use the repeated nature of a crystal lattice in combination with the scattering pattern to work backwards and build the electron density. The more reflections we measure, the more we know. Some reflections are weak, some are rare and not often repeated. Throwing out data that doesn’t have a signal:noise above two is still throwing out signal. Why would we throw out our signal?

We have more rules for when to throw out data. There is a test that measures variance within the data set. As you add more data, the variance grows. As our methods improve and we can collect more data, our statistics actually get worse. We get punished for having a stronger crystal that can handle more exposure. We get punished for having a better detector that picks up more signal.

On top of all of this, we have a series of modelling steps that check us as we model. Are we biasing our system? Does the original data still fit? Except this method of checking is itself biased.

Why do we still use these tests? It is because they are written in all the books, they are hammered into us constantly, and we simply do not use our brains. These tests are presented as our sacred way of doing things - but sacred ways tend to be outdated, inappropriate, and written for a different time.

I urge using your brain over trusting the statistics. I bet they are done wrong the majority of the time, and there is no reason to allow yourself to be idiotic and blindly trust in them. Papers use multiple t-tests to compare several groups rather than an ANOVA. We can use a variety of outlier tests on data that looks “wrong” until something says it is an outlier, but were we using the right test? Statistical tests are a tool, but not a rule. They are not the science and they do not determine everything.

Monday, April 4, 2016

What's in a correlation?

What is correlation? Or perhaps a better question is, what is a good correlation? The answer isn’t very straightforward. I made up the following data to see if there was a correlation between a person’s midi-chlorian count and the number of soft drinks they consumed during a year.



The correlation statistics are as follows:
            p=0.0002
r= 0.3483
            r2= 0.1213
            n=111

The p-value is very small, so you could conclude that there is a strong relationship between midi-chlorian count and soft drink consumption. If there really was no relationship between midi-chlorian count and soft drink consumption your chance of obtaining a correlation this strong is very unlikely. However, the r2 value is very small. This suggests that only about 12% of the variation in midi-chlorian count is explained by variation in yearly soft drink intake.


So, which measure do you look at to judge the correlation? The p-value is really small, which suggests that the correlation is unlikely to occur coincidentally. However, its important to remember that p-value is highly dependent on sample size. This study samples 111 individuals, and with a sample size that large, very small effect sizes can become statistically significant. The effect size is small since r2=0.1213, so we are left with the question, is an effect size of roughly 12% scientifically important? This is a difficult question to answer and it’s probably best left to the judgment of the scientist or the reader. I think this problem raises an interesting question about the strength of correlations reported in media. The news is full of correlation data between various categories, but how are the strengths of these correlations being judged? Do journalists and scientists look at low p-values and decide that a correlation is strong or do they look at a high r2 value and effect size? An alternative to this is for journalists to publish the actual data, and let readers conclude whether the correlation is strong enough to warrant action or consideration.