Showing posts with label Introducing Statistics and Confidence Intervals. Show all posts
Showing posts with label Introducing Statistics and Confidence Intervals. Show all posts

Tuesday, April 12, 2016

No confidence interval for car make reliability?



I'm looking to buy a used car in the very near future, and there are some important attributes I'm looking for: performance, a gorgeous interior, 2015, and, yeah, it should be made on the other side of the pond.

But then I look at my stipend provided by my PhD program and I wonder what in the world it was that convinced me to become a biologist. And then think of Toyota.

Out of curiosity, I take a peak at Consumer Reports' 2015 Annual Auto Reliability Survey. And I find the following plot:




A very informative plot. What they've done is survey subscribers on the vehicles they own, in total covering more than 740,000 vehicles. For each reported vehicle a reliability score is calculated and associated with that vehicle's model. A mean reliability for the model is calculated, and all the mean reliabilities for a make's models are averaged to generate the yellow dot, indicating the make's mean reliability. The blue spread indicates the range from the make's least reliable model on average to the make's most reliable model. But why didn't they use confidence intervals? After all, they do claim to be determining "predicted reliability scores."

A confidence interval is used when sampling a population for a value of some characteristic in order to estimate the population's mean for that characteristic. Since samples are being taken as opposed to observing every individual in the population, our calculations will not give us the true mean, but it will hopefully be close. A confidence interval informs us of the range of values surrounding our calculated sample mean that must include the true population mean with a certain level of confidence.

All that is needed to calculate the confidence interval is the sample size, sample standard deviation, and the sample mean. But I can see why Consumer Reports decided not to have their plot display that. As a consumer, once I see the average of the all the models' mean reliabilities, I don't need to see a spread telling me about the true mean. The true mean is not the goal of the reliability report nor the consumer who's in the market for a vehicle. The spread I want to see is the range from the make's least to most reliable model, indicating the variability in the reliability of the make's models. For example, if I was comparing Mazda's reliability to Subaru's, I would note that their means are very close, but the spread of Mazda's reliability is constrained and sits to the higher end of Subaru's reliability, which varies much more. This information might give me more confidence in Mazda's reliability than if I had simply compared the means. The confidence interval may resemble the worst-to-best spreads, but it simply does not mean the same thing. And once we have "outliers" that skew a spread in one direction (see Hyundai, BMW, Ford, Cadillac, Jeep), the fixed ± error of confidence intervals begin to have less meaning to the consumer.

Perhaps if I were a chemical engineer or computer scientist, that Maserati pictured above (no reliability data available, but who cares when a car is that beautiful!) wouldn't be such a distant possibility. But for now, I'm going to look at Toyota. And it probably won't be 2015, either.

Optimizing the BOT



Prior to taking this class, I had a lengthy conversation with my PI about statistics.
We debated over what statistical methods were appropriate to use for our experiments. She opted for the classic t test and I opted for anything but that.
During this debate she would often throw out statements like,
“We don’t base our conclusions solely on whether or not something is significant.”
“We should be able to tell if a result is significant or not just by looking at the data.”
“We can’t publish without stats.”
“Even if a result is significant, it doesn’t matter if it doesn’t have any biological relevance.”
Looking back on this debate now, I realize my PI was/is a follower of BOT.
BOT standing for the “Bloody Obvious Test” coined back in 1987 by Ian Kitchen. Kitchen noted that there was pressure from journals to use statistics and that p-hacking was a problem. 

                                                but it does seem that too often we
labour over their (statistics) use unnecessarily
and indeed on other occasions we
manipulate them to prove a very
thin point.” –Ian Kitchen

Because of these issues, Kitchen proposed the use of the “Bloody Obvious Test. ”
The protocol for the BOT is as follows:
                Question #1: Is it bloody obvious that the values are different?”

                  Answer: Yes.  The test is positive, proceed to “Go” and collect $200.
                  Answer: No.   Proceed to question number 2.

                Question #2: “Am I making a mountain out of a molehill?”

Kitchen really wanted to drive home the point that statistics were being abused to appease “the gods of statistics” who happened to frequently sit on journal review boards. He wanted to remind scientists that sometimes the easiest and most obvious answer is the right answer. Lastly, he wanted scientists to recognize that statistical significance doesn’t always equal scientific significance.

Sadly, Kitchen didn’t stop these issues from persisting in science today. Scientists are still appeasing “the gods of statistics” because to be successful in science, you have to publish.

As the reality of science publishing seems unlikely to change and the pressure to include stats continues, I propose we optimize the BOT with confidence intervals.

Confidence intervals are a form of statistics that provides a range in which the true population value may lie. Traditionally, we set confidence intervals at 95%. A 95% confidence interval tells us that there is a 95% percent chance that confidence interval contains the true population parameter of interest.

The addition of CIs would add a statistical robustness to the BOT, that would perhaps appease “the gods of statistics.” Also, the addition of confidence intervals wouldn’t detract from the initial step of the BOT. We could still ask Question #1 without a pesky p-value getting in the way of our conclusion. Instead, confidence intervals would be to the BOT “as a drunk uses a lamp-post; for support rather than illumination.

Monday, April 11, 2016

Quick and Dirty Significance

Thought about logically and mathematically, it is unsurprising that confidence intervals and p-values are related. In the calculation of confidence intervals, for example in a t-test, the same t statistic that is calculated in order to determine a p-value is used to calculate confidence intervals. However, it is rare to see “statistical significance” addressed in terms of a confidence interval, instead of the more commonly calculated p-value. We know that confidence intervals quantify the random error seen in experimentation, and that they represent the range of values within which the true population value lies with whatever confidence we decided to use. A value outside a 95% confidence interval is unlikely to be the true population value, since we can say that the true population value lies within our interval with 95% confidence. Expanding this to the idea of “quick and dirty significance” means looking at our calculated confidence interval and determining whether our null value is contained within it. Given a 95% confidence interval, a null value outside of it has a low (<=5%) probability of being the true value. On the other hand, if our null value is included within our confidence interval, the null is at least somewhat consistent with the observed data, and further analysis is needed. Wayne W. LaMorte of Boston University School of Public Health recommends thinking of the relationship between a p-value of 0.05 and a 95% confidence interval in terms of an “embrace.” Values within the confidence interval “arms” are “embraced,” and therefore are not rejected. Null values within these arms will have calculated p-values greater than 0.05, and so are determined to not be statistically significant. Alternatively, if the 95% confidence interval does not contain the null value, there is no “embrace,” null hypothesis is rejected, and the p-value will fall below 0.05.
Furthermore, confidence intervals in multigroup comparisons can also be used as a “quick and dirty” check for significance. If intervals do not at all overlap, the 95% confidence intervals allow you to say with at least 95% confidence that there is a significant statistical difference between the groups. Large overlaps, on the other hand, support a lack of statistical significance, and p-values above 0.05. However, this method is slightly more “dirty” in that interval overlap can be as much as 25% and still be statistically significant! In these cases, it is best to conduct general calculations of p-values and be absolutely sure of your results. 

From colored marbles to real science experiments: looking at your data with truthfully



When you hear the word statistics, what is the first thing that comes to mind? A coin that is flipped over and over again?  A die game that you seem to never win and always blame bad luck? Perhaps built up frustration about logging into Prism and not being able to figure out which graph you need to present your data?  

Statistics seems to start off easy. The professors get excited and they say, “Let’s start with a coin!” And then they progress to “Oh! We can move to pretty, colored marbles!” And by the end of it, you’re asking yourself what in the world this has to do with the mice that are downstairs that have various treatments with various time points to measure various readouts at.
Well it all starts with coins and colored marbles. Our pretty bag of colored marbles contains known distributions of different colored marbles; therefore, it is a simple probability about which one you will choose out of the bag at random. Now take your mouse experiment, in which the distributions of the various responses are unknown. The bag of colored marbles is now all the mice in the world, but a typical experiment might have only 20 mice. How are you to determine that a response in a few of your 20 mice is not just random, but indicative of a significant result?
According to the central limit theorem, the larger our sampling, the more the distribution of responses forms a normal distribution centered around a mean. This is one large assumption that has many assumptions hidden within it. First, we assume that we can measure a mean and standard deviation from our data. We also must assume our sampling is completely random and independent. Further, we assume that with increased sampling (more mice), the standard deviation of the mean becomes smaller and smaller, allowing us to be more confident that the sample mean we are measuring (our 20 or so mice) is nearing the population mean (all the mice in the world). From the beginning, we are already assuming a lot.


From there, it gets even a bit hairier. Based on the normal distribution created from our sampling (after a few assumptions), significance is then determined by comparing a preset threshold of error to the probability of obtaining a particular result. If the probability is fairly low for obtaining a result, then it is more likely to pass below our error threshold and become significant. If it does not pass below our error threshold, there is too much error involved with claiming its significance and is more likely to have occurred by chance. It's important to see here that our definition of significance depends upon error.    
From the outside looking in, statistics seems like a black box in which data go in and significant results come out, but upon further analysis, we simply make assumptions, sample populations and then infer. Although the premise is simple, it is critical to remember that all our inferences about significance are based on “unlikelihoods” that could have occurred by chance alone and consist of many assumptions that might not have been met. A proper understanding of the statistical analyses done to yield particular results is extremely important in determining how confident we can be in those results.

Counter-intuitive apophenia and common statistical problems


Introducing Statistics and Confidence Intervals

An interesting idea that I wanted to share was the concept of “Apophenia”, or the problem where humans find patterns and meaning where there is none. This idea got me thinking, and I went back and flipped through our textbook, and noticed several instances in the first section where the author discusses how humans tend toward counter-intuitive decisions. For example, Motulsky says, “Our brains have simply not evolved to deal sensibly with probability, and most people make the illogical choice.” When it comes to scientific research, the most obvious illogical choices made always seem to tie to the statistics reported to support a publishing’s claim. The two most common culprits with this problem are the P-value and confidence intervals (or more often, the lack thereof), and these are two basic concepts that were introduced early in the textbook.
But why are these two statistics problematic?
The more I researched about P-values, confidence intervals, and the statistical pitfalls of scientific research, the more I felt I was reading the same words and phrases. Everyone agrees unanimously and rallies behind this banner with the battle-cry for reproducible research, and for more stringent statistical reporting! But it was rare to find a proposal of any actual mechanism to correct the problems.
One article from Erika Check Hayden in NatureNews explored an interesting proposal from statistician Valen Johnson. Johnson developed a method to directly compare, “the P value in the frequentist paradigm, and the Bayes factor in the Bayesian paradigm”, and then re-examine select published data to see how the two juxtaposed. His results found that the common standard of p≤ 0.05 coordinates with a weak Bayesian factor, and that of the published data reviewed, “as many as 17–25% of such findings are probably false”. To ameliorate this problem, Johnson suggests, “to use more stringent P values of 0.005 or less to support their findings, and thinks that the use of the 0.05 standard might account for most of the problem of non-reproducibility in science — even more than other issues, such as biases and scientific misconduct”.
Another article I found from Nature News that was published in Medicine by Jean Baptiste-du-Prel et al. discusses when p-values and confidence intervals are most important individually, and when it’s important they’re reported combined. For example, one major point that is discussed is that, in clinical research, the over-stressed p-value is a problem most reports make, and Baptiste-du-Prel even states, “the investigator should be more interested in the size of the difference in therapeutic effect between two treatment groups […] rather than whether the result is statistically significant or not”. The take home message from this article was that both p-values and confidence intervals are intertwined statistics should be reported together to bolster credibility.
While these and other articles I found in my research were very interesting, I can’t help but think about how these more stringent statistical standards will be achieved, and what sort of impact raising the bar will have on new scientists just beginning. While I, personally, am excited at the idea of knowing all the papers I read genuinely are statistically significant, I can’t help but wander back to the problem that got me researching to begin with.

How much of the statistics published in scientific literature is apophenia that stems from lack of basic statistical knowledge?
In a “publish or perish” scientific arena, I can certainly understand why so many statistical mistakes can slip through the cracks, but that doesn’t really change the fact of why there is this statistical apophenia at all. Perhaps if instead of raising the bar for statistical stringency in scientific publishing, we should raise the standards for statistic education for scientists.  Then the problems of abusing the p-value, misinterpreting significance, and making unnecessary statistical decisions illogically, would correct themselves.

Sunday, April 10, 2016

Patterns Real and Imagined



One of the interesting topics covered in Harvey Motulsky’s first few chapters is the argument that probability is not intuitive because people tend to identify patterns even when none are present. He provides the example of basketball players being perceived as more likely to make or miss their next shot based on their current “streak” of successful or unsuccessful shots. As a side note, I think this is a rather poor example of the point, because it suggests that whether a basketball player makes or misses a basket is based on random chance as opposed to all the other factors that go into it. A better example is the randomly generated table provided on page 5, which could be interpreted as depicting patterns. As Motulsky points out, humans are adept at identifying patterns because it is evolutionarily advantageous to do so. I would propose that scientists are more perceptive than the average person to hints of patterns, as we are trained to detect regularities that point us to the underlying mechanisms that govern our world. That also means that we are decidedly prone to introduce bias into our work even with the best of intentions. If in the course of an experiment we start to see a trend emerging, we tend to look harder for more data that fit that trend. I discovered this during my first pilot experiments that measured disease severity in mice, and I have performed all subsequent experiments of this type blinded to genotype. Blinding is not a universal fix for this problem, though. Another instance of possibly spurious pattern recognition in data that comes to mind is multimodal populations. If you look at a scatter plot and see points clustered in what seem to be two groups, it is tempting to think that perhaps they reflect a bimodal response to a variable. Flow cytometry is another area where this can occur, as it is often possible to identify numerous populations that seem to express different combinations of marker intensity. In the complexity of biological systems, the possibility that these “patterns” in the data represent truly distinctive physiological entities is very real, and especially in more variable systems such as human studies or experiments with outbred animals, it is not at all unlikely that subsets of individuals could exhibit different responses to treatments that are based on underlying physiological differences. For instance, studies relevant to our lab’s work have found that a subset of depressed patients exhibit high levels of inflammatory markers and that their depressive symptoms can be improved with anti-inflammatories(Raison 2013). So it is important for scientists to recognize and pursue patterns that may lead to outcomes like this. But we must also recognize the potential for bias that comes if we choose to focus only on one perceived population of a dataset that “behaves better,” and also the potential to miss interesting findings by subdividing populations to the point that we lose experimental rigor.