Showing posts with label error. Show all posts
Showing posts with label error. Show all posts

Monday, April 11, 2016

From colored marbles to real science experiments: looking at your data with truthfully



When you hear the word statistics, what is the first thing that comes to mind? A coin that is flipped over and over again?  A die game that you seem to never win and always blame bad luck? Perhaps built up frustration about logging into Prism and not being able to figure out which graph you need to present your data?  

Statistics seems to start off easy. The professors get excited and they say, “Let’s start with a coin!” And then they progress to “Oh! We can move to pretty, colored marbles!” And by the end of it, you’re asking yourself what in the world this has to do with the mice that are downstairs that have various treatments with various time points to measure various readouts at.
Well it all starts with coins and colored marbles. Our pretty bag of colored marbles contains known distributions of different colored marbles; therefore, it is a simple probability about which one you will choose out of the bag at random. Now take your mouse experiment, in which the distributions of the various responses are unknown. The bag of colored marbles is now all the mice in the world, but a typical experiment might have only 20 mice. How are you to determine that a response in a few of your 20 mice is not just random, but indicative of a significant result?
According to the central limit theorem, the larger our sampling, the more the distribution of responses forms a normal distribution centered around a mean. This is one large assumption that has many assumptions hidden within it. First, we assume that we can measure a mean and standard deviation from our data. We also must assume our sampling is completely random and independent. Further, we assume that with increased sampling (more mice), the standard deviation of the mean becomes smaller and smaller, allowing us to be more confident that the sample mean we are measuring (our 20 or so mice) is nearing the population mean (all the mice in the world). From the beginning, we are already assuming a lot.


From there, it gets even a bit hairier. Based on the normal distribution created from our sampling (after a few assumptions), significance is then determined by comparing a preset threshold of error to the probability of obtaining a particular result. If the probability is fairly low for obtaining a result, then it is more likely to pass below our error threshold and become significant. If it does not pass below our error threshold, there is too much error involved with claiming its significance and is more likely to have occurred by chance. It's important to see here that our definition of significance depends upon error.    
From the outside looking in, statistics seems like a black box in which data go in and significant results come out, but upon further analysis, we simply make assumptions, sample populations and then infer. Although the premise is simple, it is critical to remember that all our inferences about significance are based on “unlikelihoods” that could have occurred by chance alone and consist of many assumptions that might not have been met. A proper understanding of the statistical analyses done to yield particular results is extremely important in determining how confident we can be in those results.

Type I and Type II Errors


For those of us who might still be struggling with how to remember Type I and Type II Errors :o)

Wednesday, March 23, 2016

Error Bar Misinterpretation


Nature Methods published an article fairly recently that explains common misconception surrounding error bars. If you're like me, I thought error bars was something I could easily look at and understand. Don't they just represent the likelihood of variance between my replicate samples? Well, first, what are you using your error bars to represent? Standard deviation, standard error of the mean, or a confidence interval? This article (which provides interactive supplementary data where you can see the raw data used in the discussion as well as make your own data), provides examples of how data can look very different (or even not significant) depending on the way the error bars are represented.

A very common misconception is that a gap between bars means that the data are significant while if the bars overlap they are not significant. That is not the case, and again, it all depends on the type of bars you choose to use. 


An an example, figure 1 (above, n=10) shows how error bars cannot be compared. The left graph shows what happens to the p-value when the error bars from SD, SEM, and 95% CI are adjusted to the same lengths. The right graph shows what happens to the size of the error bars when adjusting to a significant p-value (p=0.05). As you can tell, just because SEM error bars do not overlap does not indicate significance and just because SD error bars do overlap does not mean that the data are not significant. 

When showing data with error bars it is important to be clear about which measure of uncertainty is being represented in order for the reader to be able to interpret the results properly.

Short summary:
SD: represents variation of the data and not the error of your measurements.
SEM: represents uncertaintiy in the mean and its dependency on the sample size.
CI: represents an interval estimate indicating the reliability of a measurement.



Tuesday, January 19, 2016

The Anatomy of Deceit

Like many great scientists before him, Dan Ariely was inspired to answer questions surrounding what he called “deceitful behavior” from a real life experience he had in the burn ward of a hospital. From his laboratory experiments on pain and reward at MIT and CMU, he concluded a few key points, which will serve as a framework for my reaction to the accompanying articles. These key points are as followed (paraphrased from Mr. Ariely’s own words):
  1. Many people engage in deceitful behavior, but only do so a little bit at a time.
  2. When people are reminded of their own morality, deceitful behavior goes down.
  3. If someone is out-performing the rest of the group and part of the in-group, deceitful behavior increases.
  4. When there is distance from a tangible end-point, deceitful behavior increases.
  5. People have a hard time doing difficult tasks to prove they are engaging in deceitful behavior. 
Some of these points, when contrasted against the accompanying articles, brought up interesting questions for me on how and why dishonest science happens. For instance, according to the article written by Julia Belluz on supposed “miracle” drugs, there appears to be both the in- and out-groups who perform some level of deceitful behavior. That is, it is not only the doctors who use exuberant language to describe the results of certain cancer drugs, but also the journalists who are responsible for reporting on them.
                                                         
Can one then make the argument that journalists, like medical practitioners, occupy the same in-group? Or is it that the in-group and out-group have a symbiotic relationship where the in-group (doctors) can influence the out-group (journalists) and vice versa? In addition, Belluz cites immunotherapies, “the vanguard of cancer research,” as being the most frequently hyped cancer therapies. This addresses Ariely’s 4th conclusion above: that is, a cure for cancer is far off in the distance, but scientific publications exist as an immediate means of professional currency. However, it exposes an interesting question for me: would cancer researchers not working in cancer’s “hottest field” feel the need to engage in describing their therapies with such hyperbolic rhetoric?

My gut tells me the answer to this question is no, especially considering the implications of curing cancer. I believe no matter where you end up, these overreaching descriptions of therapeutic results serve as a way to move the field forward, albeit not in a very honest way. This grandiose language exposes dishonest behavior by putting a proverbial red flag to heed attention to potential results. In Jared Horvath’s article, he suggests that these mistruths are simply a consequence of science, and that reproducibility, whether it can be achieved or not, must be fully disclosed. Furthermore, the inability to reproduce serves as a helpful caveat to moving the body of research forward. In many ways, Hovarth seeks to engage more researchers in Ariely’s 5th conclusion: he hopes that researchers will undertake the difficult tasks of proving their deceitful behavior for the common good of science. This, I believe, is the future of scientific research -- engaging with our human errors.