Showing posts with label P values and statistical significance. Show all posts
Showing posts with label P values and statistical significance. Show all posts

Saturday, April 16, 2016

p less than 0.05: habit or sound clinical judgment?

The p-value is the observed significance level of a result. In the context of hypothesis testing it is the smallest level of significance at which the null hypothesis would be rejected for a given test procedure and data set.  This allows for comparison with a chosen threshold of type one error, alpha, which is the probability of the rejecting the null hypothesis when it is actually true.  Thus a p-value is the probability of obtaining a test statistic value at least as contradictory to the null hypothesis as the value observed if the null hypothesis is actually true.  In terms of applicability, the p-value is used when determining if a test statistic value is statistically significant or not where the threshold of significance is defined by the minimum level of type 1 error one is willing to accept.  In other words, the null hypothesis can be rejected (i.e. result of a test statistic is significant) if the probability of obtaining a test statistic value at least as contradictory to the null hypothesis as the observed test statistic value is less than the maximum probability one is willing to accept that the null hypothesis is actually true in this context of rejecting it.

Now the threshold for significance accepted by most in the scientific community is an alpha of 0.05.  However, is that perhaps too large? If we are trying to determine if a particular treatment is successful at treating a disease is a 5% chance that the determination that this treatment is successful was a false positive acceptable?  Now a lot of this is dependent upon the clinical context.  If this were a case where a patient had little to no other alternatives to this treatment than it may be judged an appropriate threshold.  Alternatively if there are treatments that are equally as effective a higher threshold for significance may be necessitated.  Other factors aside from alternate treatment availability could be cost, risk of complications, and contraindications for future treatment options to name a few. It is trusted that clinical judgment is used when determining a threshold alpha, yet 5% seems to be the standard so much so that I suspect that it is used more out of habit rather than out of sound clinical judgment. 

Additionally, it is important to make an assessment with regards to the goals of the study and whether a type 1 or type 2 (false negative) error would be worse.  The significance of our statistical methods is focused on type 1 error, but if type 2 error is also or more critical that must be accounted for in the study design.  For example, the power of a study needed, the compliment of type 2 error, should be determined a priori which would allow for an appropriate sample size to be determined to ensure the study is not underpowered, thus reducing the chance of type 2 error.

The p-value selected should also depend on the nature of the statistical analysis.  For example, while for one test a p-value of 0.05 may seem reasonable if we begin to perform multiple tests and comparisons the actual error will increases over the collection of studies and it becomes more likely that among our results a false positive has occurred.  Thus in the case of multiple comparisons it is essential to select a lower p-value to compensate.

As a final note, as had been alluded to throughout this blog post, the p-value, although the current foundation of statistical hypothesis testing, does not tell the whole picture (e.g. power and effect size) and is perhaps used too frequently, as a matter of habit. It should be expected that researchers determine what they actually want to show with their analysis and what the implications of their results will be and then choose the appropriate statistical test and parameters (alpha, beta, n ect.) accordingly.

Friday, April 8, 2016

Interpret the P-values and Statistical Significance Properly

The p-value is probably the most commonly used reference for decision-making in biological and biomedical sciences: A result with a p-value less than 0.05 is valuable, while a result with p-value greater than 0.05 is worthless. Clear-cut, right? Unfortunately, no. A p-value cannot tell if an experimental result is worthy or not. It cannot tell if the original research question is addressed. It even fools someone when one obtains some p-values greater than 0.05 while others less than 0.05 from different tests. To interpret the p-value properly, we may start by asking: ‘What is p-value designed for?’

Let’s say we want to know whether a coin is fair, i.e., the probabilities of getting a head (H) and getting a tail (T) are equal to 0.5. One flips the coin 10 times, and gets a result

H, H, T, T, H, T, T, H, H, H.

Most people would say this coin is fairly fair, although the number of head is not exactly 5. However, if the one gets a result

H, H, H, H, T, H, T, H, H, H.

Now some people become indecisive: Does this result show the coin is not fair, or this seemly extreme case occurs simply by chance? An intuitive way that helps on decision making is to calculate the chance of observing this result or more extreme cases given the coin is fair and compare this probability to a pre-defined threshold. 


If the threshold we employ is 0.1, then the hypothesis of fair coin is not true since it yields the chance of extreme cases lower than the threshold. This is the exact p-value, a p-value calculated from the distribution of interested random variable. Most of the tests, like Z-test and Pearson’s chi-square test, calculate the approximate p-value, which is based on the same idea but using another distribution to approximate the distribution of the test statistic.

Either exact or approximate p-values, the concept is clear: P-value is the probability that one observes the data under null hypothesis. This brings us two things. First, a complete hypothesis testing includes four parts: the null hypothesis, the alternative hypothesis, the threshold (i.e., the alpha), and the test statistic with corresponding probability distribution. All of four parts affect the p-value and statistical inferences. Second, a small p-value does not imply the scientific hypothesis is true. It still can occur by chance, especially when multiple comparison is conducted. Besides, the statistical hypotheses may depart from the scientific hypothesis by employing a surrogate maker (e.g., using blood low-density lipoprotein level as a surrogate for risk of cardiovascular disease), measurement error, or data transformation (e.g., transforming a continuous response to binary).


P-value is widely used, sometimes it’s abused. The best way to interpret the statistical significance reported on research papers is asking ‘What is the hypothesis they are testing? Why did they choose this procedure? Is there other applicable tests? How far is the testing procedure away from their research question?’. Just keep cautious: P-values can fool you.

Thursday, April 7, 2016

P-hacking and publication bias

iii. P values and statistical significance

(Note: names have been changed)
I glared at the data glowing back at me on the computer screen – the postdoctoral fellow who was mentoring me had mentioned finding a p-value for all the numbers I had. 
“Sorry, Florence, how did you want me to analyze this data again?” 
“Just calculate the SEM for each group, stick the numbers in PRISM, and then look for the p-value and see if the difference is significant.” Then she walked away. Something about a motor neuron prep and her mouse embryonic spinal cords sitting on ice for too long. 
I had no idea how to do what she wanted me to do. What is an SEM? Why aren’t we calculating standard deviation instead? Google searches brought up a myriad of statistics websites that attempted to explain how to derive the equation used to calculate an SEM, but with little context as to why you would even want an SEM in the first place. I tried to recall the statistics class I took sophomore year, but only came up with “the p-value indicates significant results.” …Right? I sat there questioning my own competence before continuing to toil over how to do the calculations on the numbers from my qPCR. When I finally got the numbers and graphs, I was dismayed that there was no significance. When Florence came around again, I showed her the results. 
“Oh. Well, that sucks. But there seems to be a trend. And we’re only at n=2, so I think if we just increase our n, it’ll probably be significant.”


Type of error bar
Conclusion if they overlap
Conclusion if they don’t overlap
SD
No conclusion
No conclusion
SEM
P > 0.05
No conclusion
95% CI
No conclusion
P < 0.05
(assuming no multiple comparisons)
 Rule of thumb provided by GraphPad's FAQ

How many of us have been put in a similar situation or have heard of a situation like this?

Without a strong background or understanding of statistics, I blindly trusted Florence’s logic and choice of statistical analyses – she was a postdoctoral fellow after all. She’s probably done more than two dozen of these kinds of statistics on her own data that granted her her Ph.D. She must know what she’s doing, I reasoned. But that was the danger of scientists who were improperly or inadequately trained to conduct statistical analyses: in hindsight, I realized that 1) few people (or even scientists, me included) actually understand what “significance” really means, and 2) as Motulsky puts it, “once some people hear the word significant, they often stop thinking about what the data actually show.” The scenario I recounted is something Simmons, Nelson, and Simonsohn (2012) termed “P-hacking,” a term that refers to attempts by investigators to lower the P value by trying various analyses or by analyzing subsets of data. Motulsky draws out two ways in which investigators do this: 1) by tweaking data (if one analysis didn’t give a P value less than 0.05, then they tried a different one) and/or 2) by changing the sample size post hoc (stopping data collection if the P value is less than 0.05, but collecting more data when the P value is about 0.05).

One study by Gotzche (2006) looked at comparing the number of publications that reported a P value between 0.04 and 0.06, hypothesizing that if results were published honestly, the number of publications reporting a P value between 0.04 and 0.05 and a P value between 0.05 and 0.06 should be similar. In the analysis, Gotzsche found that there were five times as many papers reporting P values between 0.04 and 0.05 compared to P values between 0.05 and 0.06. The emphasis on statistical “significance” equating as scientific significance ends up skewing the publication of results and data and creates publication bias. I really wonder if some scientists believe that inadvertently p-hacking is a legitimate way to conduct statistical analyses, or if some do it knowing that it is the improper way to generate "significant" results.


Perhaps the fix here is for journals to start requiring authors to submit a short cover note explaining the justification of the utilized statistics to corroborate that they understood why and how the statistical tools were chosen and used. In this way, it could force scientists to not only conduct reliable and properly designed experiments, but also to think more carefully about the interpretation of their results, rather than just trying to force or find significance that might not be there.

Wednesday, April 6, 2016

Jelly beans and statistical significance
















Statistical significance is present in the everyday life of most people, such as weather predictions, political campaigns, medical studies, quality testing,  insurance and the stock market. Most of the people in the world use statistics unconsciously by noticing patterns in daily circumstances and drawing conclusion based on those patterns. On a greater scale, researchers use statistics to represent their data in a meaningful way.
But what does “significant” means?
If you would open a dictionary you would find the definitions “important” or “meaningful”, but saying that research results are significant, doesn’t mean that they are important. Indeed, a statistical significant result means that two the difference seen between two groups is real and not given by chance. In other words, the falsification of the null hypothesis will occur by chance only under a certain percentage that appears to be set at 5%.
It is still unclear where the origin of the 5% threshold lies, but the most reliable source can be found in the discussion published by Fisher in 1926 on the theoretical basis of the experimental design.1
The real question is, what does this p-value tell us in terms of significance in research?
When conducting studies, researchers should keep in mind three main points:
1.     The dichotomization of p-values into “significant” and “non-significant” leads to a loss of important informations. Two values might be significant, but that doesn’t imply that they are the same.
2.    Statistical significance is not directly linked to clinical significance. As statistical tests are influenced by the sample size, a significant study does not always mean that the outcome is clinically meaningful. A large study might be significant and not be clinically relevant, while a small study can be important as outcome, but not statistically significant.
3.    Although it is tempting to rely only on p-values, the weight that researchers give to them should not be overemphasized. The most important question should remain on the qualitative level of the study, such as design, sample type, patients and bias.

Nowadays, we are overwhelmed by advertising for weight loss pills, miraculous anti-wrinkles creams and any other kind of aesthetic treatment stating that you will get significant results based on data collected in clinical trials. What they clearly forget to mention, it’s what they truly mean by “significant”.


1. Fisher RA, The arrangement of field experiments,

   J. Ministry Agric.,1926, 33:503-513

Tuesday, April 5, 2016

The value of the P-value

Scientist tend to place a lot of importance in the P value. PIs loose grants over P values, grad students have nervous break downs over “insignificant data”, undergraduates desperately look for outliers they can redo to correct that p-value. Is the P-value really that important?

As we have learned in class this semester, the p-value is the probability of obtaining a result equal to or more extreme than predicted with the null hypothesis. So a low p-value indicates that it was unlikely that random chance gave you such an outlying result. Many scientist would stop when they see that low number, wipe the sweat off their brows, and submit their grant or paper.

However is a p-value of 0.051 really all that much worse than a p value of 0.049? Is the difference between those p-values really worth a nervous breakdown? The p-value is an arbitrary line drawn by scientists, and often young scientists fail to look beyond it. A p-value is just one statistical number that can be used in conjunction with other statistical values such as the mean, confidence interval, and error, to make a decision about the implications of an experiment.

A scientist needs to use his/her judgment to conclude if an experiment is worth pursuing further or if it is a waste of time; a significant or insignificant p value should not make that decision. An insignificant p-value does not mean the data was worthless, but only a clear minded scientist can determine what it actually means. Likewise, but much less considered, a significant p-value does not mean that the experiment showed a result worth celebrating over. Your significant p-value might indicate that your drug decreased the stress levels of mice 10%, but is a 10% decrease in mice stress really significant for any application?

These sorts of questions are what scientists should be thinking about. In one experiment, a p-value may hold a lot of weight, but in another experiment, the confidence interval might tell much more about the results. Scientist must be able to analyze their data, not just statistically, but critically to know what their statistics are actually telling them.

Monday, April 4, 2016

The cause of feelings of hoplessness and failure in graduate school: P-Values and Statistical Significance

Scientists, especially graduate students, have become too focused and driven on results being statistically significant. We play statistical significance up to be all-important in science; most of our experiments and projects focus on finding some difference. If we don’t get the results that are “statistically significant”, we feel like failures and that something went wrong. Maybe I am generalizing too much of my own experience in graduate school, but bear with me. Graduate school is notoriously viewed (well, at least by me) as “soul-sucking.” I believe that much of these feelings of hopelessness and failure originate from the moment you press “Analyze” on Prism and see “ns.” Imagine how different graduate school would be if that feeling of failure were eliminated…how would things be if we took every negative result and no longer viewed it as a dead end or a reflection of our abilities as scientists? What if when we saw “ns” we could feel joy and not distress? I feel like our success in graduate school is defined by statistical significance; without a p<0.05, our hard work means nothing. When was the last time that any of us went to a thesis defense that focused on non-significant results? Why has it become that a statistically significant result is necessary to earn our doctorate? Would our education be at a disadvantage if were not required to present statistically significant data?
In a way, statistical significance helps to remove bias by allowing for quantification and comparison of results in order to look for a difference. Statistics and calculating a P value are what allow western blots to be informative and unbiased. Without P value, there would likely be variation in what some would say “looks” like a difference between two groups. Science needs statistical significance. However, statistical significance has also created bias in the way that we approach problems. The need for statistical significance prevents us from exploring concepts and hypotheses that may turn up to be of no significance. The need for statistical significance may also lead a researcher (without proper statistical training) to increase the n of their experiment to the point where a p value of <0.05 is inevitable. It has become unacceptable to just say no significance; we force our P value to mean something, even if it’s just “trending” towards significance. Statistical significance and p values both eliminate bias as well as create it.

I feel that people don’t actually think about what “statistically significant” means; all a P value can tell us is the probability that we could see that a result of the same magnitude if the null hypothesis were true. It cannot actually tell us how likely the alternative hypothesis is true. Thus, we need to stop defining the importance of our work by the P value. Motulsky brings up that colloquialisms may contribute this problem of focusing on statically significance a P values. We associate the term significance with importance, which is incorrect when interpreting statistics. In order to interpret statistics, one must understand the theory and definitions of the terms used. Then, and only then, can we understand that statistics does not interpret the importance of our experimental results; it only allows us to accept or reject the null hypothesis.  We can no longer define our work and goals by “statistical significance”; instead we should be seeking scientific importance.