Showing posts with label null hypothesis testing. Show all posts
Showing posts with label null hypothesis testing. Show all posts

Friday, April 8, 2016

Interpret the P-values and Statistical Significance Properly

The p-value is probably the most commonly used reference for decision-making in biological and biomedical sciences: A result with a p-value less than 0.05 is valuable, while a result with p-value greater than 0.05 is worthless. Clear-cut, right? Unfortunately, no. A p-value cannot tell if an experimental result is worthy or not. It cannot tell if the original research question is addressed. It even fools someone when one obtains some p-values greater than 0.05 while others less than 0.05 from different tests. To interpret the p-value properly, we may start by asking: ‘What is p-value designed for?’

Let’s say we want to know whether a coin is fair, i.e., the probabilities of getting a head (H) and getting a tail (T) are equal to 0.5. One flips the coin 10 times, and gets a result

H, H, T, T, H, T, T, H, H, H.

Most people would say this coin is fairly fair, although the number of head is not exactly 5. However, if the one gets a result

H, H, H, H, T, H, T, H, H, H.

Now some people become indecisive: Does this result show the coin is not fair, or this seemly extreme case occurs simply by chance? An intuitive way that helps on decision making is to calculate the chance of observing this result or more extreme cases given the coin is fair and compare this probability to a pre-defined threshold. 


If the threshold we employ is 0.1, then the hypothesis of fair coin is not true since it yields the chance of extreme cases lower than the threshold. This is the exact p-value, a p-value calculated from the distribution of interested random variable. Most of the tests, like Z-test and Pearson’s chi-square test, calculate the approximate p-value, which is based on the same idea but using another distribution to approximate the distribution of the test statistic.

Either exact or approximate p-values, the concept is clear: P-value is the probability that one observes the data under null hypothesis. This brings us two things. First, a complete hypothesis testing includes four parts: the null hypothesis, the alternative hypothesis, the threshold (i.e., the alpha), and the test statistic with corresponding probability distribution. All of four parts affect the p-value and statistical inferences. Second, a small p-value does not imply the scientific hypothesis is true. It still can occur by chance, especially when multiple comparison is conducted. Besides, the statistical hypotheses may depart from the scientific hypothesis by employing a surrogate maker (e.g., using blood low-density lipoprotein level as a surrogate for risk of cardiovascular disease), measurement error, or data transformation (e.g., transforming a continuous response to binary).


P-value is widely used, sometimes it’s abused. The best way to interpret the statistical significance reported on research papers is asking ‘What is the hypothesis they are testing? Why did they choose this procedure? Is there other applicable tests? How far is the testing procedure away from their research question?’. Just keep cautious: P-values can fool you.

Wednesday, April 6, 2016

Jelly beans and statistical significance
















Statistical significance is present in the everyday life of most people, such as weather predictions, political campaigns, medical studies, quality testing,  insurance and the stock market. Most of the people in the world use statistics unconsciously by noticing patterns in daily circumstances and drawing conclusion based on those patterns. On a greater scale, researchers use statistics to represent their data in a meaningful way.
But what does “significant” means?
If you would open a dictionary you would find the definitions “important” or “meaningful”, but saying that research results are significant, doesn’t mean that they are important. Indeed, a statistical significant result means that two the difference seen between two groups is real and not given by chance. In other words, the falsification of the null hypothesis will occur by chance only under a certain percentage that appears to be set at 5%.
It is still unclear where the origin of the 5% threshold lies, but the most reliable source can be found in the discussion published by Fisher in 1926 on the theoretical basis of the experimental design.1
The real question is, what does this p-value tell us in terms of significance in research?
When conducting studies, researchers should keep in mind three main points:
1.     The dichotomization of p-values into “significant” and “non-significant” leads to a loss of important informations. Two values might be significant, but that doesn’t imply that they are the same.
2.    Statistical significance is not directly linked to clinical significance. As statistical tests are influenced by the sample size, a significant study does not always mean that the outcome is clinically meaningful. A large study might be significant and not be clinically relevant, while a small study can be important as outcome, but not statistically significant.
3.    Although it is tempting to rely only on p-values, the weight that researchers give to them should not be overemphasized. The most important question should remain on the qualitative level of the study, such as design, sample type, patients and bias.

Nowadays, we are overwhelmed by advertising for weight loss pills, miraculous anti-wrinkles creams and any other kind of aesthetic treatment stating that you will get significant results based on data collected in clinical trials. What they clearly forget to mention, it’s what they truly mean by “significant”.


1. Fisher RA, The arrangement of field experiments,

   J. Ministry Agric.,1926, 33:503-513

Monday, March 21, 2016

To ban or not to ban the P-value: A question of qualitative vs. quantitative data raised by ASA and BASP

TJ posted an article earlier this month about how the ASA issued a statement concerning the P value, saying that "statistical techniques for testing hypotheses.... have more flaws than Facebook's privacy policies."

I found out after some click-holing from article to article about statistics and the P value that Basic and Applied Social Psychology (BASP) has actually banned the P value since 2015, and more specifically, the null hypothesis significance testing procedure, or NHSTP. BASP even states that prior to publication, all "vestiges of the NHSTP (p-values, t-values, F-values, statements about 'significant' differences or lack thereof')" would have to be removed. And this basis arises from the fact that numbers are being generated where none exist, a problem that I feel is more specific to more qualitative fields such as psychology. But the BASP raised the same concerns as the ASA statement regarding the p value: that p < 0.05 is "too easy" to pass and sometimes "serves as an excuse for lower quality research." While the ASA's concern is mainly geared towards the people and scientists who perform the research (i.e. people are not properly trained to perform data analysis), it seems that BASP's concern arises from the nature of psychological research, stating that "banning the NHSTP will have the effect of increasing the quality of submitted manuscripts by liberating authors from the stultified structure of NHSTP," even stating that it hopes other journals will follow suit.

I certainly understand where BASP is coming from - with a field where response/measured variables are more often qualitative than not, how does one effectively apply statistics to analyze whether or not an effect is real? What should we do about data generated in the "hard sciences" that are more qualitative, such as characterization of cell morphology? What about clinical data that measure subjective things such as level of pain on a scale of 1-10? Is there an existing statistical tool or procedure out there that everyone could agree would accurately measure "significance" without having to apply values or generate numbers to describe qualitative measurements?

Do you think we should abolish P-value significance testing for all research? Only psychology research? How about all qualitative vs. quantitative research?