Showing posts with label nonparametric. Show all posts
Showing posts with label nonparametric. Show all posts

Monday, May 2, 2016

Reading Boxplot used to deal with skewed Data

Recently I came across some data that were gathered from human patient samples.  At first glance, it was really difficult for me to understand how to interpret the data, and why the data was generated this way but not the other way around, for example, t-test or ANOVA (See below the figure).  And it is also confusing to read the black bold line, the whiskers and the dots.  I talked with the scientist who sent me the data and combined with what we've come across during this semester and got a general idea of how to read the data.



When we are dealing with patient expression data, often times they are not normally-distributed, but they are "skewed" as shown by the example below.  The mean of the data is at zero represented by u.  However, as it is "skewed", the u (dotted middle line) is a not as informative as a median( the line).  Also of notes, when there are a few outliers in the data collected, they couldn't be excluded when statistical analysis is performed.  We don't know for sure if outliers are caused by measuring errors or just a true responses of certain patients.  In this case of not excluding, the inclusion the outliers could change the mean of the samples and makes it quite far from the bulky of the data.  This makes the median preferred to the mean in human patients data.  


So now we know it is totally different from data drawn from the normally distributed population. With that in mind, it would be much easier to understand the Boxplot.  As shown by the following image, the box represent the area of values between first quartile (25%) and third quartile (75%). The line represents the median, which is usually not in the exact middle of the box.  The whiskers represent the minimum and maximum or 1.5IQRs(Interquartile Region).  The asterisks or hollow dots represents the outliers.  



Now it seems not difficult for me to understand the first figure.  Then how to know if there is a significance when we compare one group to the control?  It's not that robust as compared to t-test or ANOVA.  p-values could be generated using some software.  A very small p-value makes it statistically significant. However, even the p-value is larger than the type I error threshold we set, we couldn't say there is no difference.  

There could be more about boxplot.  For example, in some cases scientists could do the non-parametric tests like Wilcoxon signed rank sum tests if data variation is too large or the parametric tests don't apply to original dataset (the first figure is an example of this).  There is always more in statistical methods and that's why it needs scientists intuitive to perform good stats.   

  

Friday, April 29, 2016

Non-parametric test flow chart

In class we discussed the importance of using non-parametric tests on ordered data (i.e. data from a subjective 1-to-10 scale). We talked about the variety of tests available, but I decided to make a handy flow chart for anyone who wants a quick reference of which test to chose.




As you see above, the first deciding factor is the number of factors. This weeds out your possibilities really quickly since you can only properly do one-factor non-parametric tests. For reasons that are really complicated, there simply isn't a way to do the equivalent of a two-way ANOVA for non-parametric data; signed-rank tests aren't made to produce enough information to have any meaningful analysis across multiple factors. A search through the internet shows that some people have attempted to develop multi-factor non-parametric tests, but statisticians have deeply contentious opinions on whether they work or not, so it's best just to avoid them. So, keep this in mind when designing studies that all non-parametric data should only be tested by one factor at a time.

Next, your test depends on the number of sample groups you want to compare. If you only have one group and want to compare to a set/expected value, then use the Wilcoxon signed rank test (non-parametric equivalent of a one-sample t-test). If you have two groups, you can use the Mann-Whitney U-test for unpaired samples (equivalent of unpaired t-test) or the Wilcoxon matched-pairs signed rank test for paired samples (equivalent of paired t-test). If you have three or more groups, there actually is a non-parametric equivalent of a one-way ANOVA! We didn't talk about it much in class, but it is called the Kruskal-Wallis test, which I will address in my next blog post. If you are making multiple planned comparisons, use the Dunn's test to correct for the multiple comparisons. If you are not making multiple comparisons, you can use the Fisher's LSD test.

Hope this flow chart helps my fellow visual learners out there!