Showing posts with label t-test. Show all posts
Showing posts with label t-test. Show all posts

Tuesday, October 11, 2016

Does Misuse of Statistical Significance Give Us a Higher Likelihood of Immortality?

This past week, a letter in Nature was published by the Department of Genetics at Albert Einstein College of Medicine in New York. The letter’s title was pithy and shocking “Evidence for a limit to human lifespan.” GASP! “But this can’t be true,” some may say, “with advances in medicine, we’re only continuing to lengthen the human lifespan!” First off, that’s what I think is a bit of a teleological argument – assuming just because humans have certain intelligences about the world, our purpose is to expand life or live forever. There’s some sort of moral purpose that is implied in that answer. Anyways, that’s not the point, but it sets off an interesting point of debate

Maybe we, humans, aren’t supposed to live forever. We forget that we are but specks on the lithograph of time; we’re babies! Only 200,000 years has the species Homo sapiens existed. On a 4.5 billion-year old planet. Many things have lived before us, and many things will live after us. Perhaps this sounds a bit nihilistic, but I was happy to hear these reports from these geneticists. Humans are single-handedly taking the gauntlet to ourselves and our planet, making sure life, in any form, is going to have a hard time existing on a planet where the temperature is rising almost 2 degrees Fahrenheit each year. So yes, I was overjoyed, ecstatic, and relieved to hear we may be shortening the execution of our planet Earth -- until I heard these geneticists used statistics to come to their results. Oh, great, another potential mishap with statistics by those not formally trained in statistics, I thought.

The methods that the study employed were actually pretty simple. Most people who took a basic stats class could probably understand everything up until the cubic smoothing splines. The authors plotted the maximum reported age at death (MRAD) for 534 people over the age of 110 or what they call “supercentenarians” gathered from the International Database on Longevity. These supercentarians came from Japan, France, US and the UK. Two linear regressions were done for subsets of data between 1968-1994 and 1995-2006. The scientists found an increase in MRAD of 0.15 per year before 1995 (with r=0.68 and p=0.0007) and a decrease in MRAD of 0.28 years from 1995-2006. There was a r=-0.35 and a p-value of p=0.27 for the decreasing MRAD. The scientists didn’t expand on the statistical shortcomings of their decreasing MRAD points. The scientists conducted the same procedure with other data points, and found a similar overall trend – that is, a significant increase until breakpoint and an insignificant decrease after breakpoint -- but did not discuss the weak correlation or significance.

The initial criticisms of the work have been predictable, to say the least. Why didn’t the scientists discuss their dismal p-values? Moreover, some say that seeing a significant increase in MRAD then observing a decrease, even if it was insignificant, is still something. This is, however, an amateur cop out, in my opinion. There are so many more reasons to be critical of the work than just the terrible p-values, and it starts with study design.

The main figure in Dong et. al. and the bane of my existence these last four days...

First, it’s really not clear why the scientists decided to use the arbitrary breakpoint between 1994 and 1995. Clarification of this would be helpful, if not crucial, to understand, as the entire crux of the paper’s argument leans on this breakpoint and its subsequent data analysis (Figure 2a). It appears that the breakpoint was chosen arbitrarily to support their initial hypothetical claim that humans have reached a plateau in age advancement, and the linear regression decrease model in MRAD is used in a rhetorical sense, rather than in a strict statistical-sense of proving their ad-hoc null-hypothesis statistical test (NHST). Is it even right to use NHST here? I’d argue no. If their alternative hypothesis was that r=0, it’s increasingly hard to test effect sizes closest to zero – you probably need to have a larger sample size, and with this data, data that exists on the margins of collectability, that’s hard to do.

The authors do make an implied comment on the power of their statistical tests, noting that they probably don’t have a large enough sample size to make case for a more robust statistical model. To alleviate this, they applied a post-hoc sample expansion to include several RADs as the highest point, including the MRAD and the second through fifth RAD; this appears to be more or less a sneaky way of doing away with outliers to me. They concluded that the average RADs had not increased since 1968, and that “all series showed the same pattern as the MRAD.” But instead of applying an ANOVA to test their hypothesis about average RADs not affecting the means, they never reference any supplemental material. It appears that they just eyed it. Another suggestion would be to compare the slopes of the “plateau” regions by doing some sort of t-test.


In the end, the statistical procedures taken to prove their own point would have any data scientist’s head spinning; probably enough so that people just wouldn’t take a hard look at it for too long. This is dangerous, but we can’t say Nature hasn’t committed the crime before: many times, the journal has published funky stats simply because the title of the study was provocative (as demonstrated by our class this semester). Nonetheless, you have to tip your cap at the scientists who wanted to make a headline, they surely did it, then perhaps bonk them on the head with your cap and tell them to do better stats before they make such grand claims.

Thursday, April 28, 2016

A Stock Solution; Asset Pricing Theory



Alright ya'll it's about to get stuffy in here. Something I have strong interest in is global markets and the stock exchange (well, strong for a biochemist with no economics training past high school). Naturally, this sector is perfect for exploring the widespread use of statistics and the unique and powerful ways they can be applied. While disciples of Benjamin Graham will warn that the valuation of a company's stock is not 100% tied to potential for an upward trend, this undoubtedly plays some role in the type of trading done on the market today. More than 50% of trading on the NYSE is done via something called High-Frequency Trading, which uses complex algorithms to buy and sell securities on the millisecond time scale for an overwhelming addition of small differences, resulting in a large profit for the companies employing these buying and trading algorithms. I'm not nearly savvy enough to describe these algorithms with sufficient detail, but I do want to discuss some aspects of these algorithms and other probability/statistics related applications in stock trading.

Investors deal with an immense amount of data, and more is generated every couple of milliseconds. As a result, a combination of instinct, experience, and sound statistical models are the professional trader's bread and butter. Various statistical tests are utilized to assess risk and confidence.

The first technique (and arguably the most central) application of statistics in stock trading is in Asset Pricing Theory. This branch of investment theory uses the calculated effects (effect size!) of various macro-economic factors, or the behavior of theoretical indices (indices track many different stocks, or sectors and are like a mean value representing how a sector is doing. Common examples are the S&P 500 or the Dow Jones Industrial Average) to forecast the expected return of an asset. I realize that all sounds a bit vague, so let's focus on a specific example:

We are all familiar with the T-statistic (departure of a parameter from its notional value and its standard error) which we use in Student's T-Tests. In the case of investing, the utility of this statistic is almost exactly the same as when we would use it to compare means. In fact, one of the foundations of investment statistics is formation and testing of a null hypothesis. In Asset Pricing Theory, the null hypothesis would propose that "the expected return of the asset is not different from the risk-free rate of return". In other words, they compare the asset in question to the performance of risk-free investments like some bonds or savings accounts. Given the historical returns of an asset and the risk-free investment of choice, an investor may find a T-statistic describing the difference between their asset (our sample of interest) and the risk-free investment (background, WT, negative control, etc.). Investors will then use the T-statistic as an indicator of the probability of observing the asset's returns under the assumption of the null hypothesis. Another similarity is that, in this branch of economics, they set their statistically significant p-value as 0.05, and utilize confidence intervals to have a better understanding of how this asset is likely to behave. Similar to our experiments, these calculations also rely on a certain "N", where a single N could be individual transactions involving this asset (in this case, higher volume stocks would be advantaged, due to a greater number of values), but this is not always the case.

This type of statistical testing is vital to many stock analysts, who will use the likelihood that a stock will perform better than a risk-free investment as part of their assessment of how to score the stock (what they should advise their advisees or their firm to do with regard to the stock). It can also indicate when a stock may be "overpriced" or "on sale". In many trading circles, the direction that a stock is likely to go short-term is less important than the absolute value of the company and what a stock of that company should be "worth", hence "Asset Pricing Theory".

Tuesday, April 26, 2016

Multiple t-test in Excel



Doing multiple comparisons can easily get people in troubles like P-hacking. But sometimes we just want to know the difference between particular groups. As we know that Prism can do Tukey’s test as unplanned comparison in post-hoc analysis. It is very efficient, but not very illustrative in my opinion. I came across Student’s t-test in Excel, which helps visualize the result better, as shown below. The significant differences can be color coded automatically for you to see. Also, I liked it because you can work with raw data and you do not to process or copy and paste data in anyway. 

In this example data set, percentage inhibition to an bacterium by extractions from different parts of a plant were compared. In this case, each plant part is a group. Starting with building a table, you do not need half of them since "branch vs leaf" is redundant to "leaf vs. branch".

Next is an important step to make the result visual: conditional formatting. Set the rule as filling the cell with color when the number falls into a range of value in this cell. We can set the range to "0 to 0.05" working with traditional p-value. But notice here, we are attempting to avoid p-hacking so we need to use the p-value for each comparison after Bonferroni's Correction. So remember to divide the p value (usually 0.05) by the number of groups that you are comparing. In my case, 10 comparison, so 0.005 for every comparison. 






Do the t-test by typing the function “=T.TEST." "Array 1" and "array 2" stand for the two groups that you are comparing, again, the sequence does not matter. Then pick one or two tail; I picked two-tailed, so I typed 2. Then choose the type; there are three choices: "1" for paired t-test; "2" for two-sample t-test with equal variance; "3" for two-sample t-test with unequal variance. 

After you do the t-test for all the comparison, your significant results are highlighted in color. 
I have looked at the ANOVAs and Tukey's test in the Excel, but they are more complicated than they are in Prism. But when it comes to seeking for the difference between groups, I liked this color-coded style t-test chart (with Bonferroni's correction) so far.