Showing posts with label Fitting Models To Data. Show all posts
Showing posts with label Fitting Models To Data. Show all posts

Tuesday, May 3, 2016

Wanna be good at golf? Just swing as hard as you can!


A lot of people want to be good at golf, but to any avid golfer its pretty clear that almost nobody is. This is even further illustrated by the fact that only 1% of all golfers will play a round at even par or better in their lifetime. So why do so many flock to an expensive game at which they are for all intents and purposes guaranteed to be terrible at? I say because we like a challenge. Others might say masochism. But I bet your asking yourself what would statistics say? Well, as it turns out statistics either can’t (or won’t) answer that question.  What statistics did say was that if we showed her some data from a lot of golfers then she could maybe give us some insight about why everyone sucks. Who knew statistics were so rude? Anyhow, the message clearly got passed along the grapevine and there were a few guys in Australia at Drekin University that got super interested in why people aren’t good at golf. And let me say they were pretty convinced they had it all figured out too. They hypothesized that swing speed (how fast you normally swing any given golf club while hitting a ball) and handicap (the most commonly used indicator of golf skill) would be highly correlated.  “That’s not a bad guess and this is actually really important to the world so here’s some money to test that” said the University and so they tested their hypothesis. 
They measured the club head speed of 45 golfers aged 18-80 who self-reported handicap indexes ranged from 2-27.  What they discovered in the data was a visually striking correlation between club head speed and handicap. Fantastic they said. Now we just need to fit a model to this data and we can accurately predict how much someone will suck at golf. Sure enough linear regression was used to determine that indeed club head speed and handicap were significantly correlated (r=0.950). In my opinion it is pretty surprising that such a simplistic readout like swing speed would be so highly correlated with the complex (and infuriating) process of getting a golf ball off the tee box and to its home (the hole) 18 times. For one, the metric of swing speed
doesn’t even account for putts which could account for up to half of the strokes the handicap index is based on. Furthermore, hitting a ball hard doesn’t make it want to go in the hole. I think we can all agree Happy Gilmore definitely proved this point.  Regardless, what we can take away from this is that players with higher swing speeds are simply more likely to be better than players with lower swing speeds. Now is it that lower swing speeds are causing players to suck? I would venture to say no. Furthermore, the data clearly only speak to correlation and not causation. But just to be on safe side, next time you hit the course make sure you swing hard; that is if you care at all about playing well.

References:


Tuesday, April 12, 2016

Fitting the right model to data

When it comes to fitting models in data, we need to be careful avoiding fancy mistakes. The regression functions can get pretty complicated, which matches any of our wish for letting the data be more explanatory. At this point, we could over-interpret the data, and neglect that we can replace the models with other analysis that make more sense in our research context.

I just finished my honor thesis and I had to persuade myself for not playing around what I have learn from this advanced statistic course for an undergraduate student. The research that I worked on is about whether a plant, GBL, can inhibit growth of a bacteria, ATCC6919. We are interested in this bioactivity because it can be an alternative cure to infection caused by this bacterium.
Extractions of this plant were made from 3 different parts of the plant, leaves, branches, and seed, and two different extract solvent was used: ethanol and water. The question that my data analysis need to answer is not only whether the GBL extracts is active against ATCC6919 growth (% inhibition> 50%), but also whether the tree parts and the extract solvent contribute to effectiveness of the extracts.

The ATCC6919 culture was treated with GBL extracts at a range of concentration. So the result of the antibacterial investigation will generate many dose response curves, like the one shown below. Since the trend of the plot is clear that at the percentage inhibition is higher at higher extract dose, there is a pretty good chance that we can find a regression model which fit most of the data well. I could build a “dose response regression model” for the extracts. However, I recalled that we were interested in finding the extracts that were active (% inhibition > 50%). Therefore, a regression model could be a statistically perfect fit, but it is scientifically non-sense. I have to discard the idea of nonlinear regression curve model.

Then I thought about, could I compare the difference in inhibition result from the two extraction solvents by comparing the fit of the data to two models. The best-fit slope of the regression line should be the differences between two group means. Thus, I set the variable defines extraction method X, and assigned X=1 arbitrarily to aqueous extracts and X=2 arbitrarily to ethanolic extracts. Y axis was the percentage inhibition of the extracts at same concentration. It would look like the linear regression graph shown below. However, if so, I neglected the other factor, which is the tree parts, which can also contribute to the difference in inhibition. I could meet a problem opposite to over-fitting the data, which is over-simplify it.


If we replace the regression models for two-way ANOVA, it is easier to see whether plant parts, or the extraction method, or the interaction of them make inhibition of the extracts differ. If you want it to be more basic, multiple student t-tests would work together, too. 

To sum up, when we try to fit the models to data, before thinking about which certain type of regression model fit better, check if other (and simpler) method fits more.