Showing posts with label #statistics. Show all posts
Showing posts with label #statistics. Show all posts

Monday, January 22, 2018

Bias, Nicolas Cage, and public health

The recent discussion surrounding irreproducibility in scientific experiments should come as no surprise to those in research fields. In fact, it is a discernible responsibility of a researcher to question results, and to explore potential sources of error and bias in experiments. This should be true whether the results come from one's own work or from a literature search.

As the number of PhD-seeking scientists increases and technology advances, it is not uncommon for researchers and science writers to feel pressure to inflate the generalizability of their results. This is especially true with the advancement of fast media. The “click bait” culture of the internet could inadvertently encourage the exaggeration of scientific claims, leading to “science” articles with idiotic headlines like “thinking makes you fat” (Yes, this journalist was actually writing about their interpretation of a real study).

Yet, this increase in pressure for strong results should be (and seems to be) coupled with a strengthening of scientific rigor and scrutiny with regard to interpreting results. I’ve heard countless anecdotes about the types of papers or results that you “used to be able to get away with” decades ago that would not withstand peer review today. The scientific standard may be increasingly challenging, but if our goal as scientists is to get closer to understanding the truth about cancer biology, neurodegeneration, etc; then we should be prepared to endure these challenges. Double-blind studies, control groups, grant review, and other measures to reduce error and bias may add challenges but they are absolutely necessary in order to produce more reliable findings. I think these are a good start, but I don’t think they are sufficient. In fact, all of these human-driven solutions are to fix a human-generated problem, as mentioned  here. In other words, while peer-review certainly helps produce more rigorous studies, these reviewing humans face the same potential pitfalls as the authoring humans.

Apart from the challenge of identifying the sources of bias and irreproducibility in interpreting data, there is a tremendous responsibility for those reporting science to the general public. Laypeople without scientific training may not have as much practice critically analyzing the controls used in an experiment, or understand the potential pitfalls of a particular method. I feel strongly that people in general, not just scientists, should understand a great deal more about statistics, bias, and how to think critically about what they read and watch.

Take for example, one of my favorite websites, Spurious Correlations, graphs correlations such as “The number of drowning deaths each year correlates with the number of films Nicolas Cage starred in that year. This is an example of a time when the results of a certain statistical analysis may tempt a naïve person into a grandiose inference about the data. This misinterpretation could be innocuous, like in the example above, but it could be an outright dismissal of scientific evidence in favor of a popular non-scientific idea - e.g. “vaccines cause autism” – that could halt progress and put public health at risk. For these reasons, it is the job of scientists and educators to steer towards a better understanding of bias in our system and fight it with whatever tools we have, while understanding and communicating the limits of scientific experimentation.


Sunday, April 23, 2017

Statistically speaking, are you going to snap me back?

Snapchat is a social media platform that allows people to send pictures for a designated period of time. It is swiftly rising to the top of the social media food chain. The easy to use interface is equipped with filters that allow users to jazz up their selfies with backgrounds or costumes.

There are currently 158 million people on Snapchat. There are 80 million American users and 100 million people between the ages of 12- 34 in America suggesting that there is a 80% chance that if you send a Snapchat, as a millennial, you will get a snap back. These numbers are used to generate probabilities. In probability theory, we consider randomness and uncertainty in models but in statistics we make observations to figure out a process.  To generate statistics on whether someone will send you a snapchat in return you would need to establish a relationship. You would need to keep track of past events and establish patterns. It begs the question how can we use randomness to determine statistics.

Chevalier de Mere was a gambling man interested in seeing if the “house” truly always wins. In a game of dice a player would roll four dice and the player would win if no six appeared on any die. The Monte Carlo was created by George Louis LeClerc. The simulation is essentially a coin flip simulator. He used the principle of Buffon’s needle which can be simulated on: http://www.metablake.com/pi.swf. He estimated that the probability of the needles crossing when being thrown randomly is 2/π. The Monte Carlo method has since been modified many times but requires three things: modeling of a system of probability density functions, repeatedly sampling from these functions and computing the statistics of interest. All in all, simulation of data and performing statistical analysis is the intersect between statistics and probability.


Unfortunately, millennials, until a great statistician like you models the probability density function for Snapchats returned we remain with our sad return probability of 80%. Hopefully you become BFFs with Kylie Jenner so she can increase your probability. Until then… Happy modeling!

Tuesday, January 17, 2017

Little Cheaters

Dan Ariely in his talk, ‘The Honest Truth About Dishonesty’ at The Amaz!ng Meeting 2013 introduces the concept of little cheaters, that is, people who are dishonest in ways that they consider small enough to maintain personal morality while still reaping benefits of dishonesty. This concept was derived from studies in the general population suggesting that scientists too are privy to such behavior, but what implications does this have for science?

The most likely effect of dishonesty in science is irreproducibility. If experiments are planned, executed, interpreted, or reported with even the slightest amount of dishonesty, they are impossible to repeat by others. Consequences extend beyond those who seek to replicate to those attempt to build on the existing work as they would be working off likely incorrect information. Such deception is clearly undesirable but eliminating it can be difficult as perpetrators may not always be aware of their deception because they perform it while convinced of their morality. This is further compounded by the inherent conflict of interest that exists in all scientists. Every researcher holds stake in the success of their work: graduate students benefit from publishing papers and graduating early, senior investigators gain career advancement and increase their marketability for grant funding by presenting positive results. All these factors color the objectivity of researchers making it harder to recognize the subtle ways in which they can be dishonest such as inflating the meaning of their findings or omitting unfavorable results. Proper statistics should be able to check this bias but it is no secret that many laboratory scientists are not sufficiently conversant in statistical methods.


What then, does the combination of dishonesty, bias, and poor statistical knowledge mean science is doomed? Should presenting work be put off until these problems are eliminated? No. Rather, science needs to be redefined as the work in progress that it is and not the subject of irrefutable answers as perceived by many. Efforts should be taken certainly, to minimize blatant falsehood in published work, but it should also be acceptable to not be quite certain. Scientists will be more likely to shed their little cheater identity when it is fine to have work that does not completely make sense.

Monday, January 16, 2017

Bias is human nature

Whenever I think about the issue of bias and irreproducibility in science, there are two quotes that come to mind. 
“73.6% of all statistics are made up.” – Mark Suster 
The second quote was popularized by Mark Twain, who attributed it to Benjamin Disraeli: 
“There are three kinds of lies: lies, damned lies, and statistics.” 
How do these quotes relate to the issues of reproducibility and bias in science? The irony of the first quote is it is itself a made up statistic, meant to demonstrate that people will parrot figures without first validating their veracity. The second quote highlights that statistics can be deceitful if misrepresented. Combine misleading statistics with the repetition of false information, and you have a crisis in the validity and reproducibility of scientific data. You do not have to go far to find proof of this phenomenon: This article discusses the source of hype around new cancer drugs, which stems from both journalists and scientists repeating statistics without understanding the full context of the situation. Yet, it is not just scientists who do this. How many times have you or a Facebook friend read a statistic and then repeated it, without understanding where that number came from? Misleading facts combined with repetition without confirmation means it is very easy to fool ourselves into thinking there is something in the numbers when in actuality, there is nothing.
To demonstrate how easy it is to fool ourselves, take a look at the graph below: 
These graphs look related, right? An r value of .666 is not terrible. Let’s add some labels.


Do you believe this graph? It seems pretty reasonable, right? But let’s look at what the graph actually represents.


Surprise! This a spurious graph where the two variables have nothing to do with each other, yet look related because of the way the data is represented.
The point is, statistics is tricky. It is easy to ignore facts and justify what we want to see, especially when it benefits us. I think this may be a big reason why science is currently in a data crisis; it is not necessarily out of intentional malice, but rather because human beings are inherently bias, and we inherently make connections, false or not, between data sets. Of course, there are those who intentionally falsify data or manipulate data to fit their theories, but that’s a whole other topic.