Showing posts with label Bias. Show all posts
Showing posts with label Bias. Show all posts

Tuesday, January 23, 2018

Is The Strained Economics of the Scientific Enterprise A Significant Cause for Scientific Bias?

Economists are the most peculiar kind of scientists.  In fact, if I were not drawn to the world of medical technology innovation, I would certainly entertain my nerdy-streak by becoming an economist.  Anyone who has read the award-winning book Freakonomics might agree (aside: Freakonomics is a brilliant podcast for the intellectual-at-heart). So what exactly does the economy and the scientific enterprise have to do with bias in academic research?

I may be biased, yet I believe economics has everything to do with it.  We as human beings are susceptible to incentives, however benign or malignant. Any science-minded individual who keeps a beat on the news will know that irreproducibility and moreover, retractions of manuscripts is on the rise globally.  Indeed, while a recent article .  Wherein does the culpability lay?



I argue, incentives unduly influence individual investigators yet publishers are also to blame.  It is well known that funding for the scientific enterprise in the United States is at an all-time low when controlled for costs of inflation since the termination of the NIH budget doubling in the early 2000s.  With less access to funding and a glut of Ph.D’s entering the academic job market (a worth subject for another discussion), researchers must to more with less in order to publish. Fellow blogger Austin Nuckols is wise to note “the culture of science, especially in the academic setting, follows a mantra of “publish or perish”.  The circle of life for academic research is an ultra tenous one driven by supply and demand of the NIH dollar: Win grantàperform researchàpublish à repeat.  One break in that chain is enough to sink a mid-career academic’s productivity (not to mention salary support). When jobs are uncertain every few years, it is easy to see where bias can come top-down, influencing the un-empowered graduate student to conduct research with significant bias, leading to conclusions “in our own image”.   Publishers are similarly incentivized to avoid reducing bias, despite calls to do so in high profile journals (e.g. Nature, Cell).  “Novelty” sells; and who can remember the last time a reproducibility study was featured in the high-impact “Vanity” journals?


Looking at this dismal state of affairs for the budding researcher, I feel incentivized to begin the inaugural edition of The Journal of Research Reproducibility or better yet, The Journal of Failed Experiments (And How to Avoid Doing Them).  Perhaps then, the odds of academic success in research will be in my favor. 

Monday, January 22, 2018

Publish... and Perish?

We like to think that scientists are innately good people, spending countless hours at the bench to help catapult us into the next generation of life-saving medication or medical procedures. On the surface it seems wholesome and altruistic, but diving deeper into the scientific community it becomes apparent that there is a very large elephant in the room: the issue of bias and irreproducibility. In an article published by The Economist the author describes the cut-throat culture that academia has established and how it leads to bias; for example, the motto “publish or perish” may influence researchers to embellish their work in order to publish in high-impact journals or to even publish at all. These high-impact journals then fight back by having egregiously high rates of rejection for manuscripts, leading researchers to cherry-pick their data further to make the cut. The author then goes on to state that companies like Bayer and Amgen failed to replicate more than half of studies they found on breakthrough cancer research, a section of research that is highly esteemed by scientists and the general population alike. 

 The scientific community has created a vicious cycle that seems to keep growing. This immense amount of pressure is leading scientists to falsify or alter data to fit a specific agenda, and soon it will cost them more than their reputation in the field. Flawed research costs us time, money, resources, and the trust of the general population. This puts scientists at a bit of a crossroads, but I think it is up to us to begin making the changes necessary to fight bias. Ethics should be taken more seriously and started even before entering graduate school even though sometimes it can seem like a “no brainer.” Additionally, having a grasp on statistical analysis is imperative for all scientists and not just the PI; if more people understand statistics then it may yield more sound data or make it easier to spot falsified data instead of relying on someone’s best judgement. 

Monday, January 16, 2017

Conflicts of Interest and Personal Bias

Ultimately a discussion on conflicts of interest and personal bias is centered around the concept of ethics. How one conducts their personal and professional business. As Dan Ariely discuses in his 2011 TED talk Beware conflicts of interest, all people have inherent bias that is shifted towards their own personal interests. It is not necessarily malicious behavior or meant to harm or take advantage of others. As Dan puts it, we’re very good at being “blinded by our own incentives.” We are always interested in what suits us best.

However, in some cases this inherent bias can lead to destructive practices. In Dan’s 2013 TAM talk The honest truth aboutdishonesty, he imagines the thought processes behind they American financial crisis of the past decade. By reducing transparency on financial decisions and having been shifted so far away from the persons that would suffer from poor decisions, the possible consequences seem less severe. Furthermore, by creating incentives to do what would be considered wrong or unethical makes the wrong action seem justifiable because the personal outcome becomes a positive one. What’s more, the further we are removed from our actions the better we feel about misbehaving. He touches lightly on the idea of cognitive dissonance, without getting into it. The idea that in some cases we convince ourselves that our actions are not wrong or unethical, but in the right and justifiable.


Ultimately, personal and professional biases have impacts on others. But these impacts are avoidable and in some cases, unnecessary. Perhaps by identifying our own personal biases we can intentionally avoid them, removing conflicts of interest and possible breaches in ethics. The result would be a society of people that treat each other appropriately and conducts themselves in a respectable manner. It sounds like a pipe-dream, but perhaps if we took the time to consider how our actions affect others, it wouldn’t take very much to see positive change. 

Saturday, October 8, 2016

Heteroskedasticity is the Statistical Concept We All Took for Granted

Heteroskedasticity. Hard to say, and often forgotten, but something that should be at least considered by many data scientists and people who work with data. Data is neither good, bad nor ugly, but it can take on such qualities when we analyze that data. Newly minted-PhDs can get their first R01 grants (good), scientists can waste a new grant (bad), or scientists can gain funding fraudulently (ugly). Therefore, how we think about and interpret our data is important step to determining the data’s fate. Heteroskedasticity is one of the concepts that can determine whether your model for your data is good or bad. However, before we explore heteroskedasticity, we need to look at statistical modeling, study design and analysis principles that precede it.

From the beginning, one imperative that seems to be eluding most scientists with data these days is Hyman’s Categorical Imperative. Never met or read any of Hyman’s stuff? Well, it’s just a fancy rule for stating what most data skeptics try to follow. Hyman’s Categorical Imperative -- a maxim coined by Professor Emeritus of Psychology at the University of Oregon, Ray Hyman – states that before you choose to investigate if a phenomenon is true, you should first determine if the phenomenon is real. Essentially, violating this imperative is the scientific equivalent of flippantly adding a trend line to data based on what the data appears to behave as. This is usually manifested in scientific studies as a linear trend line, often after least squares regression analysis, and can be seen in the following examples (with captions).

Figure 1 from AJP Endo Journal. Change in prolactin secretion vs. change in leptin secretion over 24hrs. The model’s fitting of least squares linear regression states that as leptin decreases/increases over 24hrs, prolactin secretion acts similarly. However, a quick scan of the data shows points in the plot populating mostly the -20 to 0 range in both axes. It consistently under-predicts positive and some negative changes in the hormones.

Figure 2 prepared by data scientists Ng and Blanchard. Deaths per state as a linear function of obesity rate. The trend line predicts a linear relationship between obesity rates and deaths per state. However, the model under-predicts some states with high death rates and low obesity rates. The data seems to be plotting an explanatory variable with no firm correlation to the dependent variable (more on this in the blog post).

Figure 3 from data scientists at TIKD app. Tickets per capita as a negative linear function of income. The trend line predicts that as income goes up, ticket issuance goes down. However, the model under-predicts all high incomes (>~$82,000) and even under-predicts low incomes (<$40,000), showing the influence that data point density can have on linear least squares regression models.

As we have discussed in class, least squares can be a good model for some data, but not all data (see above figures). The problem with least squares regression analysis is that it happens to be unusually influenced by where large amounts of data points accumulate or where unusual points scatter in Cartesian space, i.e., the variation in variance of the data points. The figures above are all good examples of this.

Generally, when data scientists want to look at how well a model works, they look at the residuals of the observations and the model’s predictions (data point value-model prediction value). If there is a pattern that appears due to the model’s application, then the model probably isn’t the best fit for the data. In other words, the model could be consistently over-predicting or under-predicting past or future points of data.

So what exactly is this unusually long word and what does it have to do with what we mentioned? Well, as described by Stats Make Me Cry, heteroskedasticity “refers to the circumstance in which the variability of a variable is unequal across the range of values of a second variable that predicts it.” Or, more simply, when the variance of data is conditioned on some other variable not shown. For the figures above, that could be some other variable like another hormone (Figure 1), an environmental source negatively affecting health in the state (Figure 2), the uneven policing of certain communities (Figure 3), or the cautiousness of drivers in older ages who just so happen to have higher incomes (Figure 3).

Why is it that heteroskedasticity gets no love in the classroom then? Well, it’s something that can typically be ignored when scientists analyze data. That’s because the presence of heteroskedasticity doesn’t truly bias the fitting of least squares regression, but it will lead to problems when analyzing variance, like in ANOVA tests. You can’t assume homogeneity of data without first analyzing heteroskedasticity of the data. Finally, heteroskedasticity might not get a lot of love from the life sciences, but it is a huge deal in the social sciences, especially in economics, and the subfield of econometrics, or the branch of economics that uses math to study the behavior of economic systems. Economists rely on homogeneity of data to give reliable confidence intervals and hypothesis tests. Maybe us life scientists should start looking for this then, too. So remember, the next time you assume homogeneity of data, make sure you use a logarithmic transformation to account for heteroskedasticity, or use these other cool tricks found on page 6 of Richard Williams’s lecture from Notre Dame

Thursday, May 5, 2016

Hans and Ola Rosling: How not to be ignorant about the world


This blog post deviates from the topics learned in our biostatistics course, given it focuses on how we can use statistics to get a point across. Since Dr. Murphy gave great importance to bias and I have an interest on public health, I found this TED talk about reducing bias in our knowledge of global population and global health interesting. Hans and Ola Rosling use statistics to prove to the audience that they have a high statistical chance of being wrong about what they think they know about the world. I believe it was an interesting use of statistics, because, as we have mentioned in class, people pay more attention when they are presented statistics (even if these statistics are wrong), so it was smart of them to use this method to get their point across!

            The main point of their talk is to improve the knowledge people have about what is going on in the world because, as Hans states, “the best way to think about the future is to know the present”. However, this is sometimes hard because of three skewed sources of information, which they recognize as:

1.     Personal bias: the different experiences each person has, depending on where they live and the people they are surrounded by
2.     Outdated information: what teachers teach in school is usually an outdated world view
3.     News bias, which is always exaggerated

Hans then adds intuition to this equation. He states that we seek causality where there is none and then get an illusion of confidence (which, as mentioned in class, can happen to many researchers). In order to combat it, he says that first we need to measure this false sense of confidence, in order to then be able to cure it. This way, we will be able to turn our intuition into strength again.

            Hans’ son, Ola, then states four misconceptions people have about the world and then demonstrates the counterpart of each misconception, or how people should be thinking. The four misconceptions and their counterparts (shown after the arrow [à]) are the following:

1.     Everything gets worse à most things improve
2.     There are rich and poor à most people are in the middle
3.     First, countries have to be rich in order to get the personal development à first a country had to work on their social aspects, then they will get rich
4.     Sharks are dangerous (meaning: if you are scared of something, you are going to exaggerate) à sharks kill very few


Ola believes that if people change their point of view about the world by looking at the facts, they might be able to understand what is coming in the future.

Sunday, April 10, 2016

Patterns Real and Imagined



One of the interesting topics covered in Harvey Motulsky’s first few chapters is the argument that probability is not intuitive because people tend to identify patterns even when none are present. He provides the example of basketball players being perceived as more likely to make or miss their next shot based on their current “streak” of successful or unsuccessful shots. As a side note, I think this is a rather poor example of the point, because it suggests that whether a basketball player makes or misses a basket is based on random chance as opposed to all the other factors that go into it. A better example is the randomly generated table provided on page 5, which could be interpreted as depicting patterns. As Motulsky points out, humans are adept at identifying patterns because it is evolutionarily advantageous to do so. I would propose that scientists are more perceptive than the average person to hints of patterns, as we are trained to detect regularities that point us to the underlying mechanisms that govern our world. That also means that we are decidedly prone to introduce bias into our work even with the best of intentions. If in the course of an experiment we start to see a trend emerging, we tend to look harder for more data that fit that trend. I discovered this during my first pilot experiments that measured disease severity in mice, and I have performed all subsequent experiments of this type blinded to genotype. Blinding is not a universal fix for this problem, though. Another instance of possibly spurious pattern recognition in data that comes to mind is multimodal populations. If you look at a scatter plot and see points clustered in what seem to be two groups, it is tempting to think that perhaps they reflect a bimodal response to a variable. Flow cytometry is another area where this can occur, as it is often possible to identify numerous populations that seem to express different combinations of marker intensity. In the complexity of biological systems, the possibility that these “patterns” in the data represent truly distinctive physiological entities is very real, and especially in more variable systems such as human studies or experiments with outbred animals, it is not at all unlikely that subsets of individuals could exhibit different responses to treatments that are based on underlying physiological differences. For instance, studies relevant to our lab’s work have found that a subset of depressed patients exhibit high levels of inflammatory markers and that their depressive symptoms can be improved with anti-inflammatories(Raison 2013). So it is important for scientists to recognize and pursue patterns that may lead to outcomes like this. But we must also recognize the potential for bias that comes if we choose to focus only on one perceived population of a dataset that “behaves better,” and also the potential to miss interesting findings by subdividing populations to the point that we lose experimental rigor.

Thursday, April 7, 2016

P-hacking and publication bias

iii. P values and statistical significance

(Note: names have been changed)
I glared at the data glowing back at me on the computer screen – the postdoctoral fellow who was mentoring me had mentioned finding a p-value for all the numbers I had. 
“Sorry, Florence, how did you want me to analyze this data again?” 
“Just calculate the SEM for each group, stick the numbers in PRISM, and then look for the p-value and see if the difference is significant.” Then she walked away. Something about a motor neuron prep and her mouse embryonic spinal cords sitting on ice for too long. 
I had no idea how to do what she wanted me to do. What is an SEM? Why aren’t we calculating standard deviation instead? Google searches brought up a myriad of statistics websites that attempted to explain how to derive the equation used to calculate an SEM, but with little context as to why you would even want an SEM in the first place. I tried to recall the statistics class I took sophomore year, but only came up with “the p-value indicates significant results.” …Right? I sat there questioning my own competence before continuing to toil over how to do the calculations on the numbers from my qPCR. When I finally got the numbers and graphs, I was dismayed that there was no significance. When Florence came around again, I showed her the results. 
“Oh. Well, that sucks. But there seems to be a trend. And we’re only at n=2, so I think if we just increase our n, it’ll probably be significant.”


Type of error bar
Conclusion if they overlap
Conclusion if they don’t overlap
SD
No conclusion
No conclusion
SEM
P > 0.05
No conclusion
95% CI
No conclusion
P < 0.05
(assuming no multiple comparisons)
 Rule of thumb provided by GraphPad's FAQ

How many of us have been put in a similar situation or have heard of a situation like this?

Without a strong background or understanding of statistics, I blindly trusted Florence’s logic and choice of statistical analyses – she was a postdoctoral fellow after all. She’s probably done more than two dozen of these kinds of statistics on her own data that granted her her Ph.D. She must know what she’s doing, I reasoned. But that was the danger of scientists who were improperly or inadequately trained to conduct statistical analyses: in hindsight, I realized that 1) few people (or even scientists, me included) actually understand what “significance” really means, and 2) as Motulsky puts it, “once some people hear the word significant, they often stop thinking about what the data actually show.” The scenario I recounted is something Simmons, Nelson, and Simonsohn (2012) termed “P-hacking,” a term that refers to attempts by investigators to lower the P value by trying various analyses or by analyzing subsets of data. Motulsky draws out two ways in which investigators do this: 1) by tweaking data (if one analysis didn’t give a P value less than 0.05, then they tried a different one) and/or 2) by changing the sample size post hoc (stopping data collection if the P value is less than 0.05, but collecting more data when the P value is about 0.05).

One study by Gotzche (2006) looked at comparing the number of publications that reported a P value between 0.04 and 0.06, hypothesizing that if results were published honestly, the number of publications reporting a P value between 0.04 and 0.05 and a P value between 0.05 and 0.06 should be similar. In the analysis, Gotzsche found that there were five times as many papers reporting P values between 0.04 and 0.05 compared to P values between 0.05 and 0.06. The emphasis on statistical “significance” equating as scientific significance ends up skewing the publication of results and data and creates publication bias. I really wonder if some scientists believe that inadvertently p-hacking is a legitimate way to conduct statistical analyses, or if some do it knowing that it is the improper way to generate "significant" results.


Perhaps the fix here is for journals to start requiring authors to submit a short cover note explaining the justification of the utilized statistics to corroborate that they understood why and how the statistical tools were chosen and used. In this way, it could force scientists to not only conduct reliable and properly designed experiments, but also to think more carefully about the interpretation of their results, rather than just trying to force or find significance that might not be there.

Monday, April 4, 2016

Coincidences and Bias

In the "Introducing Statistics" section of Intuitive Biostatistics, the author explains how probability and statistical thinking are not, in fact, intuitive.  As humans, our brains are hardwired to look for patterns, even in data that was actually randomly generated.  As an example of this, the author points out that coincidences are actually much more common than we realize because "it is almost certain that some seemingly astonishing set of unspecified events will happen often, since we notice so many things each day" (p. 5). 

I have often noticed that soon after I learn a new word, I will repeatedly hear that word used over the next several days, even though I could never remember hearing it before in my life.  This happened to me a lot when we learned SAT vocab in high school English, I remember repeatedly hearing and reading the word "gregarious" outside of class after first learning what it meant.  Apparently this is a common enough phenomenon that it has a name; it's called the "Frequency Illusion" (or the "Baader-Meinhof Phenomenon").  Basically, learning a new word primes the brain to pay more attention when that word is heard again.  I probably heard "gregarious" no more often during my sophomore year of high school than in the previous fifteen years of my life, but because my brain was more aware of this word right after I learned it, it seemed like I was hearing it all the time. 

This example illustrates how pattern-seeking and unconscious biases can influence our perceptions, even in the most mundane of situations.  Because scientists are also human, our interpretation of our findings can also be affected by these cognitive biases.  Statistics and rigorous experimental design are imperative to prevent our biases from clouding our scientific judgement, causing us to seek results where no pattern actually exists.