Showing posts with label reproducibility. Show all posts
Showing posts with label reproducibility. Show all posts

Thursday, January 21, 2016

Planet Money Podcast

Hey everyone, after our blog posts earlier this week I found an interesting podcast regarding scientific bias and replication errors. It's about 20 minutes long and a good overview of the subject.

NPR’s planet money, episode 677: The Experiment Experiment

Tuesday, January 19, 2016

Boost scientific data reproducibility with quality standards, honor code


There is indeed a problem in the realm of scientific research when it comes to replicating data, as we learn from an article published by The Economist: Trouble at the Lab. The article discusses several phenomena that contribute to this problem, including bad statistics on the part of the scientists, poor research methodologies, peer reviewing that is inadequate, and lack of access to researchers’ methodological data and software. As a result, scientific research finds vastly challenged its reducibility, a quality which is a pillar of its ascribed objectivity.

There is a striking complementarity between this article and the assigned TED talk by Dan Ariely: Our Buggy Moral Code. Ariely discusses factors that encourage or discourage cheating in people. He discusses that people indeed do choose to cheat, and they choose to do so only a little. Additionally, one of his main points was that when reminded of morality people tend to cheat less. Let us replace Ariely’s “cheating” with the article’s phenomena that contribute to irreproducibility in research. If we do this, “cheating” is essentially not following proper standards of scientific research. From this perspective, we may posit that scientists at least aren’t too improper with their research methodologies, so there must be hope for us! As for reminding scientists of morality, the NIH could create an honor code for scientists.

That may sound funny, but at the same time I envision this: that a future generation will not only have an honor code for scientists but also an enforced formal quality standards for reporting scientific data. Firstly, that future generation with a regulated approach to reporting data will look back to ours and find it absurd that we did not have an enforced formal quality of standards. And I would not blame them, based on data such as that in the article being discussed which indicates that there is a prevalent degree of improper methodology when it comes to reporting data. I conceive that the same generation will look back at our lack of honor code with similar estrangement, for if they were taught to respect a set of morals in research, it would generally move them toward abiding by those morals, shifting the scientific culture. This would likely improve scientists’ desires to implement more proper methodologies in reporting their data, despite any cost to themselves since carrying such costs would not be taboo. Rather, it would be widespread and done in the cause of preserving the integrity of science’s reproducibility and thus objectivity. Oh, and the NIH would make sure these scientists don’t go out of business.

Monday, January 18, 2016

The Morality of Reproducible Data

            As academic scientists, we of course are invested in our research, and (ideally) want to leave our own mark on our respective fields. The data that we generate and publish do not just contribute to personal edification, but also the understanding of a topic on a global scale. With the wide accessibility of scientific journals, data is consistently reanalyzed and published findings applied to new experiments in a growing international research network. Keeping this in mind, it is more critical than ever before that the research community emphasize the importance of sound (and complete) data and experimental reproducibility.

            In his TED talk, Dan Ariely discussed how it was more likely for every person in a room to cheat a little than for just one to completely cheat. Students would give themselves a “4” in lieu of the “2” they deserved, presumably so that they would get a slightly better reward while still maintaining a degree of self-respect. This is sadly applicable in the research world, in the form of “cherrypicking” data or only withholding any negative findings from papers. Findings may often be smudged or spun in a certain light, or an incomplete story presented, not nearly enough to warrant a retraction, but just enough that the findings may not be entirely trustworthy. I agree with Jared Horvath’s Scientific American article in that funding provides a constant pressure for scientists to focus first on generating marketable data, and second on generating complete or valid data. However, while funding is a legitimate concern, I do not think that it is an excuse to perform unviable research or twist results. Granted, this is easy for me to say as a graduate student with a guaranteed stipend, but labs that produce questionable data do more than fail to contribute to science; they actually detract from ongoing research. False data can mislead other researchers who may use these findings as a baseline for their own projects, which in turn could possibly fail or lead to more misdirection. The withholding of negative data could lead other labs to pursue these and waste precious grant money rediscovering what should already be public domain.

            After reading some of these articles, it seems that it should be easier than ever to make sure that research is well-executed, given the formation of organizations, such as the PLoS ONE New Reproducibility Initiative and PubPeer. While these opportunities should be taken with a grain of salt, they seem like a viable means for experts to help fact-check or ensure that results hold true. I do think there is a critical difference between difficult and irreproducible experiments, in that some procedures may have a low success rate due to the necessity for high level of technical skill or specialized setup. However, if even a group of experts in the same field cannot recapitulate a finding, something is likely at fault with the underlying experimental strategy or the published data.

The statistics of statistics

The largest issue, in my opinion, with irreproducibility and bias is scientific research isn’t a lack of awareness, and it is not even the fact that these problems exist. The greatest problem is that there is a lack of understanding of why, and a lack of self-awareness that you personally are capable of bias. To my point of awareness of bias and irreproducibility that is not the issue. As obvious in the homework assignment, there have been many articles “shedding light” on the issue. It is so common that the term irreproducible data is more like a running lab joke then a real day-to-day concern. Most people blame the “perish or publish” culture, which puts a lot of pressure on publishing data as soon as possible, the vague (either on purpose or not) methods sections, and also the use of statistics. The pressure to publish will never go away. The concern with a lack of transparency in methods section is also an issue that arises mostly due to the high pressure to publish, but also because there are small things that make a large difference, and the research is not even aware of these. After my admittedly short time in science (6 years) I do believe that this issue has started to be addressed, and may be a more personal then systematic issue.  

The argument for blaming the use of statistics is a complicated one. This is because statistics can be extremely powerful and necessary, particularly as researchers move toward analyzing large data sets and “-omics” type studies. The main argument is that statistics is either used to freely or in a basic understanding, or that complicated statistics are used to make data seem more significant then truly are. However, as pointed out in Jeremy Berg’s blog, the issue is that scientists do not really understand how the statistics work. Importantly, they do not understand the bias that is inherent in the statistics, and why a significant fraction of experiments cannot be repeated exactly.

If you fully understand the statistics on how easily data could not be reproducible it is easier to swallow the thought that irreproducible data is common place, and has always been a part of scientific research. This is argued by John Horvath, where he stresses that if we accept that most of the data published is false or irreproducible, we can then strive to focus on what is true. As scientists we are taught always to question data or idea, and this does not end just because something is “statistically significant”.  However, if we as scientists can accept this and focus on the ideas being presented, and how we can use the small amount of “true” data to move science further then the irreproducibility is not as huge an issue as we once thought, as long as we are being honest with our data and eliminating personal biases. Meaning it is OK to accept that there is a small level of irreproducibility, but we cannot add to it with our personal biases, otherwise that small level will become very large.

This leads to an issue of awareness. Being aware of our ability to bring bias into our research. As Dan Ariely pointed out in his TED talk, “a lot of people cheat a little bit”. We all have what he refers to as a fudge factor. We are willing to cheat just a little bit, but in general most people (and here I’m really saying most scientists) do not make up data, we simply allow our biases to creep into our experiments both in design and analysis. Dan states that one of these is because of our social norms, we as a scientific community need to educate researcher on how to avoid experimental biases, and we need to accept that our intuition is not always correct.

We Need a New Paradigm for Scientific Communication


The problem of irreproducible results in science has been tied to increased competition for shrinking research funding and pressure to publish, but surprisingly, reproducibility has been an issue since the beginning of science. At the heart of the issue is the use of statistical analyses to determine which results are significant. 

While the choice of appropriate, standardized, statistical tests is an issue that needs to be addressed, I believe a more fundamental problem is how the narrative of science is taught and perpetuated. Science classes in high school and college often teach the way the world works without emphasizing the way that knowledge was acquired. At best, the seminal experiments and theories of a field are taught as elegant works carefully crafted by brilliant minds (see kekule's dream, Miller-Urey experiment, etc.) While many scientists are undoubtedly brilliant, transmitting scientific knowledge through the narrative of the genius also transmits the expectation that a specific conclusion can always be extracted from a set of data. The drive to "fit" one's data to a particular, neat conclusion can sometimes lead to false conclusions. Related to this is the idea of citation bias, or the preference to publish positive results over negative ones. This leads to an incomplete picture of science, and can obscure potentially useful data.


Overall, there is pressure in culture of science to generate data that supports elegant theories of how the world works. The publication process itself also perpetuates this idea – that scientific data is only ready to be shared once it can be neatly wrapped up in a cohesive narrative. This method worked well when routine data collection and dissemination was scarce, and "scientist" was synonymous with "natural philosopher." Nowadays, the proliferation of scientists and the abundance of data has made the circumstances of data collection as important as the data itself for drawing accurate conclusions. The publication process should involve more frequent and open publication of smaller data sets as a basis for critical discussion - a model embraced by the recent website PubPeer. Ultimately, addressing the problems of bias and irreproducibility will require more open communication about data collection and analysis.     

The Necessity of Reproducibility and Two Challenges to the Scientific Research

        The Merrian-Webster Online Dictionary defines ‘research’ as ‘the investigation or experimentation aimed at the discovery and interpretation of facts, revision of accepted theories or laws in the light of new facts, or practical application of such new or revised theories or laws’. As such, we, the scientists, usually cannot predict what the results will be (because they are probably not seen before), neither can we know if the observations reflect the facts. Therefore, logic and methodology form the cornerstones of scientific inferences. Ideally, once we follow the well-established methodology to design study, and we interpret and present the observations in a logic manner, we can claim and convince others that the results reflect the reality.

        In the modern scientific society, people starts to question if this paradigm works. The first challenge comes from the validity of methodology, which were discussed in Pannucci, C. J. and E. G. Wilkins (2010) and ‘Trouble at the Lab’ from The Economists (2013). Briefly, two types of error may be introduced to a study. The systematic errors (biases), including selection bias, information bias, and confounding, dampen the validity of a study and cannot be alleviated by increasing the sample size. In contrast, the random error compromises the precision of the measure of association, and it can be controlled by large sample size and appropriate statistical methods. The validity of methodology itself is one kind of research, which will be improved as it goes on.

        The second challenge is more about the research ethics, discussed in a TED talk from Dan Ariely (2012), Julia Belluz (Sep 2015), and Julia Belluz (Oct 2015). Science itself is simply the pursuit of knowledge about the nature or the human society, but scientific research involves politics — the pursuit of benefit, and the competition of resource. Working in such intense publish-or-perish culture, scientific misconduct or violation of regulation may be used to get seemly ‘astonishing’ results, which would earn the investigators much more funding — temporarily. We should have learned lessons from the scandal of unethical source of oocyte (W. S. Hwang, 2005), the STAP stem cell fiasco (O. Haruko, 2014), and the ‘peer review and citation ring’ (P. Chen, 2014), that the mechanism of peer review is not complete enough to prevent some researches from being overvalued. Besides the improvement of ethical education, PubPeer proposed a new paradigm of peer review — the peer review that involves all members in the same filed. Although currently this mechanism works after the research is published, it remains a good and worth trying.

        Perhaps the most essential and informative debate is the reproducibility. In the article of Jeremy Berg (2013), he stated the low reproducibility in modern scientific researches is due to the low prevalence of ‘correct hypothesis’. While I agree with his analysis in general, I want to point that first, the percentage of correct hypothesis can be increased by carefully designing and interpreting the pioneer studies. Second, the specificity can be improved by replicating the experiments by other lab members or even other labs. A significant result occurring by chance would not be reproduced, a phenomenon called ‘regression to the mean’.

        Above all, I still argue that the reproducibility is an essential part of scientific research. Especially for the natural science, the necessity of reproducibility is embedded in the mechanism of knowledge formation and application. What we should work on is keep improving scientific and statistical methods to avoid biases, and educate the scientific society on research ethics.

How to change a field that seems doomed with bias

With the many recent articles siting bias as an extreme impediment to the purpose of science, it is difficult to not stand on top of the cafeteria table and call B.S. to the entire field. Although this seems drastic, faith in science and the scientists that perform it become more and more futile as knowledge of how bias has infiltrated research becomes known. An article published in The Economist lists several ways that biased and unrepeatable research can become published. How can we change an entire field that consists of researchers partial to find their hypotheses correct, reviewers that don’t have enough time to complete an accurate assessment of the science, or journals that rarely publish needed and essential repetition of previous results?  

The psychology of science is complex, and it seems that bias is unavoidable, especially in the “publish or perish” mentality that exists in the field of academia. Jared Horvath, in an article in Scientific American, states that bias and the lack of reproducible research is not just a recent phenomenon, but can be seen for many centuries previous and among even the most lauded scientists, including Galileo, Dalton, and Einstein. So again, how are we to avoid something that has been innate in our field for centuries?

We as scientists must acknowledge that science not properly designed, analyzed, reviewed and repeated exists and is pervasive in our field, even among our own institutions, departments, and laboratories. We also must acknowledge that this kind of science is not trustworthy and can mislead not only the scientific field but the general public to hope for cures that might not exist—wasting hope, time and funding on biased hypotheses and results. Although changing a field seems impossible, I believe it starts in one laboratory that is willing to fight for good scientific practices (click here for a list of biases common in design and analysis of research written by Drs. Pannucci and Wilkins).


If we truly want to change the field to one that is trustworthy and contains good, unbiased scientific practices, we must become our own “bias police” in which we are adamant about self-critiquing and repeating experiments from our laboratories. As we begin to transform our own laboratories, we can begin to hope for the transformation of our entire field.   

Sunday, January 17, 2016

The bias of priming and replication experiments

           In the TED talk given by Dan Ariely, (Source), he describes several experiments where he tests individual circumstances to test the frequency of cheating among college students. The one that struck me most was where individuals were given twenty math problems to complete in an insufficient amount of time. One member of the crowd was a hired actor, who stood up quickly, announced that they had finished the problems, thus providing an example of someone who appears to have cheated. The results Ariely found were that people in the crowd were more likely to cheat if the actor had a sweatshirt from their university.  Providing the idea that the frequency of cheating is based on our concept of our “in-group”. Ariely says, “if someone from our in-group cheats, we see them cheating, we feel it is more appropriate as a group to behave this way” (Source). This concept reminded me of the idea of priming from the Trouble in the Lab article from The Economist (Source)
            Priming studies are based around the idea that  “decisions can be influenced by apparently irrelevant actions or events that took place just before the cusp of the choice” (Source). In Ariely’s experiment, the actor was that irrelevant event, and the response his actions elicited was based on whether people perceived the actor as one of their own. I couldn’t help but question why this occurs. In a scientific context, where graduate student and mentor conduct experiments, it seems understandable that the actions of one person (the mentor) may set a benchmark for the other (the graduate student). Regardless of whether that standard is negative and conducive to cheating, or positive and a high ethical stance, we as student scientists cannot deny the impact of mentors on our individual approaches to science and data. As a current rotation student, I thought my personal approach to this was decently defined, and hadn’t given much thought to how my decisions are influenced by potential mentors or even lab-mates. Now, I could say I’m in the very least more aware of how others in the lab group impact my approach to science overall.
            Along the same thought of the impact of others on an individual’s actions, Trouble in the Lab also raised the issue of how most published research findings are flawed in one of three ways. The most interesting to me was, “the pervasive bias favoring the publication of claims to have found something new” (Source). This ‘pressure to publish’ concept seems to be the primary motivation factor for much research, and is disenchanting for me. The concept of novel, groundbreaking discovery is always exciting in science, but I don’t feel it should be the target goal. Setting the target so loftily seems to create an environment where this “publish or perish” concept has arisen from, and it primes scientists to chase game-changing discoveries at the sacrifice of personal and scientific standards. An over-arching theme of both articles by J. Belluz (Source 1, Source 2), was how many miraculous cancer drugs lack the significant data to back their claims, as shown by the replications of the original studies.

In regards to reproducibility in research, I can easily see the need for replication studies, but from a funding and career standpoint, I can also agree against them. Not only would it be extremely difficult to get funding for a replicate study, but, as Trouble in the Lab phrased it, “Most academic researchers would rather spend time on work that is more likely to enhance their careers” (Source). I do still agree that, “the community should find effective mechanisms for sharing the results of replication experiments, both successful and unsuccessful” (Source), but my agreement stems more from a desire to have more acceptances of negative/unsuccessful data in science communications that there currently is, rather than to have more communication of replication experiments as a whole. Additionally, the problem with replication experiments presented by Trouble in the Lab that, “only people with an axe to grind pursue replications with vigour” (Source), introduces a problematic sense of bias to the replicated work and thus, at least from my perspective, diminishes the intent behind the replication.
The inciting incident of the actor in Ariely’s experiment primed the group to feel at ease with cheating, and thus increased the frequency that cheating occurred. Overall, I feel these articles have provided a sense of individual ethical standards, where a sense of alertness is necessary, in regards to how the behaviors and choices of those around us impact our own biases. In our cases, these would be our fellow scientists and lab mates in regards to how they conduct experiments, or handle, present, and manipulate data, and how those decisions thus color our individual practices and the bias we apply to them. The quote from these articles that best encapsulates my feelings in regards to priming and bias in science is, “each researcher has a responsibility to ensure that his or her own published work is as reliable as possible within the limits imposed by resources and other constraints” (Source).