Showing posts with label publication bias. Show all posts
Showing posts with label publication bias. Show all posts

Tuesday, April 12, 2016

A first look at meta-analyses

I’m going to confess that I had never given meta-analysis much thought before reading this section of Motulsky. I can also confess that I’m really, really glad this has never applied to my research so far.
The subject itself is conflicting. Motulsky introduces meta-analysis as a way to combine evidence from multiple studies, usually clinical trials that are used to determine the effectiveness of some therapeutic. At first read, it doesn’t sound like the worst idea—the fact is, not every study has the resources to follow and collect the thousands of patient samples necessary to determine efficacy of a treatment. Pooling many, well designed, smaller trials offers a solution for researchers interested in meta-analysis. It also offers a giant problem.
A quick google search of “meta-analysis and research bias” yields over one million results. It seems from reading this chapter, that there’s a good reason for this. Meta-analysis lends itself to publication bias, and not necessarily because people bloodthirsty, competitive, publish-or-perish, nightmare monsters. Not that this doesn’t happen, but honestly, performing a non-biased seems incredibly difficult. There’s a huge table of challenges in Motulsky’s chapter on it (seriously, check it out: p. 412 table 43.1). Some of the challenges involved in performing a meta-analysis include needing to seek out ALL relevant data, including unpublished studies and studies published in other languages.
One survey published in The BMJ states that of the 31 most recent met-analyses they examined only 9 included participant data from unpublished studies. Many of the studies included in this survey didn’t list any limitations to their analysis, which leads the authors of this particular survey to strongly caution reviewers when reading meta-analyses.
It seems like there are some ways to detect bias in your meta-analyses by the generation of funnel plots, where small sample bias can be seen as asymmetry on an axis:



As opposed to: 


But even this isn’t uncontentious.

At this point in my researching of meta-analysis, I’m genuinely glad that I have no interest in clinical efficacy of therapeutics. Honestly, I think I would stick to well written narrative review articles.  



Thursday, April 7, 2016

P-hacking and publication bias

iii. P values and statistical significance

(Note: names have been changed)
I glared at the data glowing back at me on the computer screen – the postdoctoral fellow who was mentoring me had mentioned finding a p-value for all the numbers I had. 
“Sorry, Florence, how did you want me to analyze this data again?” 
“Just calculate the SEM for each group, stick the numbers in PRISM, and then look for the p-value and see if the difference is significant.” Then she walked away. Something about a motor neuron prep and her mouse embryonic spinal cords sitting on ice for too long. 
I had no idea how to do what she wanted me to do. What is an SEM? Why aren’t we calculating standard deviation instead? Google searches brought up a myriad of statistics websites that attempted to explain how to derive the equation used to calculate an SEM, but with little context as to why you would even want an SEM in the first place. I tried to recall the statistics class I took sophomore year, but only came up with “the p-value indicates significant results.” …Right? I sat there questioning my own competence before continuing to toil over how to do the calculations on the numbers from my qPCR. When I finally got the numbers and graphs, I was dismayed that there was no significance. When Florence came around again, I showed her the results. 
“Oh. Well, that sucks. But there seems to be a trend. And we’re only at n=2, so I think if we just increase our n, it’ll probably be significant.”


Type of error bar
Conclusion if they overlap
Conclusion if they don’t overlap
SD
No conclusion
No conclusion
SEM
P > 0.05
No conclusion
95% CI
No conclusion
P < 0.05
(assuming no multiple comparisons)
 Rule of thumb provided by GraphPad's FAQ

How many of us have been put in a similar situation or have heard of a situation like this?

Without a strong background or understanding of statistics, I blindly trusted Florence’s logic and choice of statistical analyses – she was a postdoctoral fellow after all. She’s probably done more than two dozen of these kinds of statistics on her own data that granted her her Ph.D. She must know what she’s doing, I reasoned. But that was the danger of scientists who were improperly or inadequately trained to conduct statistical analyses: in hindsight, I realized that 1) few people (or even scientists, me included) actually understand what “significance” really means, and 2) as Motulsky puts it, “once some people hear the word significant, they often stop thinking about what the data actually show.” The scenario I recounted is something Simmons, Nelson, and Simonsohn (2012) termed “P-hacking,” a term that refers to attempts by investigators to lower the P value by trying various analyses or by analyzing subsets of data. Motulsky draws out two ways in which investigators do this: 1) by tweaking data (if one analysis didn’t give a P value less than 0.05, then they tried a different one) and/or 2) by changing the sample size post hoc (stopping data collection if the P value is less than 0.05, but collecting more data when the P value is about 0.05).

One study by Gotzche (2006) looked at comparing the number of publications that reported a P value between 0.04 and 0.06, hypothesizing that if results were published honestly, the number of publications reporting a P value between 0.04 and 0.05 and a P value between 0.05 and 0.06 should be similar. In the analysis, Gotzsche found that there were five times as many papers reporting P values between 0.04 and 0.05 compared to P values between 0.05 and 0.06. The emphasis on statistical “significance” equating as scientific significance ends up skewing the publication of results and data and creates publication bias. I really wonder if some scientists believe that inadvertently p-hacking is a legitimate way to conduct statistical analyses, or if some do it knowing that it is the improper way to generate "significant" results.


Perhaps the fix here is for journals to start requiring authors to submit a short cover note explaining the justification of the utilized statistics to corroborate that they understood why and how the statistical tools were chosen and used. In this way, it could force scientists to not only conduct reliable and properly designed experiments, but also to think more carefully about the interpretation of their results, rather than just trying to force or find significance that might not be there.

Monday, January 18, 2016

We Need a New Paradigm for Scientific Communication


The problem of irreproducible results in science has been tied to increased competition for shrinking research funding and pressure to publish, but surprisingly, reproducibility has been an issue since the beginning of science. At the heart of the issue is the use of statistical analyses to determine which results are significant. 

While the choice of appropriate, standardized, statistical tests is an issue that needs to be addressed, I believe a more fundamental problem is how the narrative of science is taught and perpetuated. Science classes in high school and college often teach the way the world works without emphasizing the way that knowledge was acquired. At best, the seminal experiments and theories of a field are taught as elegant works carefully crafted by brilliant minds (see kekule's dream, Miller-Urey experiment, etc.) While many scientists are undoubtedly brilliant, transmitting scientific knowledge through the narrative of the genius also transmits the expectation that a specific conclusion can always be extracted from a set of data. The drive to "fit" one's data to a particular, neat conclusion can sometimes lead to false conclusions. Related to this is the idea of citation bias, or the preference to publish positive results over negative ones. This leads to an incomplete picture of science, and can obscure potentially useful data.


Overall, there is pressure in culture of science to generate data that supports elegant theories of how the world works. The publication process itself also perpetuates this idea – that scientific data is only ready to be shared once it can be neatly wrapped up in a cohesive narrative. This method worked well when routine data collection and dissemination was scarce, and "scientist" was synonymous with "natural philosopher." Nowadays, the proliferation of scientists and the abundance of data has made the circumstances of data collection as important as the data itself for drawing accurate conclusions. The publication process should involve more frequent and open publication of smaller data sets as a basis for critical discussion - a model embraced by the recent website PubPeer. Ultimately, addressing the problems of bias and irreproducibility will require more open communication about data collection and analysis.     

Sunday, January 17, 2016

Is the Pressure too much?

The pressure to publish is ruining science. Don't get me wrong, publishing data and sharing it with other scientists is critical for the advancement of science. However the issue comes when scientist feel so pressured to turn out X number of publications a year that they cut corners. 

Scientific journals focus on novel findings that can advance the world of science to include in their journals. In general this means a hypothesis confirmed for each article. An article in the economist took out the calculators and did the math, showing that up to 35% of published confirmed hypothesizes are false based on the very statistics that scientists use to confirm or reject a hypothesis. Of course this is assuming the statistics were done properly in the first place.

Many things can go wrong leading to an experiment that biases the research to a certain result. Often things that may bias research are looked over simply because the research would not realize that it could bias the results. In the situation where a scientist feels overly pressured to publish, bias may come about through desperation to confirm a hypothesis.

In a blog on Scientific American, Jared Hovarth mentions that part of science is learning from mistakes and using those mistakes to advance research forward. However with today’s journals focusing on success and limited money for replication, it is hard to weed out those 35% of falsely confirmed hypothesizes. I thing that for science to become more efficient and progress more swiftly, journals should publish more articles that disprove hypothesizes. Not only would this decrease the desperation to confirm a hypothesis, but it would also help other scientist focus their own research.



One of the reasons I have always admired science is because I considered it a pure art to seek out a true answer. When biased gets involved, science becomes less pure and less reliable. Though there are many sources to introduce bias, the pressure to publish is a completely unnecessary pressure to influence results. Publishing should be where the truth is spread and data is shared, not a pressure to doctor results and skew data. 

Saturday, January 16, 2016

Negative Data versus Irreproducible Data



The Economist article “Trouble at the Lab” raises several interesting points about the current state of scientific endeavor.  This article addresses the bias towards publishing “groundbreaking” data, the lack of incentive to conduct reproducibility studies, and the shortcomings of the peer review process in vetting erroneous papers.  A rebuttal article by Jared Horvath published in Scientific American, “The Replication Myth: Shedding Light on One of Science’sDirty Little Secrets”, asserts that irreproducibility is inherent in the scientific process, and therefore not as big of a problem as The Economist says.  However, this article confuses negative data with unreliable data, when the two are fundamentally different.
The Scientific American article uses the discovery of Viagra as an example for the importance of publishing irreproducible data.  Viagra was originally invented as a treatment for cardiovascular disease, but was rebranded when it was found to be more effective at treating erectile dysfunction than heart conditions.  Horvath uses this story as an example for the utility of “unreliable” data.  However, this example misses the point of The Economist article.  The Economist article focuses on the publication of data that is fundamentally flawed, mainly through improper statistical analysis and a bias towards only believing dramatic results.  Rather than being an example of unreliable, irreproducible data, the Viagra study is actually a good example of scientists overcoming their biases and believing their negative data.  Rather than manipulating their data to support the initial hypothesis that Viagra would make a good cardiovascular drug, these scientists rejected their hypothesis and allowed their data to lead them to new conclusions about their drug.
While the Scientific American article makes a valid point that irreproducibility is a longstanding phenomenon in the sciences, I do not agree with the conclusion that irreproducibility is therefore not an issue.  There need to be more incentives for scientists to reproduce each other’s work, so that erroneous findings can be identified and the field can move forward.  Both the Scientific American article and The Economist article illustrate the importance of publishing negative data from well-conducted studies.  By disproving previously held hypotheses, negative data can still be useful in contributing to scientific knowledge.

Friday, January 15, 2016

Publication bias in my own research



Unbiased research is the “gold standard” of science and aims to prevent scientists from achieving “self-fulfilling prophecies” and deceiving themselves into believing an idea or hypothesis is true when it is actually not true or completely due to random chance. Properly applied statistics, experimental design, replication, among other methods are ways to prevent bias from happening, but preventing bias is actually against our normal human nature as discussed in Dan Ariely’s TED Talk.  Additionally, the many forms that bias takes whether it be known or unknown to the person performing the experiments can confound studies and/or hinder scientific progress. Panucci et al. (2010) reviews numerous types and examples of bias in tangible research situations that most biomedical and/or clinicians will find themselves involved in during their careers, whether it be conducting research with such biases or reviewing it. Even though there are many types of research bias, I think that one of the major forms that I currently face and/or am frustrated with as a fourth year graduate student is publication bias, which in my view is the preference of publishing positive results and/or the inability to publish data that goes against the current paradigm of a particular field.

Recently, I designed an experiment to analyze the cytokine response in nonhuman primates with malaria using two different species of malaria parasites. I went through the standard process of generating my samples, developing and designing the protocol, etc. Then, I proceeded to perform the experiment. In one set of samples, I found that Interferon gamma, a hallmark cytokine in malaria involved in immune responses, was upregulated as we had predicted. In my other set of samples, the cytokine was undetectable. I immediately thought, “How could this be? The assay must be wrong because this has been seen before and is found in all malaria infections from mice to humans to monkeys.” Like a good student, I proceeded to validate the data on a different platform and confirmed my results, which took about three more weeks considering the need to standardize yet another assay, place orders, etc. Whenever I presented the results to my advisor, he or she responded, “I am not surprised; I’ve seen this before.” I was shocked because in the literature, this cytokine is always reported, and my advisor had even advised me to use this cytokine as a “calibrator”. However, it turns out that this cytokine has a temporal dynamic during the infection so if one misses a certain time frame, the cytokine may not be detectable. To my knowledge, not a single paper addresses this observation, and this is largely due to the need to have a positive result to get the data published. Obviously, it is much more interesting to talk about an increase in a particular cytokine than no change at all. Furthermore, the dogma and published literature are skewed to presenting this cytokine as central in malaria so if it is not present in detectable levels, researchers and reviewers begin to question the integrity and quality of the experiment and samples.  
            Publication bias thus caused me to lose time in finishing this component of my project. If the literature was not biased based to current opinions and also by the need for a positive result, I and others likely in my situation would not be surprised by such results and spend time triple checking all aspects of an assay, repeating results too many times, etc. Publishing negative results and results that may go against current understanding or dogma should be welcomed in science and are ultimately what will help current scientists address new and interesting questions in novel way. Of course, this opinion is qualified by the fact that the experiments are designed appropriately and replicated.