Showing posts with label reliability. Show all posts
Showing posts with label reliability. Show all posts

Tuesday, April 12, 2016

No confidence interval for car make reliability?



I'm looking to buy a used car in the very near future, and there are some important attributes I'm looking for: performance, a gorgeous interior, 2015, and, yeah, it should be made on the other side of the pond.

But then I look at my stipend provided by my PhD program and I wonder what in the world it was that convinced me to become a biologist. And then think of Toyota.

Out of curiosity, I take a peak at Consumer Reports' 2015 Annual Auto Reliability Survey. And I find the following plot:




A very informative plot. What they've done is survey subscribers on the vehicles they own, in total covering more than 740,000 vehicles. For each reported vehicle a reliability score is calculated and associated with that vehicle's model. A mean reliability for the model is calculated, and all the mean reliabilities for a make's models are averaged to generate the yellow dot, indicating the make's mean reliability. The blue spread indicates the range from the make's least reliable model on average to the make's most reliable model. But why didn't they use confidence intervals? After all, they do claim to be determining "predicted reliability scores."

A confidence interval is used when sampling a population for a value of some characteristic in order to estimate the population's mean for that characteristic. Since samples are being taken as opposed to observing every individual in the population, our calculations will not give us the true mean, but it will hopefully be close. A confidence interval informs us of the range of values surrounding our calculated sample mean that must include the true population mean with a certain level of confidence.

All that is needed to calculate the confidence interval is the sample size, sample standard deviation, and the sample mean. But I can see why Consumer Reports decided not to have their plot display that. As a consumer, once I see the average of the all the models' mean reliabilities, I don't need to see a spread telling me about the true mean. The true mean is not the goal of the reliability report nor the consumer who's in the market for a vehicle. The spread I want to see is the range from the make's least to most reliable model, indicating the variability in the reliability of the make's models. For example, if I was comparing Mazda's reliability to Subaru's, I would note that their means are very close, but the spread of Mazda's reliability is constrained and sits to the higher end of Subaru's reliability, which varies much more. This information might give me more confidence in Mazda's reliability than if I had simply compared the means. The confidence interval may resemble the worst-to-best spreads, but it simply does not mean the same thing. And once we have "outliers" that skew a spread in one direction (see Hyundai, BMW, Ford, Cadillac, Jeep), the fixed ± error of confidence intervals begin to have less meaning to the consumer.

Perhaps if I were a chemical engineer or computer scientist, that Maserati pictured above (no reliability data available, but who cares when a car is that beautiful!) wouldn't be such a distant possibility. But for now, I'm going to look at Toyota. And it probably won't be 2015, either.

Monday, January 18, 2016

Bias: Inevitability and Mitigation

As a graduate student in a microscopy lab, bias in data analysis and presentation constantly lurks in the back of my mind. Deciding how to fairly quantify and represent data that is highly qualitative is a constant struggle for our lab. In fact, the problem of research bias has troubled me since my very first hypothesis-based research project, a high school science fair project in which I struggled with selection bias, confounding variables, and my own flawed expectations that my data should match my hypothesis. These types of bias and myriad more are profiled by Pannucci and Wilkins, but simply recognizing our often inevitable biases will not be enough to minimize its damage. As Jared Horvath discusses, “In actuality, unreliable research and irreproducible data have been the status quo since the inception of modern science”; as such, we have a responsibility as scientists to confront our innate bias, the greatest threat to our credibility. As Dan Ariely says in his TED Talk, checking our expectations and intuitions should be the first step to improving our morality and our research. Is there any perfect solution beyond being aware of our bias? I don’t know. Bias is a huge and multifaceted challenge. But, I think working to improve education in this area and increasing public access/publication of negative data is a good place to start.

I would also like to share a resource not listed in the suggested readings that will be of great interest for those wanting to learn more on these topics. This past Friday, NPR’s Planet Money podcast released and episode entitled “The Experiment Experiment”. In this episode, the hosts talk with Brian Nosek, a researcher at the University of Virginia, about a massive study he led to examine the reproducibility of psychology studies. Briefly, Nosek found that only 39% of 100 psychology experiments in the top journals were able to be reproduced. Like many of the other sources in the recommended reading list, Planet Money explores some of the causes of unreproducible data, such as the pressure for scientists to “publish or perish”, and the bias of journals to only publish positive data, and the lack of publicity for negative data. If you enjoyed Ariely’s TED Talk, this podcast episode is well worth a listen as it is another engaging platform in which to explore the inevitable challenges of bias.

Irreproducibility: a bug or a feature?

Scientific research has an image problem.  Recent identification of irreproducibility issues in biomedical and psychological research has brought to light a seeming epidemic of unreliability in academic research.   However, in order to evaluate this problem, it’s important to think about how reliable science is designed to be in the first place.  As odd as it sounds, irreproducibility is a necessary part of the scientific process.  With any hypothesis test, even in the most strongly designed studies with potential for causative outcomes, there exists the possibility of missing a real effect or having a false hit.  For historical reasons, a 5% false positive rate and 20% false negative rate have become a standard for study design in academic research, but adhering to these characteristics alone is not sufficient to limit the potential for risk in publishing.  In an ASBMB blog, Jeremy Berg does a nice job of summarizing how seemingly stringent statistical conditions can still lead to an unreliable experimental finding.
As with any complicated problem, there are many potential changes to start curing the “reproducibility epidemic” and a solution will probably utilize intervention in many places.  First and foremost, there is a necessity for more statistical understanding among scientists and more peer review on statistical methodology needed.  In addition, as mentioned in the Vox article, the rise of post-publication peer review services like PubPeer and Pubmed Commons allows for continued public discussion and editing related to study design and statistical rigor following a paper’s release.  The rise of independent science media outlets like Retraction Watch allows the scientific community and general public to be better aware of malfeasants and plagiarism in publication, and has also brought to light the extent to which the retraction process is arbitrary and unregulated across journals.  And finally, the rush to publication likely prevents the necessary completion of corroborating experiments that would limit the amount of poorly supported science that is published.