Showing posts with label selection bias. Show all posts
Showing posts with label selection bias. Show all posts

Tuesday, July 23, 2019

How p-values of underpowered studies select for exaggerated effect size

When we think of experimental power, we're usually in the mindset of setting things up to have  enough power to detect a treatment effect. In fact, we should also think about detecting an effect size accurately. It turns out that underpowered experiments not only make it less likely to detect effects, but are also more likely to lead to an exaggerated effect size.

Here's a simulation to illustrate this point.

We simulate a run of unpaired two-sided t-tests that compare a drug effect to a placebo effect at each of two different power levels. Because we define the parameters of the sampled population, we know for a fact that the response level to the drug is 350 response units (SD = 120). The response level to the placebo is 200 response units (SD = 120 also). The variables are assumed to be normally distributed.

Thus, in testing over the long run, the drug would be expected to have an average effect of 150 response units over that of placebo. How accurately do experiments conducted at low and high power estimate this effect size between the drug and placebo?

The experiment is simulated at low power (~50%, sample size per group = 6) and at high power (~90%, sample size per group = 15). An unpaired two-sided t-test is done for each of 1000 random samples under each of the two power conditions. At 50% power about half of the t-tests have a p-value of 0.05 or less, whereas 90% of the tests at higher power are positive.

Now let's look at the averages of the mean group response values only in those tests that correctly detect a "significant" difference between drug and placebo treatment.

In the low power simulation, the average drug effect is 375 response units and that for placebo is 176, for a difference of 199 response units. Thus, in the long run, the average estimate from the low power is inaccurate by 49 response units.

In contrast, the high power simulation yields response unit estimates of 353 and 197 for the drug and placebo groups, respectively. They are accurate to within 6 response units from what we coded in for the groups.

When using p-value based cut off criteria to make decisions on whether an experiment worked or not, we are at risk of making inaccurate estimates of the effect size of a treatment. The high powered tests not only detect more "positive" results, but they are also much more accurate!

Some in the biomedical world have called for increasing the stringency of type 1 error tolerance, setting the "significance" threshold at p< 0.01 rather than p<0.05 alpha. What effect do you think this will this have on effect size estimation of underpowered experiments?



Saturday, January 16, 2016

Scientists as Idols

When I tell people I have a degree in chemistry, they often respond with a variation on the theme of “Wow, you must be so smart, I could never do that.” I also find that people will defer to me in a number of areas, regardless of whether they’re in my field, only on the basis of me being a scientist. Sorry, friend. I know nothing about mold…Yes, I know it’s science, but I study immunology...No, I won’t come look at your bathroom to see what species is growing…Call an exterminator.

This idea of a scientist having above-average intelligence and thus being worthy of more respect is so pervasive in our culture that we tend to put them up on a pedestal. Look, here, a glittering example of erudite logic, someone who we can trust to be correct and awe-inspiring and know which toothpaste is the best toothpaste. While I sometimes enjoy this, I also think it’s a factor in a number of the problems pointed out in the reading for this week.

At the individual level, scientists often feel a huge amount of pressure to live up to the standards imposed by society. We need to study something high-impact, something worth the grant money, mentorship time, and resources, something that our parents can brag about— and something that can get published. As pointed out in the Economist article, people don’t want to pay for replication, no matter how necessary it might be to the core tenants of empiricism, and people don’t want to publish repeats or failures.

This selection bias in publications leads to a false idea that if you’re doing things right, it’ll all work out. If it doesn’t, then your perceived competence, funding, and perhaps graduation date are in jeopardy. Indeed, the only real failures of scientists we see are those of eminent researchers being ripped from their pedestals when they’re caught with false data or their drug trial goes poorly. These incidents exemplify the idolization of scientists when the popular media harps delightedly on the fall of the mighty and the dissolution of the trust the public had in science overall. Against a backdrop of such high standards and fear of failure, scientists themselves perpetuate the need for flawlessness and continuous “miraculous breakthroughs” among themselves. (Recently, a study found that much of the hype surrounding mundane results originates in the press releases from institutions themselves.)

As Jared Hovarth wrote in his Scientific American article:
“In reality, science progresses in subtle degrees, half-truths and chance. An article that is 100 percent valid has never been published. While direct replication may be a myth, there may be information or bits of data that are useful among the noise. It is these bits of data that allow science to evolve. In order for utility to emerge, we must be okay with publishing imperfect and potentially fruitless data. If scientists were to maintain the ideal, the small percentage of useful data would never emerge; we’d all be waiting to achieve perfection before reporting our work.”
We all strive for perfection in our work, to live up to the expectations set by our peers, our mentors, our funders, our publishers. Unfortunately, our inherent imperfections and inability or aversion to recognizing them leads to bad science and bad publications.