Showing posts with label hypothesis. Show all posts
Showing posts with label hypothesis. Show all posts

Monday, April 11, 2016

Statistics helps explain challenges in cancer prevention research

Last Friday in Winship Cancer Institute auditorium, there was the speaker Dr. Yang talking about tea in cancer prevention. Below is one of his slides which illustrates that tea might help prevent skin, oral, esophagus, lung cancers etc..
It could be very wrong to say that tea definitely prevents these cancers in "potential" human patients based on what have been taught in biostats class.  Thankfully, Dr. Yang did not make that statement in his talk but instead listed findings or research that has been done across the globe to test ingredients in the tea that might help prevent cancers in lab settings or clinical trials.  Even at the very end of Professor's Yang's talk, he did not jump to the conclusion of cancer prevention effects of tea.  This might sound frustrating to young researchers who has the ambition to prevent cancers someday in the future.  But it is a good talk to me if I combine biostats concepts and what Dr. Yang talked about or what others in this field has been pursuing for many years.  

The first thing I learnt was to keep both scientific hypothesis and statistical hypothesis in mind before you actually perform an experiment or start a research project.  It's easy to get lost when you gradually learn more about your research subject but without a clear science question to answer.  For example, in cancer prevention research of a lab setting, if you want to test if ingredient A has the effects of lessing prostate tumor burden in mouse model, you should stick to it even later in research you found that ingredient A somehow magically decreased the mouse weight and may have effects in preventing obesity.  Then you suddenly changed your hypothesis in order to get your research published as soon as possible.  This is the so-called HARKing (hypothesizing after results are known).  In this way, there is greater possibility that you are biased to jump to the false positive results without careful statistical decisions beforehand.  

Also, it's also critical to keep effect size in mind rather than just rely on p-value when you want to extrapolate a lab finding to the clinical trials.  Nowadays, we've seen so many failures in clinical trials where the drug used has been shown to have "significant" effects in lab research.  Part of the reason is that we consider p value to have magic power to draw the conclusion but ignoring the effect size, especially in the case of cancer prevention.  We've seen that 8 out of 10 mice in our lab setting responded to the drug and showed a delay in tumor formation (treated group showed five days delay of developing the same size of tumor).  Statistical test was run and it had low p value.  However, irrespective of the fact that mice are different from humans, would the clinical data capture the difference which corresponds to difference of five days delay of tumor in mice?  Not quite. Confounders in clinical trials are difficult to control, which make the prevention study even more challenging. 

It's still questionable to persuade individuals to have certain supplement or drink tea even if you've seen the significant benefits in clinical trials.  The study in population may not apply to every individual.  In one research presented by Dr. Yang, the clinical trials showed that the benefits could only been seen in females but not males and they proposed that smoking might be the confounder.  If smoking is really the problem here, then drinking tea might not do any good to a smoking individual. We need to treat the clinical data with caution before we make any conclusions or approve any supplements that's said to prevent cancer.

It needs both statistics and sound judgement to do good science, especially in cancer research.            

Sunday, April 3, 2016

Chicken or Egg?



Every discipline seems to have its own version of the chicken or egg debate. For statistics, the debate could be: Which came first, the model or the data?

The answer would initially appear to be quite obvious. The data must come first, of course. As Harvey Motulsky states in Intuitive Biostatistics, “Regression does not fit data to a model…rather, the model is fit to the data.” In other words, the data are used to calculate the parameters of the model. That basic premise is quite clear.

Where it gets tricky, however, is in the selection of which type of model to use. There are linear, logarithmic, dose-response curves, binding curves, higher order polynomials, etc. To navigate these many options, Motulsky again has sage advice: “Choosing a model should be a scientific decision based on chemistry (or physiology, or genetics, etc.).” So if your data are measurements of radioactive decay, you know that you should use an exponential decay model. Or if your data are measurements of the effect of a drug, you know that you should use a logarithmic dose-response model. Even if the R2 value for a different model were higher, it would be inappropriate to try to fit your data to a model that you know does not make sense biologically, chemically, etc.

But that is where the chicken or the egg question comes in to play. How does the initial model for a particular biological or chemical system get established? Surely someone, at some point, had to try several models and find the one that fit best for that type of data. If no one has ever established a model for a particular system, you cannot follow the statistical best practice of selecting a priori that you will fit a particular type of model to your data (as in (A) in the figure below). Instead you have to try several types of models. Rather than using the data to determine the best-fit values for parameters of a model, you are now using the data to determine which parameters even need to be calculated in the first place. Certainly this process will still be driven by the data, and rigorous statistical tests can be applied to determine which model fits the data better. Nevertheless, as the figure below demonstrates, this situation (B) fundamentally alters the workflow needed to test your hypothesis.  Once the appropriate model for a system is established, the workflow can return to that outlined in (A), but the initial need to establish a model flips the data-model relationship somewhat.




So although the data do always inform the model (and a model is always fit to the data, not vice versa!), there are situations where the data come first and situations where the model comes first. As with most chicken or egg debates, perhaps this one does not have a definitive answer either.  

Tuesday, February 2, 2016

Is bias a necessary evil?

Biases are ubiquitous. And because they are ubiquitous, we must embrace them for we cannot escape. This is especially true to science. I see biases as vital and necessary components of science. I often see bias portrayed as being bad--rightfully so as bias can really screw us over as portrayed by the articles I read. But from reading these articles, I wanted to find reasons for how bias can be good. Perhaps I am misunderstanding the connotation of the word "bias" or applying the word subjectively? No matter what, here are some reasons why I think bias could be a good thing:

  1. Biases, aka hypotheses, are the impetuses for projects. Scientist must have a central belief, which are biased by our expertise, past life experiences, our colleagues, mentors, and etc, to which we frame our scientific questions. I believe that these biases provides the momentum for the creation of projects and propel discovery. Without our constantly changing biases science would not be in perpetual motion. 
  2. Biases allow us to be more critical, allowing for the advancement of science. We are taught as scientist to always question what we see, what we read, and what we hear. We would not have such critical minds if we did not have a bias that something published is not always true or causative. I think that such skepticism pushes science forward. 
  3. Biases force us to do better science. The point of publishing is share your discoveries with supportive evidences that try to minimize biases. Because we have these biases and want our results to be as objective as possible, we design "controlled" experiments. Thus, bias forces us perform scientifically valid experiments and analyze data that can best confirm our hypotheses.

Here are my thoughts on how bias is a necessary evil in science. Without them it may be hard to pose a scientific question, make it impossible to be more critical, and most importantly, perform and analyze truly honest and objective experiments. Since there are many examples of how bias can be detrimental to science, I just wanted to be a devil's advocate and provide some reflections about how bias can actually be good for science.