Showing posts with label biostats. Show all posts
Showing posts with label biostats. Show all posts

Monday, January 16, 2017

Wasted Money on Clinical Trials


Training to be a scientist includes the ability to figure out which science articles might not have enough results to back-up its claim. It also includes the ability to identify if there are any results that could have been affected by bias. Being able to figure out if something might not be reliable is important because it might protect you from being steered in the wrong direction.

This might be a frightening fact but it has been reported that 80% of Chinese clinical trials have been fabricated in some way (Fiona Macdonald). These are clinical trials, not scientific discoveries! This means that these are the studies that may eventually result in a drug approval and be used for a particular condition. I am shocked that the values are so incredibly high on irreproducible/fabricated clinical data in China. The Chinese State Food and Drug Administration (SFDA)  reported that results were written before the clinical trials took place, statistical data was designed to look significant, and/or data was completely left out (to make positive results look consistent). Fabrication is a form of bias because it is in the self interest of the researcher to publish and gain recognition. But why do these individuals not get exposed if the fabrication is so obvious?


Reviewers in China seem to be more passive with “Sketchy” science and it is their responsibility to change the vicious cycle of publishing fraudulent or bad science. Jeremy Berg emphasized the importance of catching clear flaws and making sure enough information is included to allow experimental replication. Reviewers and editors of manuscripts should always have the highest expectations for the journals they wish to publish. The everyday person should be concerned about the amount of fraud, malpractice, fabrication. The government uses tax money to fund research and billions are lost each year to fraudulent research. An article from Jocelyn Kaiser stated that $28 billion a year (in the US) is spent on irreproducible biomedical research. People are more likely to follow rules that are strictly enforced and this includes rules about methods described in journal articles.


Monday, April 11, 2016

Because Most Things Change... (Continuous Variables)

Since I was a little girl, I have had one thing clear: most things change. I have grown older, taller, and gained more weight.  This is because none of these characteristics are constant. When thinking of numerical values, we hardly ever see them as a variable amount that describes a measurement. Notice the importance of the word “variable” when used as an adjective, given it is a number that may be assigned more than one value. For example, the last time I weighed myself, the value on the scale was different than the previous time (although I wish it wasn’t). This is because weight is a variable measurement. Among other variables we may have height, age, time, distance, and temperature. However, it is not as simple as deciding if something is constant or variable.

A variable can be either continuous or discrete. As opposed to the discrete variable, the continuous variable can assume an infinite number of real values. So from the variables described above, time, height, distance, age, and temperature would be continuous. Harvey Motulsky explains in his book Intuitive Biostatistics that one way to summarize continuous values is to calculate the mean, median, mode, geometrical mean, or harmonic mean. On the other hand, if it can only take a finite number of real values, it is discrete. Thus, examples of something that cannot be divided, as Dr. Murphy stated in class, could be a person or a pregnancy. The type of variables we have will determine the type of graph and type of statistical analysis we will perform.


In Harvey Motulsky’s Intuitive Biostatistics, he states that when graphing continuous variables, we should consider creating a graph that shows the scatter of the data. He suggests for us to either show every value on a scatter plot or show the distribution of values with a box-and-whiskers plot or a frequency distribution histogram. The fact that he emphasizes on the way we should graph the data is an indicator of the importance of differentiating between these two types of variables. The main importance is to know which statistical analysis you will use for your data. In Michael Cheatam’s A Practical Guide to Biostatistics, he states that the t distribution is frequently used to evaluate hypotheses regarding the means of continuous variables, assuming the data is normally distributed. However, the most commonly used nonparametric methods for non-normally distributed data are the sign test, the Wilcoxon signed-ranks test, and the Mann-Whitney U test.


It is because of the difference in the way the data is presented and the analysis being made to the variables that it is crucial to understand what a continuous variable is. It amazes me how something so simple in definition can be so crucial for correctly exposing your data to others in a way that it actually conveys what the crude data was showing. 

Statistics helps explain challenges in cancer prevention research

Last Friday in Winship Cancer Institute auditorium, there was the speaker Dr. Yang talking about tea in cancer prevention. Below is one of his slides which illustrates that tea might help prevent skin, oral, esophagus, lung cancers etc..
It could be very wrong to say that tea definitely prevents these cancers in "potential" human patients based on what have been taught in biostats class.  Thankfully, Dr. Yang did not make that statement in his talk but instead listed findings or research that has been done across the globe to test ingredients in the tea that might help prevent cancers in lab settings or clinical trials.  Even at the very end of Professor's Yang's talk, he did not jump to the conclusion of cancer prevention effects of tea.  This might sound frustrating to young researchers who has the ambition to prevent cancers someday in the future.  But it is a good talk to me if I combine biostats concepts and what Dr. Yang talked about or what others in this field has been pursuing for many years.  

The first thing I learnt was to keep both scientific hypothesis and statistical hypothesis in mind before you actually perform an experiment or start a research project.  It's easy to get lost when you gradually learn more about your research subject but without a clear science question to answer.  For example, in cancer prevention research of a lab setting, if you want to test if ingredient A has the effects of lessing prostate tumor burden in mouse model, you should stick to it even later in research you found that ingredient A somehow magically decreased the mouse weight and may have effects in preventing obesity.  Then you suddenly changed your hypothesis in order to get your research published as soon as possible.  This is the so-called HARKing (hypothesizing after results are known).  In this way, there is greater possibility that you are biased to jump to the false positive results without careful statistical decisions beforehand.  

Also, it's also critical to keep effect size in mind rather than just rely on p-value when you want to extrapolate a lab finding to the clinical trials.  Nowadays, we've seen so many failures in clinical trials where the drug used has been shown to have "significant" effects in lab research.  Part of the reason is that we consider p value to have magic power to draw the conclusion but ignoring the effect size, especially in the case of cancer prevention.  We've seen that 8 out of 10 mice in our lab setting responded to the drug and showed a delay in tumor formation (treated group showed five days delay of developing the same size of tumor).  Statistical test was run and it had low p value.  However, irrespective of the fact that mice are different from humans, would the clinical data capture the difference which corresponds to difference of five days delay of tumor in mice?  Not quite. Confounders in clinical trials are difficult to control, which make the prevention study even more challenging. 

It's still questionable to persuade individuals to have certain supplement or drink tea even if you've seen the significant benefits in clinical trials.  The study in population may not apply to every individual.  In one research presented by Dr. Yang, the clinical trials showed that the benefits could only been seen in females but not males and they proposed that smoking might be the confounder.  If smoking is really the problem here, then drinking tea might not do any good to a smoking individual. We need to treat the clinical data with caution before we make any conclusions or approve any supplements that's said to prevent cancer.

It needs both statistics and sound judgement to do good science, especially in cancer research.            

Challenges: Paper Compilation and Mad Dollar $igns



When I first looked at my assignment for this exam and found "Challenges in Statistics" next to my name I was a little dismayed. Isn't everything in statistics a challenge in statistics? I asked myself, but I realized this was the section of the text which dealt with gaussian distributions and outliers, probably the two statistical challenges I've actually faced since beginning this course. I have encountered both of recent in looking through a large old data set which I am trying to re-interpret for a paper I am writing. The data was generated by an old tech in the lab who isn't in the state anymore, and one challenge I had in approaching the data statistically and anticipating how I could display my data was that I didn't know about the distribution. I hadn't worked with these cells before, and I had only done a similar, but not exactly the same protocol. Reading about normality tests reminded me that I had that option, and now I remember what the "skewness" and "kurtosis" values are useful for in prism. Similarly, I struggled with outliers and how to treat potential outliers, because normally I would only reject an outlier if I had reason to believe that my own experimental error or some other thing outside my control had biased this outlier. However, since this wasn't my data, I also didn't perform the experiments! I had to make some educated choices with my PI, trying to remain informed with the Grubb's test.

 

In addition to the lab, the issue of distribution also affects my day-to-day money management. I'm a grad student who likes to invest my modest stipend in stocks. As someone who is fairly new to trading securities, one comfortable way into the world of Wall Street was via statistics. The same challenges in statistics which face scientists also face economists and stock traders. When there is real money on the line, it helps to have some of these statistical tools on my side, to help temper any bias. Given that I am trading in real companies with significant reputations, it's hard not to get emotionally involved in any of my holdings. However, if I can set some rules and use some of the tools in this section of the text, I can increase my overall profit through bullish and bearish. For example, check out this Forbes article which underscores how a normal distribution of the DJIA 30 indicates a healthy market (also see figure).

Displaying an individual stock's performance in what we hope is a normal distribution or finding standard deviation can indicate the volatility of a stock, correlated to risk. This value can also be similar to beta, the value many investors use to indicate volatility/risk, although it should be mentioned that this is not how beta is calculated. Lastly, many trading algorithms on Wall Street and financial theories anticipate a gaussian distribution of stock prices in order to maximize alpha or gains adjusted to the overall performance of the market. Thus, skew and kurtosis are important for investors to keep track of! The role of normal distribution, skew, and kurtosis in investing is simply summed up in this post, for those interested.