Showing posts with label #analysis. Show all posts
Showing posts with label #analysis. Show all posts

Monday, January 22, 2018

Should scientist be synonymous with statistician?

Big data, which refers to a huge number of diverse data created at a high rate, should be leveraged by scientists to gain knowledge previously inaccessible via traditional tools and techniques. The scientific field is rapidly advancing; researchers are performing increasingly difficult experiments and creating exponentially greater data sets. Complex mathematical techniques for crunching big data have been developed, but these statistical techniques are not well understood by the greater scientific community (read more here). In fact, the McKinsey group asserts that “the United States alone faces a shortage of 140,000 to 190,000 people with analytical expertise and 1.5 million managers and analysts with the skills to understand and make decisions based on the analysis of big data" (read the full report here).  This uncertainty forces scientists to either 1) rely on outdated and often inappropriate statistical techniques, or 2) apply new statistical techniques without formal training, increasing the likelihood of errors.  
I find myself at this crossroad, searching for a third path where I can gain the statistical training I need to analyze the complex data I am generating in my laboratory. As a graduate student in the field of neuroscience, I expected my classes to adequately prepare me for everything I would encounter at the bench top. However, I have realized I need additional training in statistics and programming that goes beyond the scope of my required coursework.  This begs the question: how many of my fellow graduate students feel the same as I do? I believe that more rigorous statistical training should be required of all scientists, regardless of whether they are in the beginning stages of their training or approaching their retirement, because the future of science lies in big data. 

Sunday, January 21, 2018

Representation to Interpretation

      As bias effects all stages of scientific research, scientists should actively contemplate common falterings and look to minimize or at least acknowledge their presence. Bias must be considered in experimental design and implementation as to collect accurate data. Such forethought allows for a closer representation of the intended variable: hence, decreasing the scope of confounding variables while increasing reproducibility. Appropriate statistical methods deal with the type of variables measured and provide analysis. One area of bias I did not readily think of was hidden within the presentation of those statistics.  
  For example, I found the traffic data in the snow intriguing. Google maps advertises the feature as a source traffic information. Therefore, we interpret the data as a measure of how many cars are on the roads, presented in ordered categories of increasing traffic being clear, orange or red. The raw data includes location and the output (speed in which iPhones on the given road are moving in relation to the speed limit). The colors then serve as a proxy where the correlation to traffic is an inference from the data. As you mentioned in the snow, the slow movement of iPhones translates to a high traffic reading. In this case, the data is still accurate (because it’s measuring iPhone motion); however, due to the limitation of this model, the outcome variable not longer directly correlated to the interpreted conclusion.  
If not given access to the data, information can be lost within the interpretations and representations of statistics. Although the published work may be accurate, if readers analyze the work further, questions of bias and inherent limitations might arise. I am a proponent of open sharing of all data when possible (as suggested in this article). This along with the proposed more extensive methods section not only gives scientists the ability to critique the data and conclusions, but may also generate conversations pertaining to mistakes, bias, and limitations.

Thursday, January 18, 2018

Preventing Population and Gating Bias


One of the primary tools used to generate scientific data in the field of immunology is flow cytometry. The data generated by this technique is then analyzed using a program called FlowJo, which requires user input to look at specific populations generated by the experiment.  Due to this requirement for user input, there may be some inherent bias generated while looking at specific populations through a method called “gating” which allows you to isolate a specific population and further compare that population with other experimental parameters which then require further gating, leading to the possibility of more bias.

These biases can stem from the thoughts along the lines of, “Oh, this is where I should end this gate to include [x] amount of the population,” or, “Gating around this section of this population will make my data significant enough to be published.” In addition to this gating bias, those analyzing flow my also generate a bias on what they would like to analyze. For example, the scientist may want to investigate the relationship between variable x and y, but decide not to look at x and z even though z is on their panel for the sake of gating. This form of bias may prevent the analysis of data that could be significant and play a key role in what is being investigated. For both types of biases, there are methods that are being and have been developed to prevent this from occurring.

There are types of scripts in R that can generate gating for analysis based on control samples, which would prevent gating bias (unless the controls were biased, but that’s just bad science). There are also programs like CITRUS that allow users to input their data into a system that will use an algorithm to analyze each variable to every other variable and report which variables are significant when compared to each variable. These types of methods can help prevent biased generated in the previously mentioned ways, and as scientists we should continue investigating methods that allow us to analyze our data in ways that prevent us from overlooking valuable information or skewing data in our favor.


Flow gating example for you non-flow users:


https://www.researchgate.net/profile/Barbara_Shacklett/publication/228099465/figure/fig4/AS:195882004815878@1423713317377/Flow-cytometry-gating-pathway-for-T-cell-activation-markers-Initial-gating-was-on.png



Citrus: https://support.cytobank.org/hc/en-us/articles/226940667-Overview-of-CITRUS
Flowjo: https://www.flowjo.com/

Monday, April 25, 2016

Bias with Western Blot quantification on Image J

As a molecular scientist, I perform a lot of western blots. As you may know, western blots are pretty useless unless you can quantify the difference between the lanes. Image J (or FIJI) is frequently utilized to quantify the results from western blots. However, I have always been weary that there are multiple ways to introduce bias into your quantification with this program. I have wondered if the same person were to analyze the same western blot multiple times, would they get the same values each time? To explore this, I decided to utilize an old western blot image and see what different values I could get without trying to be biased.
To start, I utilized the western blot below. Since this was just an exercise of intrigue, i only analyzed the top bands on the last three columns.


To use Image J, you must select the areas that you want to quantify:
I believe that this first step can cause variability. If a researcher were to make the selected area more narrow or more wide, they may pick up noise surrounding it and make their signal seem greater than it is. To examine this further, I selected those bands three different times and produced the following intensity graphs:
"Attempt 1""Attempt 2"
"Attempt 3"

As of now, there does not look like too much of a difference between these results. To quantify your results on Image J, you use a selection tool to calculate the area under the curve. After quantification, I performed a two-way ANOVA with Tukey's multiple comparisons. I wanted to see if 1) the quantification of the same lane altered between "attempts" and 2) if the relationship between the lanes was the same throughout the different "attempts". It is important to note that the following quantifications do not take noise from the western blot into account.
I was quite surprised to find that the three attempts significantly differed from one another. However, the relationships between the lanes was preserved. 

As mentioned, this method of quantifying did not take noise into account. To correct for noise, Image J allows you to draw a line at the level of noise to act as a threshold. The problem with this is that you cannot really standardize it. Typically, a researcher has to "eyeball" the average level of the noise. To complete this section, I attempted to be as unbiased as possible.The plots will then look like this:
"Attempt One" "Attempt Two"
"Attempt Three"

I then took the area under the curve for these noise-corrected plots and performed another two-way ANOVA with Tukey's multiple comparisons. 
With these normalized values, there was NOT a significant difference between the values for each attempt. Interestingly, normalizing the values revealed a difference between Lane 2 and Lane 3 that were not significant prior to normalization. 

Obviously, we cannot overreach these observations to all wester blot analyses. It is safe to assume that to be as accurate and unbiased as possible, one most control for noise in there western blots. However, I was very surprised that the effect betweens the lanes was preserved amongst all three "attempts." I would like more examples demonstrating this before I get too comfortable with western blot analysis. The next step to examine bias with Image J analysis would be to have multiple people analyze the same western blot. Until then, we must all be careful when analyzing western blots and try to be as consistent as possible.