Showing posts with label randomization. Show all posts
Showing posts with label randomization. Show all posts

Tuesday, May 17, 2016

Replication, Pseudoreplication and Cell Culture


This is not an easy subject. 

Here's an email thread from the other day between Rick Kahn, TJ Murphy and Nancy Bliwise. Since both TJ and Nancy direct courses at Emory for PhD students and undergraduates, respectively, Rick consulted each with his question.

I've injected a final thought at the end.

_________________

Rick: An issue arose at lab meeting yesterday when I suggested to a student that they include a second set of samples, same as the first, to provide a biological replicate. This is a control experiment demanded by a reviewer that is very unlikely to change the point of the paper, but of course that is not really the point here. You were quoted in our discussion as saying that it is not a biological replicate unless performed on a different day - which is pretty much in agreement with this article, which I like a lot but disagree on this point:

http://labstats.net/articles/cell_culture_n.html

Given that the cells are all "the same" or come from one vial or even separate vials all frozen down together, I see no value in adding extra days and it is impractical to obtain multiple aliquots of cell lines from different vendors for example. Certainly if doing rats or mice or human patients you gather from different animals. But the assumption in cell line experiments is that they are clonal and therefore supposed to be identical.

I think it is a discussion worth having and I am thinking of soliciting views across the GDBBS, for those who use cell culture in experiments. Thoughts?


TJ: A sample is comprised of replicates.  Statistically, here are the two unbreakable rules on replicates within a sample:

1. Each replicate should be independent from every other replicate.
2. They should be randomized (in some way, shape, or form).

The problem with continuous cultured cells is that they are immortalized clones. Another problem is that it is hard to randomize culture wells to treatments given the need to be efficient with pipetting, etc.

Yes, you can argue that cultured cells are so similar from day to day that it seems absurd to consider a replicate taken on Monday as independent from one taken on Tuesday.  I see that point.

It is harder to argue there is anything different at all between two replicates that are taken on the same day, side-by-side.

The day-to-day rule introduces a modicum of random chance into the design, given a system where the homogeneity is so heavily stacked against random chance.

Monday and Tuesday are more independent form each other more than are the left side and right sides of the bench on Monday.


Rick: I think we agree in principles here. As the article says it is incrementally better. But I would argue impractical or not worth the increase in terms of time and resources. I suspect if put to a vote of others my way is the overwhelmingly popular solution, though admittedly incrementally less independent.

TJ: I always made cell culture passage number the mark of independence. In my own system, a continuous primary cell culture line, there was good variation from passage to passage, much more so than within passage.

I think one of the big reasons we're in the unreliable research mess we're in is because when its been a choice between sound, unbiased statistical practice vs efficiency/cost, the latter wins too often.  You may not like to hear that, it is surely an unpopular viewpoint for the go-go-go PI set, but I don't have any doubt that is what underlies the problem.

"I'm not cutting corners, I'm saving money."

The thumbs are on the scales in all kinds of ways in the name of efficiency.

It's like the culture of conducting unbiased scientific research has been replaced by a manufacturing culture. We should have a few beers just to hammer out the latter point.


Nancy: Rick, Do you know if the reviewer was requesting a biological replicate (e.g., lines from different animals) or a technical replicate to show that the procedure/phenomenon can be reproduced by others?  While the article you provided suggests that replicating over different days or weeks, depending on what is appropriate for the experiment provides a purer technical replicate, I am not particularly convinced by that argument.  I would argue that any replicate depends on the conditions/issues that you think need to be demonstrated.  For example, if the experimental manipulation is highly technical and requires great skill, I would think a replication by different technicians/experimental staff would be important.  If lab conditions are important, then different days/weeks might be meaningful.  If it requires particular equipment calibrated carefully, then having another lab replicate it might be important.  If the concern is just that this might be a fluke of a particular set of circumstances, then a single replication that shows the same phenomenon would be ok.

TJ was addressing independence and "n".  It is possible to do a replicate on the same day/week and model the analysis to show that there were multiple cultures created nested within the original line.  Doing the experiment on different days/week produces some independence but only if the lab conditions across days/weeks are relevant to the question.

TJ: I'd argue that statistical independence is only critical when doing significance testing. Not every experimental observation needs a significance test. For example, a central observation related to the main scientific argument should be designed for significance testing. Observations that are ancillary, adding only texture to the central argument, should be repeated to ensure reliability. 
___________________

Let's distill this down to something of a more general take away.

We want to know either or all of three things about our data: Is the observation accurate? Are the findings precise? How reliable are the results?

It's up to the researcher to decide which of these is the most important objective of an experiment and then to go about collecting their information accordingly. 

When accuracy is a key finding, replication has the effect of improving the estimate since standard error reduces with more replicates. 

We standardize our techniques and analyze variance when precision is important. 

When reliability of the observation is important, then the biggest threat are biases, both expected and unexpected. The device invented to deal with that problem is randomization.



Thursday, April 21, 2016

A Basic Guide to Randomization in Excel

Throughout this semester, we've discussed many many many times how important randomization is to proper experimental design. In order to infer anything about a population, your experimental sample must randomly choose samples from the population. Any bias in choosing some samples over others will invalidate one of the fundamental assumptions of proper statistical analysis. So, on a more practical note, I wanted to provide a (hopefully) quick and easy guide to randomizing a list of numbers in Excel.

For this tutorial, I made up an imaginary data set. Let's assume I'm examining individual cells in live cell microscopy and timing how long they stay in prophase (one stage of the cell cycle). In my preliminary data, I counted all 40 cells imaged in this group, but I only want the data from 20 randomly selected cells. Below is the table of all 40 cells and their time in prophase.

Next, make a new column to the right of the data points titled "RANDOM". In the first cell of the column, add the formula "=RAND()". Excel will produce a random number between 0 and 1 in this cell.


Apply this same formula to every cell in the column corresponding to a data point by double clicking or dragging down the highlighted bottom right corner of the first random cell (or just copy/paste).


Importantly, you must convert all of the random formulas to permanent numbers that won't keep changing every time you refresh the page or perform another function. So, copy all cells in the Random column, and right click on those same cells to choose Paste Special --> Values. This will replace the formula with a permanent number of the same value. This is really really important to do or the next steps won't work.


Next, go to the Data menu in the top bar and select the Sort button (see button in upper right corner in image below).


A window will appear. Choose to sort the RANDOM column by values, smallest to largest.




Now, your list will be sorted by smallest to largest random value, meaning the data points are now in a completely random order. The entire row is maintained when you sort in this way, meaning that the same ID number, data point, and random number are sorted together as a group. Now you can select the first n number of cells that you need for random selection. Since I wanted 20 random data points in this scenario, I highlighted the first 20 entries that appear in the list.



And voila! You have a randomly selected sample from your list! Have any of you used Excel for randomizing data before? Do you prefer a different program? Any other tips or tricks for people learning to randomize data? Comment below!

Sunday, April 10, 2016

Political Polls: A Quagmire of Sampling Errors

It is once again a presidential election year, which means mainstream media is practically covering our daily lives in a bloodbath of tortured poll results. Everywhere you look, there is yet another new opinion poll showing how the political pundits are faring against each other. However, before even looking at any of these results, we should be questioning how the survey was designed before it even began and how the data was actually collected.

The end goal of these political opinion surveys is to infer how the whole population of voters in the US will vote in November. However, sampling the opinion of every single American voter is completely impossible given the constraints of time, money, and willingness/truthfulness in every voter answering. So instead, polls must base their opinions on a smaller sample, a subset of individuals taken from the greater population. If the sample is chosen by random sampling, then the distribution of opinions of those people will be similar to that of the entire population. Randomization is key here. If there are any parts of the population that are intentionally (or unintentionally) excluded or over-represented in polling, then the sample distribution of opinions will not necessarily resemble that of the whole population. Problems in proper random sampling have been hotly debated for decades as discussed by Jill Lepore in a recent article in The New Yorker, and the debates have only grown more contentious in recent years with changes in technology. For example, many opinion polls are conducted by random dialing of phone numbers and collecting responses of whoever picks up. However, telephone response rates are currently in the single digits, and people who willingly pick up the phone often have very different political behaviors than the general public. Furthermore, random calling to cell phones was banned in the 1991 Telephone Consumer Protection Act, so calls are only going to land lines and not cell phones. And once again, the demographics of people who have land lines is not fully representative of the whole population, excluding younger voters in particular. After the UK elections in 2015, Nate Silver, the widely beloved political analyst behind FiveThirtyEight, highlighted a few recent examples of inaccurate political polls and some of the reasons to worry about bad sampling, such as the inability to rely on contact by phone, unreliable online polling methods, Americans withholding their true opinion in polls, and herd opinions swaying voters.

Although political polls aspire to have completely representative polling, an accurate sampling is nearly impossible in our current culture given the reasons above. To compensate for this, many political polls use methods that could be considered scientific heresy to scientists working in hypothesis-driven experiments. For non-probability sampling methods, pollsters make the assumption that each sample selected does not have an equal chance of being chosen from the population (i.e. younger voters with cellphones but no land lines are far less likely to be chosen in phone-based polling methods). There are various methods to compensate for the changes in variability, such as weighting the results of certain samples more than others. For example, a phone-based poll could weight the opinions of younger voters or minority voters more so than the opinions of older white voters more likely to have phone lines. But remember, although this seems like a practical method to “fix” political polls, it makes for really really bad science. Hypothesis-driven scientific studies only work on the assumption that all samples were RANDOMLY chosen from the population, i.e. the probability of choosing any one sample is equal for all samples in the population. Using non-probability sampling or not using true randomized selection throws this assumption out the window, runs it over with a semi-truck, and lights the remains on fire. Never trust scientific experiments using non-probability methods, they are scientifically unsound and completely useless when it comes to predicting anything of value about the whole population. And for political polls, make sure to investigate the weighting methods used and determine yourself how much you actually trust the reported results.

Another important feature of political polls is their sampling error. By pure chance, any sample poll of a population will have a slightly larger or smaller proportion of subjects voting a certain way compared to the true population vote. Pollsters only know the opinions of the small sample, not of the whole population, so the best way to show how sure you are that your sample represents the whole population is to calculate confidence intervals. A poll result percentage by itself is completely useless for making assumptions about the population without a confidence interval. By pure tradition, almost everyone uses a confidence interval of 95%. This means that the range of values reported has a 95% chance of containing the actual population value and a 5% chance that it does not contain the population value. For example, let’s pretend you polled 100 people about if they would vote for Bernie Sanders or Hillary Clinton in a state’s Democratic primary. If 54 of 100 voters in the sample say they would vote for Bernie, then the 95% CI for the sample is +/-9.77%, which means that the percentage range of 44.23% and 63.77% has a 95% chance of containing the population value of the proportion of voters who will vote for Bernie.


Increasing your sample size decreases your confidence interval, as illustrated below by PewResearch.This is because the confidence interval is roughly proportional to the square root of 1/n where n is the number of samples. So, bigger n leads to smaller confidence intervals. 

If you repeat the poll above for Hillary vs Bernie in 1000 sampled voters, and find 54% of voters +/-3.09% support Bernie to 95% confidence, then the percentage range of 50.91% to 57.09% is 95% likely to contain the population percentage of voters supporting Bernie. It’s still the same sample percentage, but the 95% confidence interval is much smaller.


If you want to take a closer look at current political polls, I highly encourage you to look at the compiled list of national polls on FiveThirtyEight and search through the methods sections to see how each poll plays with non-probability sampling weighting, random sampling methods, and determination of confidence intervals. The lack of standards between polls is both horrifying and fascinating and makes me question the results of all political polls, especially when methods aren’t clearly reported.

I’ll leave you with this one interesting snippet of the methods section of an NBC/SurveyMonkeypoll that actually acknowledges inherent bias in their selection methods and how they compensated for that in their methods.  It’s rare that a poll will actually acknowledge their bias like this, so enjoy this unusually honest example.














Saturday, January 30, 2016

Quick thought on randomization and normalization