Sunday, August 26, 2007

Optimization, RNAi, and Lessons Learned from the Summer

I had met Milan Chheda at an outreach event a few weeks earlier, and my final day at the Broad started with a tour of his workspace. Milan, a neurologist, works in William Hahn's Lab, which focuses on how human cells transform into cancer cells. Milan's work is to optimize a technique that is currently applied to study the development of glioblastomas, a particularly virulent type of brain tumor.

The technique in question uses RNA interference to knock down the expression of certain genes. Using a lentivirus, the researcher introduces a hairpin structure to candidate nerve cells to reduce the expression of a specific gene in vitro. After culturing the nerve cells that have received this structure, the researcher typically uses a microscope to check for glioblastomas or other types of cells that may form.

What does it mean to optimize this technique? Milan is part of an effort to create methods that would enable RNAi experiments to scale up the way the Broad's sequencing center has industrialized gene sequencing. In one example, members of his group are working to use software that automatically checks microscope images for glioblastomas and other features. By identifying trouble spots in the experimental pipeline, the hope is that these experiments can be conducted at a larger scale, resulting in reliable and larger quantities of data to analyze.

I left the Broad with a greater appreciation of the interplay between data gathering and data analysis. Writing about my discussions with other researchers greatly enhanced my ability to gain this appreciation. With this in mind, expect further updates to this blog once I return to Berkeley.

Wednesday, August 8, 2007

Science on Wednesday

During junior high, I attended a weekly lecture series at the Princeton Plasma Physics Laboratory designed to introduce students to science research. The program was called Science on Saturday, and it featured scientists in fields as diverse as cosmology and forensics talking about their work to a largely non-technical audience that consisted of students and their parents.

The Broad Institute had a similar program this summer called Midsummer Nights' Science, an apt name given this year's Shakespeare on the Common production. For the four Wednesdays following Independence Day, scientists at the Broad would describe their work to the greater Boston community. Each Wednesday featured a different researcher describing his or her work. While the projects and interests described each week were quite different, all of them implicitly promoted the idea that large databases of data can enable new kinds of research.

The first talk featured David Reich, who discussed how he and his colleague conducted a comparative analysis of the DNA of humans, chimpanzees, and gorillas, which have led them to a new model for the evolution of these species from a common ancestor. The way I first learned about evolution was that it starts when two groups of the same species are physically isolated from one another. Then, under appropriate environmental conditions, the two groups would eventually evolve into different species, after which any hybrids between these two species would be less fertile and die out. This is called allopatric speciation. If this is true, then one can model the DNA sequences as following a branching process, so the evolution of species would look like a tree, where each fork in the tree indicates one species dividing into two. Reich and his colleagues discovered that this model is not a great one to describe the evolution of humans, chimps, and gorillas. In fact, if one constructs a phylogenetic tree for these three species using DNA from one section of the genome, one tree emerges indicating that the most recent split was between humans and chimps split, but if the same analysis is performed using a sequence from another section, which comprises between a fifth and a third of the genome, a different tree emerges, indicating the most recent split was between humans and gorillas. An alternate hypothesis the group proposed was that hybridization among these species took place, and by a careful analysis of the sequence data they had, they were able to confirm this was a better model by which to describe the speciation of these three species. Indeed, the study probably would not have been possible without all the DNA sequence information available for these three species.

During the following week, Pardis Sabeti explained how the HapMap project, another data gathering effort, is enabling researchers to determine the role natural selection has played in humans and pathogens. The HapMap project collects DNA samples from different populations around the world. The samples of DNA they collect account for 90% of the genetic variation among humans. These samples are divided into haplotypes, which represent sections of DNA that are inherited as a group. If there is no selective pressure on an organism, one would expect the prevalence of a particular haplotype to decay as its size gets larger. By similar reasoning, if a larger haplotype is highly prevalent in a population, then there is evidence that the corresponding section of DNA is under selective pressure. Sabeti explained how this has allowed researchers to track lactose tolerance in European populations, who domesticated cattle relatively early, and link the sickle cell trait to malaria resistance. Once again, the availability of this data enabled such an analysis.

The third talk, given by Todd Golub, was about cancer research in the era of genomics. He started the talk by describing two patients, both the same age, both diagnosed with the same type of leukemia in a similar stage of progression, and both given similar doses chemotherapy. However, Patient A lived and Patient B did not. Golub then explained how the mutations that had occurred at the genomic level for these patients were actually quite different, and if one were to look at patient survival by isolating these two different mutations, the group with the same mutation as Patient A had a survival rate much closer to 1 and those treated with the same mutation as Patient B had a survival rate close to 0 within a few years of diagnosis. Golub then went on to describe a treatment that had been customized to target the mutation in groups with Patient B's mutation. The result led to Gleevec, a drug now available to patients with this version of the disease. Since it's introduction, patients diagnosed with this specific mutation have had a 100% survival rate with minimal side effects from the medication. Golub appeared optimistic that similar treatments could be developed for most of these mutations that results in cancer.

Unlike the the preceeding talks, the final speaker barely mentioned genomics in his talk. Vamsi Mootha described mitochondria and his group's research efforts on understanding them. Mitochondria are found inside the cell and produce much of the energy a cell uses. Unlike most other organelles, mitochondria contain their own DNA. However, proteins found in mitochondria are mix of genes derived from within the mitochondrial DNA and the cell's nuclear DNA. It turns out that metabolic diseases are closely related to problems with mitochondrial function, which are apparent in changes to their protein composition. Mootha's group is building an atlas of the protein content of mitochondria in different parts of the body. Their hope is that this data will enable researchers to characterize the specific problems associated with certain metabolic diseases. In this instance, the hope of a future payoff inspired his group to procure a large data set.

Midsummer Nights' Science showcased how biological and medical research have benefited and can continue to do so when certain kinds of data are available in significant quantities. Hopefully this message reached the students who attended and will capture the imagination of those who decide to pursue research in the future.

Wednesday, August 1, 2007

Cultural Learnings

When I told a friend I would be interning this summer, he was surprised.

"Why are you doing an internship?" he asked.

"The idea," I responded, "is to get introduced to a new environment, so I return to grad school with a broader perspective."

"Sounds like Borat."

Like a foreign correspondent reporting to his home country, I gave an informal talk to the Stochastic Systems Group about my summer project. The resulting feedback helped me improve my results this summer. However, once the problem was described, there were a lot of similarities with problems familiar to the group. It was hardly Borat.

That said, there are practices at the Broad outside of my work that I would be surprised to see in my own research community. Perhaps the most surprising thing I have discovered is that people are willing to share their ongoing research with people at the Broad. Weekly seminars feature researchers from outside the Broad discussing their as yet unpublished work. Broadies see data that has yet to be made public. I was particularly surprised by this since there is some controversy that Watson and Crick's paper about the structure of DNA used unpublished data from Rosalind Franklin.

There is a catch. Attendees of the seminar must agree not to work on anything they pick up during the course of the presentation. This understanding and the honor system are what make people comfortable enough to discuss work they might otherwise keep private.

The presentations may also be a way to start collaborations. In a field driven by data, if someone provides the data for a figure on a paper, that person frequently becomes an author, even if the idea of the paper came from others. Thus, advertising results before they are published might allow other researchers to avoiding running the same experiments.

A consequence of this practice is that one rarely finds single authored papers and often finds papers with four or more authors. How does one delineate the contributions of each author? Author ordering may only give a coarse indication of an individual's contribution. An existing solution in some journals is to include an author contributions section. This section typically follows the acknowledgments and may read some like the following:
S.B.C. conceived and designed the experiments. B.S. conducted the experiments. S.B.C. and B.S. performed the analysis. S.B.C. and B.S. wrote the manuscript.
What happens if the work is primarily by two authors? The practice described to me for these instances is called co-first authorship. To do this, one simply places an asterisk next to each author's name with a footnote that reads: "These authors contributed equally to the work."

While some biologists I spoke to joked about some of these practices (one described how an author contributions section might read if each individual's contribution were described honestly), almost all of them were comfortable with the idea that providing data is a legitimate way to become an author on a paper. The same might not be true for my community, but I wonder if any of these practices would transfer well.

Friday, July 27, 2007

Genome Factory

About ten years ago, I spent a summer with other high school students for a summer program at the Waksman Institute of Microbiology. The program's goal was to introduce us to protocols to extract plant DNA and isolate regions of interest for sequencing. We learned how to use restriction enzymes to cut the DNA into smaller fragments, bacterial transformations to make copies of the DNA within E. coli, PCR to make copies of DNA without the help of E. coli, and gel electrophoreses to separate the DNA fragments by size and isolate the one(s) we wanted. Finally, the DNA had to be sequenced, and for this, we were introduced to the Sanger method, developed in 1975 by Frederick Sanger and his colleagues.

The Sanger method involves adding modified nucleotides called dideoxynucleotides, which can only form bonds at one end. Think of a Lego piece with a flat top. Thus, a DNA chain that has such a nucleotide will immediately terminate. If these nucleotides are mixed in with regular nucleotides during a process like PCR, it creates fragments of the DNA sequence with the same starting point and varying endpoints. If only a particular type of dideoxynucleotide such as dideoxyadenine (ddATP) is used, then all the resulting fragments terminate with an 'A'. If these fragments are then separated by gel electrophoresis, one can get a rough idea of the positions where 'A' shows up in the DNA sequence of interest. If 'C', 'G', and 'T' wells are adjacent to the one for 'A', one can just read off the DNA sequence from the gel electrophoresis. This is the basic principle of the Sanger method.

By the time school started again, we had become familiar with the techniques and protocols. We continued to return to the Waksman Institute periodically and apply these techniques. We would eventually use the sequence data from these visits to construct a phylogenetic tree of the Allium (i.e. onion) genus. Unfortunately, the data collection process could often be slow and annoying. There were many stages in which something could go wrong, and I would have to return to the beginning. All of this work produced just a tiny fraction of sequence information from these genomes.

A lot can happen in ten years. Thanks to my friends in the Broad's Outreach Program, I had a chance to visit 320 Charles St., the location of the Broad Institute's DNA sequencing facility. It is sometimes called a high-throughput production facility because of the rate at which they manage to sequence DNA. The facility was responsible for many of the sequences that were part of the Human Genome Project, and I was about to find out how they did it.

We entered 320 Charles St. and sat down for a presentation. Before we could start our tour of the facility, one of the scientists wanted to describe the process. To my surprise, she described the Sanger method. How could this be the process of a high-throughput production facility? Once the tour started, it became clear how: they industrialized the process. We had entered a factory, complete with conveyor belts, robotic arms, and computers. A group of technicians oversaw that the work on this genome assembly line went smoothly. Others, including the scientist leading the tour, were working on ways to industrialize new and improved sequencing methods developed by Solexa and 454.

It was interesting to learn that part of the rate increase has come from engineering solutions to scale up production. The amount of sequence data now available is enabling some researchers to ask questions that may previously have been too time-consuming to answer. I have talked to biologists this summer that have told me how challenging data collection can be, and I am starting to realize how those difficulties play a role in the questions they ask. How might these questions change if other protocols for data gathering were similarly industrialized?

Wednesday, July 25, 2007

Sergio Servetto

It was a few weeks into the start of the semester, and my schedule was set. As I went to the lab printer to pick up a problem set, I accidentally picked up one for ECE 445. Some problems required tools from signals and systems or probability to answer questions about quantization. The final problem was to design a primitive image compression algorithm. Although I would have to switch my schedule, I wanted to take the class. An e-mail to the professor was met with an enthusiastic response, so I made the switch.

Sergio's class was one of my favorites at Cornell. The course mixed theory and programming and made me appreciate the important role theoretical questions have in the design of practical systems. Sergio's teaching style was also one that encouraged questions. He would often pause before answering as if the question being asked were important. Even if I later realized I had said something incorrect or the answer to my question was self-evident, Servetto never sounded dismissive when he answered.

Part of the reason Sergio was able to relate with students was how comfortable he was around them. The first time I walked into his office was just after someone had brought him a freshly baked chocolate chip cookie. Without a second thought, he immediately split the cookie and handed me half. I still remember seeing one of the melted chips stretch between the two halves of the cookie and thinking what a generous thing to do.

Given my experiences in Sergio's class and others, I wanted to pursue information theory and communications after starting graduate school. Sergio and I would periodically meet at conferences. As we were catching up during ITA 2006, he mentioned that he had looked over my Master's thesis. It was great a feeling to know that one of my former professors was still interested in my progress.

Most recently, I saw Sergio at ISIT 2007. He had recently agreed to oversee the Information Theory Society Student Committee, and we talked a bit at one of their events. I last saw him among the audience at my talk.

I found out about the plane crash this evening. It is much easier to reminisce about the past than to describe how I am currently feeling. I've been fortunate to have professors like Sergio Servetto who have encouraged my interests.

I remember we tested our primitive image compressors from that first problem set on a photo of one of Sergio's sons. My thoughts are with the family.

Saturday, July 21, 2007

Variations on a Theme

Information theory, statistical decision theory, and game theory have developed methods to analyze what some may consider adversarial situations. Lessons in these fields have certainly influenced how I model problems involving adversaries. Perhaps it should come as no surprise then that such models were in my thoughts as I attempted to read about host-pathogen interactions. The work in question was a review paper from Hidde Ploegh's lab. Ploegh's lab studies mechanisms by which seemingly simple bacteria have been able to infiltrate our complex immune systems.

How does the lab conduct this research? Renuka Sastry, a researcher at the Whitehead Institute and one of Ploegh's graduate students, gave me a tour of the lab. Our first stop was at what appeared to be a cylindrical dark room.

"It's used for western blots," Renuka said. As she explained, a western blot is a technique to test for a specific protein in a tissue sample. The results are represented as lines on a plastic page, where a line indicates the presence of said protein.

Surprisingly, it takes some effort to extract this bit of information. Part of the process was unfolding on Renuka's workbench. A gel electrophoresis was running, but there were a couple differences from ones I had seen for DNA. First, the gel was positioned vertically instead of horizontally. Second, the gel looked significantly thinner than an agarose gel. Renuka's electrophoresis was one stage in an experiment to test for a particular protein. She was hoping the result would validate an observation she had made earlier.

Noticeably absent from the workbench was a computer. While there was a computer in that room, our next destination was filled with them. The mass spectrometry room is used to identify proteins, and computers are used to crunch numbers and consult databases for protein matches. Of course, the proteins come from living cultures, and in the final part of our tour, I saw one under a microscope.

What struck me during the visit was that these methods and techniques could also be applied to problems that did not involve hosts and pathogens. What drew Renuka to the work?

"I wanted to do biochemistry research," she responded. I left with a better appreciation for this research. Whether this appreciation might inspire new ways to model aspects of these host-pathogen interactions is an open question.

Wednesday, July 18, 2007

Cat's Cradle in a Hard-Boiled Wonderland and the End of the Brave New World

During the ITA workshop in January, Desmond Lun and I had the following exchange.
Me: So what are you doing these days? Are you a post doc?
Desmond: Actually, I'm at the Broad Institute.
Me: What's the Broad Institute?
By the time I started looking for internships, I knew what the Broad Institute was and sent Desmond my resume. I have been at the Broad now for two months, and Desmond and I work in the same group. It helps working with someone here who hails from the same research community, and our conversations span topics that include information theory and biology research.

I had a taste of the future of biology research during lunch when Desmond described his project with George Church's lab. The goal of the project is to study ways to use biology to produce renewable fuel sources. One fuel source is ethanol, and there is a well-known biological recipe to produce it. Add yeast to a sugar solution. Mix. Let it ferment.

The approach Desmond described was a little different. It turns out one can modify the E. coli genome and use the modified E. coli to produce ethanol. Driven by this success, there is an effort to see if alkanes or other fuels can be created by hacking the genome. Indeed, some start-ups are trying to capitalize on this idea.

The technology that enables such genome hacking falls into the field of synthetic biology. What is synthetic biology? The answer can vary depending on who answers, but to my understanding, synthetic biology is the study of how to design and fabricate living systems that do not exist in nature. In addition to adding and removing genes from a genome, Desmond said there exist techniques that allow one to increase the mutation rate of certain organisms. Once enough mutations accrue over the population, a researcher can then create conditions that select the mutations most suited to a task of interest. This may be the only truly parallel implementation of a genetic algorithm.

Of course, such technology also generates concern. The ETC Group is a public watch-dog for synthetic biology. They have been vocal in challenging Craig Venter's attempt to patent synthetic life and oppose the idea of scientists creating synthetic life without regulations. "Playing God in the Galapagos," the title of one of their publications, reflects this position.

These concerns are also in the public consciousness. Desmond mentioned a recent online poll asking about such technologies. The response choices ranged from complete opposition to regulations to complete opposition to the research. How do scientists feel? It turns out Church's lab took a similar poll. Surprisingly, the group was in favor of more regulations.