Showing posts with label genetics. Show all posts
Showing posts with label genetics. Show all posts

Wednesday, August 1, 2007

Cultural Learnings

When I told a friend I would be interning this summer, he was surprised.

"Why are you doing an internship?" he asked.

"The idea," I responded, "is to get introduced to a new environment, so I return to grad school with a broader perspective."

"Sounds like Borat."

Like a foreign correspondent reporting to his home country, I gave an informal talk to the Stochastic Systems Group about my summer project. The resulting feedback helped me improve my results this summer. However, once the problem was described, there were a lot of similarities with problems familiar to the group. It was hardly Borat.

That said, there are practices at the Broad outside of my work that I would be surprised to see in my own research community. Perhaps the most surprising thing I have discovered is that people are willing to share their ongoing research with people at the Broad. Weekly seminars feature researchers from outside the Broad discussing their as yet unpublished work. Broadies see data that has yet to be made public. I was particularly surprised by this since there is some controversy that Watson and Crick's paper about the structure of DNA used unpublished data from Rosalind Franklin.

There is a catch. Attendees of the seminar must agree not to work on anything they pick up during the course of the presentation. This understanding and the honor system are what make people comfortable enough to discuss work they might otherwise keep private.

The presentations may also be a way to start collaborations. In a field driven by data, if someone provides the data for a figure on a paper, that person frequently becomes an author, even if the idea of the paper came from others. Thus, advertising results before they are published might allow other researchers to avoiding running the same experiments.

A consequence of this practice is that one rarely finds single authored papers and often finds papers with four or more authors. How does one delineate the contributions of each author? Author ordering may only give a coarse indication of an individual's contribution. An existing solution in some journals is to include an author contributions section. This section typically follows the acknowledgments and may read some like the following:
S.B.C. conceived and designed the experiments. B.S. conducted the experiments. S.B.C. and B.S. performed the analysis. S.B.C. and B.S. wrote the manuscript.
What happens if the work is primarily by two authors? The practice described to me for these instances is called co-first authorship. To do this, one simply places an asterisk next to each author's name with a footnote that reads: "These authors contributed equally to the work."

While some biologists I spoke to joked about some of these practices (one described how an author contributions section might read if each individual's contribution were described honestly), almost all of them were comfortable with the idea that providing data is a legitimate way to become an author on a paper. The same might not be true for my community, but I wonder if any of these practices would transfer well.

Friday, July 27, 2007

Genome Factory

About ten years ago, I spent a summer with other high school students for a summer program at the Waksman Institute of Microbiology. The program's goal was to introduce us to protocols to extract plant DNA and isolate regions of interest for sequencing. We learned how to use restriction enzymes to cut the DNA into smaller fragments, bacterial transformations to make copies of the DNA within E. coli, PCR to make copies of DNA without the help of E. coli, and gel electrophoreses to separate the DNA fragments by size and isolate the one(s) we wanted. Finally, the DNA had to be sequenced, and for this, we were introduced to the Sanger method, developed in 1975 by Frederick Sanger and his colleagues.

The Sanger method involves adding modified nucleotides called dideoxynucleotides, which can only form bonds at one end. Think of a Lego piece with a flat top. Thus, a DNA chain that has such a nucleotide will immediately terminate. If these nucleotides are mixed in with regular nucleotides during a process like PCR, it creates fragments of the DNA sequence with the same starting point and varying endpoints. If only a particular type of dideoxynucleotide such as dideoxyadenine (ddATP) is used, then all the resulting fragments terminate with an 'A'. If these fragments are then separated by gel electrophoresis, one can get a rough idea of the positions where 'A' shows up in the DNA sequence of interest. If 'C', 'G', and 'T' wells are adjacent to the one for 'A', one can just read off the DNA sequence from the gel electrophoresis. This is the basic principle of the Sanger method.

By the time school started again, we had become familiar with the techniques and protocols. We continued to return to the Waksman Institute periodically and apply these techniques. We would eventually use the sequence data from these visits to construct a phylogenetic tree of the Allium (i.e. onion) genus. Unfortunately, the data collection process could often be slow and annoying. There were many stages in which something could go wrong, and I would have to return to the beginning. All of this work produced just a tiny fraction of sequence information from these genomes.

A lot can happen in ten years. Thanks to my friends in the Broad's Outreach Program, I had a chance to visit 320 Charles St., the location of the Broad Institute's DNA sequencing facility. It is sometimes called a high-throughput production facility because of the rate at which they manage to sequence DNA. The facility was responsible for many of the sequences that were part of the Human Genome Project, and I was about to find out how they did it.

We entered 320 Charles St. and sat down for a presentation. Before we could start our tour of the facility, one of the scientists wanted to describe the process. To my surprise, she described the Sanger method. How could this be the process of a high-throughput production facility? Once the tour started, it became clear how: they industrialized the process. We had entered a factory, complete with conveyor belts, robotic arms, and computers. A group of technicians oversaw that the work on this genome assembly line went smoothly. Others, including the scientist leading the tour, were working on ways to industrialize new and improved sequencing methods developed by Solexa and 454.

It was interesting to learn that part of the rate increase has come from engineering solutions to scale up production. The amount of sequence data now available is enabling some researchers to ask questions that may previously have been too time-consuming to answer. I have talked to biologists this summer that have told me how challenging data collection can be, and I am starting to realize how those difficulties play a role in the questions they ask. How might these questions change if other protocols for data gathering were similarly industrialized?

Thursday, July 12, 2007

Outreach

I had been at the Broad for a little over a month, but I had yet to meet the co-worker standing next to me in the elevator. To avoid my tendency to shift between staring awkwardly at the elevator doors and the lighted floor number, I introduced myself. "I'm Megan," she responded, and we started a conversation.

Megan Rokop is Director of the Broad Institute's Educational Outreach Program. In addition to the research that goes on at the Broad, the Institute also sponsors a series of programs to engage with students, teachers, and the general public in the Boston area. A main feature of the program is the opportunity for high school classes to visit the Broad, where students get to conduct experiments using Broad facilities.

Megan wasn't always interested in biology. She started college at Brown as a foreign languages major, but a scheduling error placed her in a biology class. Unlike her previous experiences with the subject, which primarily involved memorizing a list of facts, the professor for this class presented the material in a way that inspired Megan's interest in the subject. "I wanted to be like him," she said of the professor.

Sure enough, Megan switched majors and eventually received her PhD in biology from MIT. After teaching at MIT for a few years, a fellow biology instructor told her about an opening for the Outreach position at the Broad Institute. Although she enjoyed teaching undergraduates, Megan recognized that not everyone benefits from a scheduling error, and saw the position as an opportunity to reach students while they were still exploring interests. When I told her I was interested in learning more biology, she was more than happy to oblige.

Our first lab involved identifying and mating different strains of Caenorhabditis elegans, a worm that serves as a model organism for investigators with interests ranging from genetics to neuroscience. C. elegans are only a millimeter long, so we needed a microscope to observe them. Once under the microscope, the distinguishing characteristics of mutant strains and sexes were clearly visible.

C. elegans
are divided into two sexes: male and hermaphrodite. Mating two of the mutant strains requires the transfer of a male and hermaphrodite onto the same dish. The offspring can later be counted to determine whether their traits were dominant or recessive. After a few false starts, I was able to use a special hook to transfer a wild-type (WT) male onto the same dish as an uncoordinated (UNC) hermaphrodite. While we couldn't see the worms without a microscope, we could see the tracks the wild-type was making as he searched for his uncoordinated partner.

The second lab involved running a gel electrophoresis with an application to paternity testing. Not all DNA code for proteins, and in the non-coding regions, certain strings repeat. The number of times these strings repeat can be used to distinguish individuals and determine heredity.

One way to distinguish the number of repeats is via gel electrophoresis. The idea is to load the DNA into different wells on one side of a gel and run a current through the gel. Since DNA is negatively charged, this current causes the strands to move across the gel towards the positively charged end. Since longer sequences diffuse more slowly, sequences with more repeats don't travel as far away from the negative end.

Unlike the first lab, I worked on the second lab with a group of high school students. They were visiting from the National Youth Leadership Forum on Medicine, a summer program for aspiring doctors. After the lab, I had a chance to talk to some of the students, who were curious what a non-biologist was doing at the Broad. In turn, it was interesting to hear from the students, some of whom weren't completely set on a career in medicine. While I wasn't sure whether their experiences that week would increase their interest in medicine, mine certainly increased my curiosity about biology.