Ref: https://www.technologyreview.com/2007/09/04/223919/craig-ven...
Ref: https://www.technologyreview.com/2007/09/04/223919/craig-ven...
And it is worth noting that despite all the critiques of shotgun sequencing at the time as "cheating", chromosome walking sequencing is dead, and shotgun is what we use today. If anything replaces that it will be long read technology like that done by Oxford Nanopore sequencing, although we aren't really there yet (you still need to assemble that data and don't really get end-to-end sequences yet).
Mother: "Amy, dear, you hardly touched any of your vegetables. Most of them are still on your plate."
Amy: "That's not true, Mom; I completely sequenced them into my tummy using the shotgun method."
And the cost of a T2T using a combination of these technologies is already well under $10,000 in reagent costs. The assembly is still a complex art especially for a messy chromosome like Y.
Does that mean that every possible polymorphism is accounted for?
Does that mean that 100% of possible polymorphic sites are known?
What about weird edge cases. You have a tandem repeat that is known, but in some small population of humans, there's two tandem repeats with an island of something else in there.
Imagine it like putting together a jigsaw puzzle, except that the picture on the front of the puzzle has lots of repeating motifs, so you end up with multiple pieces that look identical. You won't be able to tell how the whole thing fits together, but you can assemble the bits that are well-behaved and unique.
Modern technology gives us larger jigsaw pieces, which allows us to distinguish between almost-identical parts of the puzzle better. But I would note that the project linked did a huge amount of sequencing using very expensive methods to be able to resolve the whole thing.
Modern sequencing techniques allow us o read 200 to 500 bases at a time. So after hat we need do find a way o arrange these short sequences into a single sequence. And this can be prety hard, especially when you are doing 'de novo' assembly[0].
Besides that, there is the fact that some regions of DNA are repeated[1].
[0] - https://en.wikipedia.org/wiki/Third-generation_sequencing#De...
My knowledge cut-off period in this domain is around 2018. So it's not suprising that things moved on.
Because the sequencing uses short reads, it will not be able to resolve parts of the genome that are repetitive with a repeat unit longer than ~150bp. You'll have all the jigsaw pieces from the puzzle, but you won't be able to reconstruct some parts of the picture. Long read sequencing can help with those, but that's more expensive.
35 years ago: https://www.nytimes.com/1987/12/13/magazine/the-genome-proje...
33 years ago: https://www.nytimes.com/1990/06/05/science/great-15-year-pro...
29 years ago: the competition gets fierce https://www.nytimes.com/1994/02/22/science/scientist-at-work...
24 years ago: the race is alive! https://www.nytimes.com/1999/03/23/science/who-ll-sequence-h...
23 years ago: draft is complete https://archive.nytimes.com/www.nytimes.com/library/national...
20 years ago: https://www.nytimes.com/2003/04/15/science/once-again-scient...
2 years ago: https://archive.nytimes.com/www.nytimes.com/library/national...
The celera approach using shotgun was tricky because assembling the bits required a great deal of computational finesse and horsepower; "The final assembly computations were run on Compaq’s new AlphaServer GS160 because the algorithms and data required 64 gigabytes of shared memory to run successfully."
At the time the GS160 was the monster machine, a classic big iron UNIX, which could be tightly clustered- unified IO/filesystem between all the machines. The primary author of the shotgun assembly was Gene Myers, who previously had written BLAST, and invented the Suffix Array with Udi Manber.
The public project had its own issues, as many of the teams assigned to work on it still had the "cottage industry/artisinal/academic" approach, then Eric Lander came along and turned it into an industrial process, and parlayed that into running the Broad Institute, a privately funded MIT/Harvard research institute in Boston.
These days petabytes of sequence data are generated every day and stored in clouds. The genome has been a fundamental tool for shaping our studies of humans, although its true potential for understanding complex phenotypes remains elusive.
So... um... BLAST processing?
As someone who is in the genomics world now as a software person, I find it amusing that I've gotten more into perl over the last year or so. It has its place, in the way that grep/sed/awk/etc does.
IMHO the person who "saved the genome project" was WJ Kent, who developed the assembler, BLAT, that the public project needed. I strive to point out that he wasn't a sole hero, nor was Lincoln. What I really like about BLAT is that while Celera was using Big Iron UNIX (massive 64-bit 64GB machines with 10s of terabytes of central high performance storage), BLAT ran on a cluster of linux machines, right around the time that people were waking up to the fact that linux was becoming a useful tool for scientific data processing. BLAT's design allowed it to work on a cluster of cheaper/smaller machines, while celera's algorithm really needed a massive shared memory single machine. it was sort of a microcosm of the larger battle being fought between Big iron UNIX and little intel linux at the time.
smaller organizations were realizing they could get slower computers power pretty cheap, and they could get a whole lot of them. There was a little flurry of activity with flocking algorithms, objective-c had a little renaissance due to swarm computing. probably the most famous and lasting was map-reduce, from alphabet, but they were called google back then.
there were a bunch of clever little tricks, like channel bonded network cards to make multiple cards look like one fast card, so you could double or triple bandwidth.
The beowulf cluster joke was kind of the spirit of the scrappy, make something cool out of junk approach while calling companies like sun, hp, compaq dinosaurs. and it continued for a while with people building stuff like this - https://ncsa30.ncsa.illinois.edu/2003/05/ncsa-creates-sony-p... I think there was some weirdness with export controls of ps2's for this kind of stuff.
I don't remember what pc chip was the new hotness back then, it was a good 20 years ago and my memory is dim.
but that, I think, captures the gist of the meme.
However, I got into the genomics world in the early aughts, and the perl hung on and on and on. I remember starting a new job in the mid-teens, and one of the first tasks I had was to port over a legacy perl script. Such is the world of scientific software. 99% of the people I know who have touched perl since ~2003ish are in the bioinformatics space.
And if I remember correctly BLAT was so useful because it could be run on machines with less cpu power by loading more data into memory…or was it the other way around?
> The primary author of the shotgun assembly was Gene Myers
No conflict of interest there!http://news.bbc.co.uk/2/hi/science/nature/2940601.stm
> The remaining tiny gaps are considered too costly to fill and those in charge of turning genomic data into medical and scientific progress have plenty to be getting on with.
> The decoding is now close to 100% complete. The remaining tiny gaps are considered too costly to fill and those in charge of turning genomic data into medical and scientific progress have plenty to be getting on with.
OP is being silly. Nobody was fooling anyone.