How Perl Saved the Human Genome Project (1996)
bioperl.org
bioperl.org
Any supporting evidence to that? If you look into it, you may be surprised.
Here's a hint, Perl doesn't necessarily optimize for the same things other languages optimize for.
The very purpose it has survived three decades is because it solves a great variety of problems, in a very centralized design philosophy which no other tool has addressed till today.
In terms of finding genetic causes of disease, this happens every day, and is practically mundane, but there are two complications towards getting to cures. First, most disease is far far more complicated than a single gene; any single gene may account for just a percent or two of what we call the same disease. Second, knowing which gene is broken does not provide a cure for that disease; even for a given small molecule, determining if it will interact with a gene's protein or have any effect on that protein's structure or function is a task that physics has not been able to tackle. Additionally, the genome has only been available for a mere decade, and for many if not most diseases, the process of going from a known gene target an approved drug is going to take far longer than 10 years.
So the HGP has fueled a huge amount of discovery, is the foundation of nearly all human biology research, and is completely indispensable, but in terms of new cures for various diseases it has not delivered, yet, but really it shouldn't have to.
http://techcrunch.com/2013/04/23/counsyl http://techcrunch.com/2013/04/25/sv-angel-health-informatics
http://www.nytimes.com/2012/07/08/health/in-gene-sequencing-...
This was just one person, sure, but there are a lot of smart people working very hard on making this scale.
Sequencing technology has gotten to a point where it's just blown Moore's Law absolutely out of the water and we can't throw more compute at the analysis problem, we have to make it smarter. The reference genome is used in how that's been made smarter.
It helps to discuss a little bit about how the HGP reference was produced, and why producing it took 10 years and three billion dollars.
The HGP process first had a map made, where the genome was broken into lots of smaller segments. The idea was that this reduced your problem space; any segment of DNA produced from a sample from that portion of the genome came from that area. Then that segment was broken into lots and lots of smaller chunks and then read on the sequencing machines in 600-800 base segments. By the time that sequencing technology reached "max level," the state of the art machine could generate 96 of those segments in an hour's time.
Then you'd calculate overlaps and assemble those smaller "reads" back into a sequence of that chunk you chose from the map. Then someone would audit the computer-generated assembly by hand, possibly ordering up more lab work to fill any gaps or resolve areas of crummy data. Repeat for the next chunk from the map.
Now here's how things work, when we need to do any sort of genomic analysis on an individual:
New technology has the ability to sequence human genomes at deep coverage in 11 days[1], and cranks out 6 billion reads 100bp long from places all over the genome. Computationally, this is an absolutely different animal. You can't feasibly try to re-assemble these reads into a human. So, what we do is use string matching algorithms to "map" a 100bp read back to where it most likely came from, using the HGP genome as a reference.[2] Since obviously your DNA does not match the HGP reference base-for-base, and mismatches/insertions+deletions are really where the interesting data is anyway, there's some leeway for mismatches in the mapping.
At that point, by mapping reads back to where they came from, we end up with a data file that represents an individual's genome. You're able to walk across the genome base for base and ask "So, base 347 of Chromosome 7 is a T in the reference, what is the most likely base on Joe's genome at this point given the reads we have that span this base?"
Mapping things to the reference also allows us to attempt to find really interesting stuff that can cause disease, such as structural variations in the genome. These are instances where large segments are removed, duplicated, inverted, or picked up and moved somewhere else relative to where they "should be."
[1] http://www.illumina.com/systems/hiseq_comparison.ilmn
[2] http://bio-bwa.sourceforge.net/ is the tool that's most popular these days.
Too much perl code is essentially write once and forget. It gets results quick but it is a disaster for repeatability which is an essential part of science. I've worked on bioinformatics perl projects where bugs canceled each other out (ie code that was supposed to clear an array and repopulate it with corrected values did neither so the original values were returned). And I've spent far too many hours trying to figure out what a perl script that is the reference implementation for a certain procedure actually does.
Their certainly still are scientists who use it but python and R are gaining ground for good reason.
Wiring together analysis pipelines with pipes as they describe is, however, an excellent technique regardless of language.
Please stop repeating this misconception. People who put little effort into learning how to program write "write-once" code in any language. Perl had the "misfortune" of being the only dynamic language on the block for a long time, leading to many people reaching for it to get things done without bothering to actually learn the language, thus creating a vast corpus of low quality code.
(It does not help that the definitive resource of Perl for bioinformatics people, which i've seen in libraries like those of the Genome Campus in Cambridge, isn't worthy of being used as toilet paper, yet influenced a whole generation of scientists.)
> I've spent far too many hours trying to figure out what a perl script does
How often do you reach for perltidy when you do this?
Here's the book I've been using: http://bixsolutions.net/
As for learning Perl, the best resource is http://perl-tutorial.org
You will find an explanation there of how to judge a learning resource and a list of the most current texts that teach modern Perl.
* chromatic's "Modern Perl" - which has a free CC licenced version (although it would be nice for you to pay some money to support the good work of the author and publisher ;-) http://www.onyxneon.com/books/modern_perl/index.html
* Curtis Poe's "Beginning Perl" http://ofps.oreilly.com/titles/9781118013847/
(Bias warning: Curtis is a mate - but it is a very good book on the topic. Honest ;-)
Perl is the language with the highest occurrence of "subtle" and "ambiguous" in its documentation and tutorials that I have ever seen. Humans may be subtle and ambiguous, but programming languages should be clear and precise.
Any language can be used to write a riddle. Even Python[1][2]. It's the community's values that determine code quality. The banner of "impossible to read" hangs over the Perl community and it has been my experience that it reminds us to write much better code.
[1] http://www.python.org/search/hypermail/python-1993/0232.html [2] http://p-nand-q.com/python/obfuscated_python.html
It hangs there, put there by other people. Sadly it correctly would hang above people writing perl but outside the community. We have many efforts going on to get people to write nice and readable Perl.
Python is barely better.
None of these languages come anywhere near a reliable verifiable formal model of what they claim to compute.
In my experience as a global contractor, there are at least big players who use Perl exclusively, and in Perl irc channels i see bio informatics people very regularly.
UCSC has switched from teaching Perl to teach Python in intro bioinformatics classes (they get 5-10 new PhD students a year, plus at least that many masters students, I think?). Nobody that I encountered at my postdoc institution used Perl, maybe about 60 people who programmed, and I knew enough of ~20 people's habits to know their programming choices. R, Python, and C/C++ were commonplace, with a few weirdos using Ruby, and of course lots of shell, awk, and sed glue to keep it all together. The Broad has been somewhat successful on shoving Java into people's pipelines with GATK and Picard, but it's not a welcome addition to many people's habits, and I haven't encountered any significantly used project around next gen sequencing that is based in Perl. For example, in the RNA world, all the common tools like Tophat, Mapsplice, and ChimeraScan all use Python in ways that Perl would have been used 10-15 years ago.
I have active collaborations with about five different labs that also have in-house informatics for everything from microarrays, to ChIP-Seq, to RNA-seq, exome to whole-genome resequencing, and nobody uses Perl for any of it.
That said, it's really easy to collaborate with lots of informatics people and never even need to know what they use internally. What's this about contracting in bioinformatics? I've never encountered such a thing.
As for contracting: Oftentimes bioinformatics people will realize that they're not good programmers at all and call in people to help make their code less of a rat's nest and learn from them at the same time. It happens quite often here in Europe.
BioPerl: https://github.com/bioperl/bioperl-live
BioPython: https://github.com/biopython/biopython
People using BioPerl just use it to get something done and don't care to join the larger community. In Perl we call it the DarkPAN.
Meanwhile anyone wishing to use Python will at least have needed to make an effort to learn the language, since there is nothing else quite like it; and as such probably more likely to invest further effort.
But, i do admit that this is conjecture.
You should learn Perl because you're likely to encounter a lot of Perl code in the wild, and because you can learn something from it. Knowing how to generate a Perl one-liner that does something incredible will take your CLI skills to a new level in a way that knowing Ruby will not.
Perl is still one of my go-to languages for sysadmin scripts, because it's so concise and powerful. It's a long-beard language.
"...because you can learn something from it" is a far better reason, because it's actually valid. Now, specifying what you're suggesting could be learned would be a good step, but I'd take it as-is without a problem.
The former, though? No, Perl is no longer occupying visible roles to the point where any new programmer can assume they'll be seeing it at all, much less "encounter a lot of Perl in the wild." It's out here, sure, but it's hidden in commands we use or scripts invoked by background tasks non-systems-programmers have absolutely no reason to touch.
Learning to do irresponsibly illegible things in Perl will teach you some great regex and... Perl. Learning Ruby (or Python, or Java, or even PHP) will get you actual work -- that doesn't involve being sequestered in backwards academia -- far more often now than Perl.
Please only say such things when you can back them up. Yes, learning Java will get you more jobs (though, who really wants to do Java for a living?), but you'll still enounter Perl in more jobs than the other languages you mentioned:
http://www.indeed.com/jobtrends?q=Perl%2C+Ruby%2C+Python%2C+...
When programming in a language you hate, don't take naming arguments for granted. It could be worse: you could be programming in Perl, where all arguments come in a flattened array.
sub named_params {
my %P = @_; # named params in hash %P
my $foo = $P{bar} + $P{baz};
return $foo * $P{multiplier};
}
my $rvalue = named_params( foo => 1, bar => 2, multiplier => 3 );
Or, some people prefer to just pass in a hash reference: sub named_params_hashref {
my $P = shift;
my $foo = $P->{bar} + $P->{baz};
return $foo * $P->{multiplier};
}
my $rvalue = named_params_hashref({
foo => 1,
bar => 2,
multiplier => 3,
});
But of course, why bother letting common sense, and a single line of code stand in the way of your chance to deliver a pithy remark?>>where all arguments come in a flattened array
That is optional, through the power of CPAN (and kbenson's methods, of course). Examples of alternatives:
http://search.cpan.org/~barefoot/Method-Signatures-20130222/...
http://search.cpan.org/~flora/MooseX-Declare-0.35/lib/MooseX...
http://search.cpan.org/~ether/MooseX-Method-Signatures-0.44/...
(Depending on if you use Moose or not.)
One of the cool parts of Perl is that the language is extensible, that is why it e.g. has a better OO than Python, Ruby, etc. (Or simple typing of parameter arguments, see the modules above. Or why MooseX::Declare has ... ah, look yourself, troll.)
Edit: The possibility to extend your environment is generally seen as a good thing, unless you believe in cleanliness of bloo... languages for ideological reasons.
You seem to be proving what I said earlier[1]. You think you understand what's going on, but in truth you don't. Feel free to continue making assertions regarding things you don't know, I can't stop you, but be aware that to do so is dishonest.
Sure, couldn't have anything to do with the language. The whole rest of the world just doesn't get it.
Never claimed that. If you read the quote you chose, you will find that i claimed that many did not even bother to try and get it, because they got things done and did not have any reason to get it.
If it looked like Lisp, people would be less likely to think that it's just a matter of applying their C experience, but alas, it generally looks pretty familiar, if a bit messy, to users of other imperative languages.
If you are trying to understand, write or change a Perl script, and you don't know what context a statement takes place in, or what I mean by context in this case, then you don't know what you are doing. (I mean you in the general sense, not as an indictment against the parent).