Perl Saved the Human Genome Project (1996)
foo.be
foo.be
> Perl is remarkably good for slicing, dicing, twisting, wringing, smoothing, summarizing and otherwise mangling text. … Perl's powerful regular expression matching and string manipulation operators simplify this job in a way unequalled by any other modern language.
Indeed! I still think in terms of PCRE (Perl-compatible regular expressions), and I love that Perl makes regexes a first-class citizen.
> Although the biological sciences do involve a good deal of numeric analysis now, most of the primary data is still text: clone names, annotations, comments, bibliographic references. Even DNA sequences are textlike. Interconverting incompatible data formats is a matter of text mangling combined with some creative guesswork.
This is still true! One common format (arguably the most common format) for sending around bits of sequenced DNA is the FASTQ format (https://en.wikipedia.org/wiki/FASTQ_format). FASTQ files are (ASCII) plain text, making them really easy to parse. Of course one byte per letter of DNA is wasteful, so FASTQ files are commonly exchanged GZIP-compressed, with the .fastq.gz extension. Many platforms & tools read in or write out .fastq.gz automatically, saving you the (de)compression step.
-P, --perl-regexp Interpret I<PATTERNS> as Perl-compatible regular expressions (PCREs). This option is experimental when combined with the -z (--null-data) option, and grep -P may warn of unimplemented features.
(Scroll down to this plot: https://swtch.com/~rsc/regexp/grep1p.png
But sometimes you just need to throw something together fast. Eg search and replace in Vim uses regexps.
Like match in functional langs or lisp matchers https://www.cliki.net/pattern%20matching Unification is so much better than regex hacks.
Regexes are not important. But they are a tiny bit of string algorithms.
Not the op. Yes, I've done some work on this, then tested it against some of the software used in various environments for this kind of work and more than once spotted alternative, more efficient alignments. The practical upshot of that is that I ended up wondering if there is ever a serious bug found in such a piece of software if it shouldn't automatically cause all papers that used the software for their work to be at a minimum flagged for an additional round of review as well as potentially from being disqualified.
What also struck me is that the people using this software treat it like a black box, they have absolutely no way of verifying that what it did it did right.
It must, but it doesn't always do so.
As mentioned in other comments, sequence analysis is probabilistic, so "matchers" instead tend to be statistical models, like HMMs. There is a rich relationship between statistical models like HMMs and parsing theory.
[0] https://virological.org/t/novel-2019-coronavirus-genome/319 [1] https://www.ncbi.nlm.nih.gov/nuccore/MN908947
> Perl programs are easy to write and fast to develop. The interpreter doesn't require you to declare all your function prototypes and data types in advance, new variables spring into existence as needed, calls to undefined functions only cause an error when the function is needed. The debugger works well with Emacs and allows a comfortable interactive style of development.
I think each and every language that could undo this second "billion dollar mistake", did ("strict mode").
Of course over time, these same people learned to program better, and its looseness and general wackiness grew into a liability.
But the important point here is: language design decisions (just like product decisions) aren't so much intrinsically right or wrong; but right or wrong at certain times.
In its heyday, Perl was, for many people, definitely the right way to go.
Interestingly almost all of the petabyte scale and beyond processing of genomes (whole or exome) is done on the JVM as the library and toolkit ecosystem is extremely mature and it is significantly more performant than just scripting things in Perl. Having access to the big data ecosystem that runs on the JVM is also another reason why languages like Java and Scala are found in the high performance areas of genomics.
How Perl Saved the Human Genome Project (1996) - https://news.ycombinator.com/item?id=5655165 - May 2013 (63 comments)
How Perl Saved the Human Genome Project - https://news.ycombinator.com/item?id=1568109 - Aug 2010 (26 comments)
How Perl Saved the Human Genome Project - https://news.ycombinator.com/item?id=631683 - May 2009 (8 comments)
Curtis Poe's Beginning Perl is a good all around introduction but I'm unsure if it's a good intro if you're aiming specifically for the text mangling side.
For an uber-quick-start, I'd go with https://qntm.org/perl_en and then getting up to speed on the regexp side using the man pages linked from here: https://p3rl.org/RE
We had to make do with less. Perl was one of those tools that let us get our work done, as it enabled us to automate our tasks. Easily.
Later, I worked at SGI when all the kerfluffle about DVD decoding/playback was about. My recollection was that Perl was used for that as well[1].
When I started my company, most of my code was in Perl, with a little in C. I was using it daily until we closed in 2017. Now its a bit more sporadic.
At the day job, its Python everywhere. I've been told people would laugh at using Perl. Kind of a shame, as I see multiple pages of code that could be easily reworked into a far more readable, easy to reason about and comprehend small set of Perl code.
Perl really isn't great at mathematical operations, but then again, neither is Python without its C/C++/Fortran extensions. To Perl's discredit, the whole FFI bit took too many years for them to get right. Its there now, but the momentum is now behind other languages.
Again, a shame, as Perl is unmatched for its data wrangling capability. I didn't need 2 languages for the work I did in Perl. One sufficed, as I was happy to have simple language to reason about/support. This is curiously, why I like Julia so much.
[1] https://www.computerworld.com/article/2800097/seven-lines-of...
Could you give an example? I would be curious to better understand python’s limitations.
Python in contrast emphasizes the "one statement per line" rule which gives somewhat more elaborate codes. And I am not even talking about "boilerplate codes" as in Java. Nevertheless, python is infamous for its very well readability, which is beyond the average Perl code.
That line tickled my funny bone.
Having worked in insurance, this piece makes me wonder what languages are typically used to write software in insurance or other business medical type settings.
Insurance is just drowning in data and entry level claims processors have to interact with various databases all day long. Every time they update the software, they introduce new glitches.
I have a certificate in GIS, which involves some exposure to how databases work. So I was apparently more talented than average at parsing exactly what went wrong and how to get around weird new glitches.
Not really the best use for such spiffy training as I was not appreciated and would have been making vastly better money had I ever managed to get a job in GIS instead of insurance.
It would probably be simpler if it were like a language.
I'm not sure if it's different for medical insurance, but I worked with a company that did other types of insurance (mostly life with a smattering of others), and it was Java as far as the eye could see. Most things that pre-dated Java were done in COBOL.
Most people in this forum probably work in the internet industry. But Perl was huge in nearly every industry back then.
Things come and things go, but some tools leave behind an immense legacy and culture, Perl is one of those tools. It's not going anywhere and will come installed default in a lot of unixy distributions for years to come.
I don't think that's accurate. I've been doing Perl professionally for over 20 years, and I've run a Perl-centric recruitment agency since 2014, and I don't think Perl has been the defacto choice since circa 2004?
"Perl is remarkably good for slicing, dicing, twisting, wringing, smoothing, summarizing and otherwise mangling text. Although the biological sciences do involve a good deal of numeric analysis now, most of the primary data is still text."
Which seems fair. Perl is very good at that sort of thing.
(I worked in bio as we were moving away from Perl)
I don't honestly think it's a bad thing on the science side that python's won for -programs- simply because maintainable perl at scale requires focus and discipline that is frankly energy the average scientist would be far better of expending elsewhere, but even if it's only really in commercial settings where large scale perl is still worthwhile it definitely still has its place in pipelines.
Then again, Java, PHP etc. programmers seem to have the same experience on a regular basis, and these days the naysayers seem to have come for Rails as well.
Programming is, as ever, a pop culture, and given our tendency towards self deprecatory humour as a community the worst part of the whole thing for us a lot of the time is that 99% of the people criticising/insulting perl are so bad at it.
[1] https://wiki.bash-hackers.org/syntax/expansion/proc_subst
Many years ago I read a story of how a (lowly) grad student watched the battle between private industry sequencing the genome and the university/research team doing so.
He was worried the genome would end up in private hands.
And so stayed up several nights and wrote tight C++ code to line up the data and raced the private team over the finish line.
Maybe it ended up the core of BLASE?
EDIT: I had completely forgotten about BLAT which was in fact developed by Jim Kent!
I don't understand. Why would the university team have to finish it before the private team? If the private team finished it and encumbered access to the output, and the university team finished a short while later, wouldn't that have the same effect? Or would there some sort of funding drop-off once the task was achieved, that wouldn't account for whether the output was made accessible?
Although this was a race it's fun to note that both sides we're learning from the other and using resources produced by the other. The assembly techniques were worked out by Gene Myers for instance, who worked at Celera.
As someone who is currently a bioinformatics PhD student, Perl has become the vinyl of programming languages. While some older folks script with Perl, most packages I’ve seen published recently use either R, python, c++, and more recently, rust.
Our new stuff tends to be in R or python though.
Turns out it's not an accident, though! There were a lot of good decisions that kept the language stable, and test frameworks and code coverage have been a critical part of Perl and its modules since the 1980's. In fact, pretty much all of CPAN gets tested on a matrix of Perl versions and operating systems (ex: http://www.cpantesters.org/distro/N/Net-Amazon-EC2.html)
$ perl -Mwarnings -E'sub f { $c++; say $c if 0 == $c % 1_000_000; f() } f'
Deep recursion on subroutine "main::f" at -e line 1.
1000000
2000000
3000000
4000000
5000000
6000000
7000000
8000000
9000000
10000000
11000000
12000000
13000000
14000000
15000000
16000000
17000000
18000000
19000000
20000000
Terminated
You get a warning after 100 calls which almost always indicates a bug. In case it's a genuine deep recursion, the warning can be easily suppressed with `no warnings "recursion"`.On my computer, the program continues to run for about 10 seconds, consuming 9 GB virt./res. after which it is killed off by `earlyoom`.
i remember silly java consultants rabbiting on about TDD and agile while dismissing oss.
meanwhile, the curation of both testing and documentation as well as overall code quality on CPAN was light years beyond the best corporate code i've ever seen.
i'd argue that perl (with CPAN) was the first internet native programming environment.
Most of the EDA CAD tools use Tcl and some of the younger people write their scripts in Python but most of the other 40-50 year old engineers still write Perl. I write all of my small home Linux sysadmin scripts in Perl as well. I started to learn Python. I'll probably get around to it sometime but I don't write big systems and 90% of what I do is text processing.
TY LW and Bruce Winter.
It’s definitely a language that was ideal for its user base… admins and early web people mostly. All of my serious coursework in college was C, C++ and Fortran, so just being able to bang out code and get shit done was amazing.
i must have been a small kid back then but i remember thinking "boy, i wish i had those so i could check mine". in hindsight, they must've done a pretty good job of explaining the project because kid me could understand what they were trying to say. kudos to them. i have never seen it in the last 15 odd years, maybe more but today i can recall "elephants" in the ad for some reason.
My lab, that does a mix of bioinformatics and population genetics, has people who know bash, python, and R. Honestly, I think bash is probably the one that gets used the most. So maybe my command line wouldn’t look that different to the ones used in the human genome project. Although technically I am using zsh most of the time.
edit: ah yeah, the mod_perl book.
[1] https://metacpan.org/dist/Mojolicious/view/lib/Mojolicious/G...
I'm grateful Rust now exists and is able to match, and even beat, Perl in regex text surgery.
The simple answer is no. There were Wx bindings at some point, bindings into Cocoa, but as a strong Perl developer who uses a Mac -- and was once pretty handy with Tk -- and would check back every year or so, nothing ever really seemed to stick or be easy to set up.
IMO, it is an inane language with amazing regular expressions.
https://wiki.c2.com/?PerlIsNotAnAcronym
I've been using Perl and Python for over 20 years. Both are delightful, in their own way.
Disliked is horrible syntactic and semantic complexity then as I do now.
However, that didn't stop people from doing great things with it.
Now, I have to find a bug in someone's Perl program