How Lisp-family languages facilitate building bioinformatics applications
bib.oxfordjournals.org
bib.oxfordjournals.org
There is a large amount of really, really bad Perl code still in use in biology. Trust me when I say to those outside of the field, I've seen things you people wouldn't believe. I try to discourage Perl use between it makes it so easy for biologists to write terrible code.
I've found a compromise in Python. People will happily use it because it's trendy and ubiquitous. And I think it makes it easier to write OK code. But I miss using Lisp. The style of programming when you have a REPL is perfect for this kind of thing. The ability to quickly write prototypes which can later be improved is an advantage which cannot be stressed enough. C is fantastic when you have a good idea of the algorithms and data structures you want to implement and are concerned with implementing them efficiently. But that's rarely the case when writing bioinformatics code. I need to know if the damn thing works first before I spend time making it work fast. And why would you not use a language where both things are equally easy?
I'm happy to see this paper because it means that I have a way to justify my usage of Lisp in the future. Big thanks to the authors.
Other communities (e.g. statistics, machine learning) have a culture of ramming all their tools in a toolbox-library/framework (R, sklearn). This makes for a much tightly coupled environment, making it much harder to introduce Lisp work.
* https://github.com/kovasb/session
(Also my own early abandoned experiments: http://celeriac.net/iiiiioiooooo/ http://celeriac.net/iiiiioiooooo/public/ http://celeriac.net/iiiiioiooooo-dom/ http://celeriac.net/io/public/)
It doesn't address the browser-based point.
I used to write a decent amount of Clojure. I would be surprised if you would have problems with reviewers complaining about a nonstandard language -- unless you're publishing a library, then you might rightfully get dinged because the target audience would be small. But if you're just running a routine analysis, I have always found the reviewers don't care what language it's in.
Anyway, I dropped Clojure because the numeric support is just terrible. There's Java libraries, of course, but you lose the conciseness and functional aspect. You just can't performantly use arrays or matrices as if they were generic sequences, as appealing as it appears from a distance.
Considered Racket, but I consider lazy evaluation a must (a la Python's generators), and I could never figure out how to use Racket's lazy mode.
So, Python it is... Maybe choosing a programming language is like marriage: better to have something you can tolerate for a long period with all its warts than something you're infatuated with.
> I have a way to justify my usage of Lisp in the future
Haha, OK. I found this paper really unconvincing. No mention of the terms "numeric" or "array", and the only mentions of "statistics" and "machine learning" are in the form of "Lisp would be a really great foundation for stats/ML if someone would just write a library for it".
Furthermore, the disconnect with traditional mathematical syntax has also turned off many users. There was a predecessor to R called Lisp-Stat [2] which never achieved the acceptance that S/R did, and the Lisp-based syntax of Lisp-Stat was cited as one of the reasons.
Following the example of a high-profile Lisp supporter like Peter Norvig who accepted Python as a "Lispy" alternative early on, I think there are a lot of us who have come to accept Python as a necessary compromise. While Python is not homoiconic, functional programming is encouraged through its libraries rather than its core language definition, and the power of macros are missing, the inherent popularity and all the benefits that come with it is tough to beat.
[1] http://winestockwebdesign.com/Essays/Lisp_Curse.html
[2] http://homepage.divms.uiowa.edu/~luke/xls/xlsinfo/xlsinfo.ht...
So, exactly like bioinformatics, then. Trivial tasks are supported by packages of libraries like BioPerl, but all the original research is sui generis, with the attitude in general being that each problem is unique and only results get published, so who cares if the code is a write-only mess as long as it does the one job after which no one will care about it anyway. In such an environment, the "Lisp Curse" already exists anyway, so why not get the benefit of a more powerful and expressive language when you're already paying its cost?
(Granted, Lisp for bioinformatics is never going to become a generally accepted thing. But if you're not releasing code anyway, no one cares what language you write it in...)
This is exactly the kind of attitude that projects like 'Software Carpentry' (https://software-carpentry.org) are pushing back against.
The syntax referred to is S-expressions. They are a simple and elegant way of writing trees. Lisp was originally supposed to have M-expressions which would be more like mathematical notation, but it was found that Lips users actually liked S-expressions.
The reason becomes obvious after some amount of Lisp exposure. S-expressions are so simple that before long one does not even see them any more. The parentheses give way and one can just look straight at the Lisp code itself. For this reason many Lisp programmers do not even edit the S-expressions. They instead use something like paredit to edit the underlying tree directly.
It's similar to how readers don't really "see" the letters that make up a word in many natural languages. And, further, they don't see the words that make up certain expressions. We have this fantastic ability to internalise language but so many are choosing not to use it when it comes to programming.
That is the same with most things, we say the verb before the object(s).
I agree that the concepts behind S-expressions are simple and logical, but expressing an entire program using S-expressions is an additional thing that must be learned. Other languages chose a syntax that is closer to what people are used to and that helps with the adoption of the language, even if the language loses other capabilities.
On the other hand, I have taught courses at the high school level at the same school based on "How to Design Programs" [1] using the Racket language. Although I can explain the concept of (+ 6 (* 1 0) (/ -2 2)) rather easily, the students make mistakes writing expressions for several weeks.
Plus when you're used to trees, linearization and parsing, you *fix the way you see fit, pre post or in ...
I've used XLispStat, S-PLUS (a commercial version or R), and SAS. SAS had the most comprehensive statistical libraries, but I never liked the language. I liked xlispstat but the statistical libraries weren't as good as those of SAS or S-PLUS. For data mining, I settled for S-PLUS, which was sufficiently lispy, and had a REPL. So it wasn't just the syntax (or lack of it).
Beyond being dynamically typed and having a REPL, Python isn't particularly Lispy: it's less functional, and it has no (linked) lists or symbols. It's quicker to develop in than most of the non-Lisp alternatives, though.
There's also Pixie, a python lisp to LLVM system, that was specifically made for high performance.
worth it
95% of my python programming time these days is spent inside a repl.
Principles of Biomedical Informatics, Second Edition, 2013 https://www.amazon.com/Principles-Biomedical-Informatics-Sec...
EDIT: (I consider Racket separate from Scheme)
BioLisp would be a great project.
However, in Bioinformatics, the reality is that many users are still using Perl. It would be get to have them migrate to, modestly more readable languages.
Outside of scripting, alignment and other performance sensitive code is usually written in C (or C++). As with many academic projects, the code quality unfortunately is quite low.
Stringy code thats hard to decipher can be written in any language I have discovered.
My take, Perl is kinda on the way out, though it was used extensively. BioPerl is a nice package.
Biologist like R, it pretty quick and behaves like they aren't programming. It can make graphs nicely.
Python seems to be the go to compromise. BioPerl is a really nice package. But people start to want to use it for big things and to get it performing adequately requires a lot. There is confusion from the researchers that scares them away: between python 2/3 numby, pypy. BioPython packages are pretty excellent.
If it needs to be fast and crunch large sets of data (fairly common) the tool is in C or C++. We should start using Rust..
We use Php (Silex) to deliver a front end and some quick database lookup and display. Its replacing perl for this.
This is definitely all relative, I would argue lisp is easier to read given that the philosophy of its syntax is that there is no syntax...
Though if you've only ever read classical languages I can see how lisp would seem alien
There are some people who do write beautiful Perl, but they are few and far between (again, in my experience).
I spent a year as a staff member of a genomics institute, and researchers occasionally came to me for help getting the local sui generis tooling stack set up and working sanely. (Well, I say "stack"; jwz's bookcase-from-mashed-potatoes metaphor hastens to mind.)
Having seen the miseries they went through, I don't think it would necessarily have to be all that hard to make migration look good, especially to a language whose syntax is other than wildly irregular and whose performance is other than usually abysmal.
We have some customers with zero lines of Python code, but they do use R and Tableau.
All the software used to talk to the devices and do ETL processing or graphical analysis is then written in Java or .NET, depending on the department and their set of OSes.
Functional versus procedural programming is first and foremost a usability problem for developers. The two are interchangeable as far as being Turing complete.
I just spent a week at a Jakolb Nielson's education course on usability. I asked the executive vice-president at the company if they had branched into software development itself, sadly the answer is still no.
Functional programming for the majority of developers is not usable. The article mentions they are are targeting DSL, domain specific languages. Which is fine. For example, Haskell has seems to have found a boutique community in async, MQ world. But, from a usability perspective one is severely limiting ones pool of developers choosing a functional language.
It is fascinating from a human nature perspective that people who think in functional programming are brains are wired differently from the procedural programmer. There is very visceral reaction when either camp is asked to program in the manner they find least usable.
Ultimately I think compilers will get to a development stage where independent of functional programming or procedural programming approaches the same optimized code gets implemented under the hood.
Postgres was originally written in LISP. Ultimately Stonebraker had to make the call that if Postgres was going to get adopted widely, LISP had to go and it was rewritten in C/C++ before being open sourced. As a point of research you all might want to look into the Postgres experience.
There are other cases where an original Lisp implementation has been rewritten: ViaWeb and Reddit. I think the reason is that the teams which took over the projects were unfamiliar with Lisp.
My experience is different, though I should mention that I'm a sole developer. I'm roughly equally experienced in C and Lisp, but find I'm far more productive in Lisp. My first language was FORTRAN. Unlike C, Lisp is memory safe and you don't have to manage the memory yourself, interactive, strongly typed, you never need to write a parser, the built in symbol and list types are incredibly useful, and there are fewer "gotchas".
Lisp is multi-paradigm rather than purely functional, so like the Algol derivatives it has loops and destructive assignment.
But Perl is still used more widely than Lisp in bioinformatics. or lisp can be neglected in bioinformatics compared to Perl.
Check the status of BioPython, BioPerl. BioLisp? not even exists.
For example, this nonsense statement reveals the author's bias: "Clojure is a rising star language in the modern software development community." Huh?
Could selection bias be skewing these results?
Proficient Lisp programmers can certainly create shorter and faster programs with Lisp. Who would ever contest that? Average programmers, on the other hand, can probably develop in a similar amount of time and write faster programs with C++ (given the amount of libraries and information available - in comparison to Lisp - and the ubiquity of tooling).
on average, the Lisp programs ran significantly faster
than the C/C++ programs and much faster than the Java
programs (mean runtimes were 41 s for Lisp versus 165 s
for C/C++).
The only thing I can say is that their C/C++ code has serious problems.http://nar.oxfordjournals.org/content/37/suppl_2/W28.full-te...