Fast high-level programming languages
lh3.github.io
lh3.github.io
I think this is the crux of it. Python may not be the fastest language out there, but it's definitely one of the fastest languages to develop with, which is why it's such a hit with the academic research crowd.
Static typing and compilation are great features for many use cases, but if you are just prototyping something (and most academics are never going to write "production" code), it's nice to be able to try something, test it and then iterate and try again. When you have to compile, that iteration loop takes longer.
A few people here mentioned GoLang as a good alternative. Personally, I love the language, but it's really not very good for quick prototyping. As an example, the code won't compile with unused variables. This is great for production code but very impractical when you are testing some new algorithm.
As an anecdote, a friend of mine is an applied mathematician specializing in PDE and fluid dynamics. He used to code with C++, because that was what was used in the labs where he worked. Then he discovered Python and he thought it was so much easier, and still quite fast thanks to optimized libraries (written in C++ with a Python interface). But after a few months his enthusiasm had disappeared because of the runtime errors of his Python code. He didn't want to go back to full C++, though, and he still codes mostly with Python.
Exactly. It seems like he understands the tradeoffs and prefers quick iteration over "correctness".
In my experience, for computational "academic" code, the hard part is often coming up with the correct algorithm, not implementing that algorithm correctly. In software engineering, it's often the opposite.
Also, for what it's worth, I recreated his benchmarks in Julia and found a startup time of 5-6 seconds on my laptop. Not nothing, but compared to the time any FASTQ-processing operation usually takes, it's insignificant.
Julia' compile time latency does prevent you from e.g. calling a short Julia script in a loop, or from using Julia scripts that both have large dependencies and do very short-timed tasks. But I find that to be not very common in my workflow - especially since I usually don't use Julia as a small glue layer between C programs the way Python is often used.
https://github.com/Morgan-Stanley/hobbes
It’s kind of a structurally-typed variant of Haskell, integrates closely with C++, produces very fast code.
I'd like to grow it into a successor to kdb, though kdb is very entrenched in finance and the company is extremely litigious (they threatened to sue me).
F# is great too, and all of the other ML variants. We'll get there eventually, it's inevitable, but there's a lot of institutional inertia.
https://github.com/cncf/cnf-conformance/
Its been a dream.
You can pretty much write the ruby you are used to but it performs like a compiled language and you can make binaries.
Was SUPER skeptical at first when the team decided on it but its been a pleasant surprise.
[1] https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
[2] https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
[3] https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
- Lua programs cannot use shared memory concurrency or subprocesses with 2-way communication with the master process.
- Lua programs run on a very slow runtime compared to the fastest Lua runtime.
My impression after this is that for languages that aren't super fast and don't include all the primitives one could want, benchmarks like reverse-compliment are mainly measuring whether the language's standard library includes some C function that does the bulk of the work.
To a great extent any "language" benchmark (for languages that don't compile to efficient machine code) is certainly a benchmark of the language's standard library. I'm not sure there's a way around that reality. Are there external Lua libraries that allow shared-memory concurrency? If so, it's probably worth opening an issue asking whether those libraries could be allowed[2]; it might just be that nobody has submitted a program making use of Lua shared concurrency.
[1] https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
[2] https://salsa.debian.org/benchmarksgame-team/benchmarksgame/...
Isn't that the same situation as Python, Perl, PHP, Ruby… except for those languages, programmers have converted the programs to use multicore ?
This means that the proposed work sharing model at https://benchmarksgame-team.pages.debian.net/benchmarksgame/... cannot be used, because it requires workers to submit results to be aggregated and accept new chunks of work.
The pfannkuchen-redux is just a bit hampered by uneven work sharing.
For reverse-compliment, it's a bit more trouble to work around the lack of 2-way communication. My implementation writes the entire input to stdout, then workers use fseek on stdout, which only works if you are piping the output of the command to a file. That is, it generates correct output if you run "lua blah.lua > out" but not if you run "lua blah.lua | cat > out" Additionally, since there's no pwrite and no way of getting a new open file description for stdout, I must cobble together a mutual exclusion mechanism to prevent workers from seeking while another worker tries to write.
I think people don't usually write Lua programs intending to run them inside the binary you get when you build PUC-Rio Lua without any additional C libraries. Libraries like LPeg and lua-gumbo are Lua wrappers around C code. For C libraries that do not have Lua wrappers, people can more or less paste the preprocessed C header file into their Lua source file and use Luajit's FFI to use the library. This last approach is similar to how the Python regex program mentioned elsewhere in these comments works :). It's also common to use frameworks like Openresty or Love2d that provide the innards of some complex threaded program to user Lua code.
Outside of benchmark games and work, I'm working on some code that uses threads and channels, but the threads and channels are provided by liblove.
So I guess I can say, it has been addressed, but it won't be addressed in the standard library.
I hope you're doing that because you think it's fun.
Take a look at the source code for the top Python program for regex-redux.
But yeah, just wanted to make sure people are actually looking at the code. The top submission for Python, for example, is not how I've ever seen anyone use regexes in Python. That's an important dimension to evaluate in these discussions! (But not the only one, of course.)
> Rust is very cool in that one for actually using a regex engine implemented in Rust
One wonders how long it will be until someone submits a Rust program that uses PCRE2.
Yeah, like "Always look at the source code."
> … not how I've ever seen anyone use regexes in Python.
More like this…?
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
Yes. Not sure what your point is though. The GP wasn't talking about that submission. ;-)
Plus measures of the other programs, for all non-C/C++/Rust languages, that don't obviously use FFI.
I pulled the code down, placed the data on a ramdisk so as to not impact benchmark measurements, with physical issues.
I built the C code, and ran the two julia codes. My timing looked like this:
version t(raw) t(gz)
c1 1.47s 8.31s
jl1 3.80s 15.82s
jl2x 5.85s 17.86s
py1 fails
py2 6.92s 29.62s
I don't have lua, nim, or crystal on my machine. This is Julia 1.4.1 BTW. Running Linux Mint 19.3 on my laptop, 5.3.0-51-generic kernel.
Beyond putting the data and compressing it with pigz for the compressed version, no optimizations were done. Putting the data on ramdisk optimizes all codes.
My thoughts: The author noted that Julia has long startup/run times. This is true, for the first compilation of modules you use. As a reflex these days I (and I am guessing most Julia users) do a "using $MODULE" after adding it. This makes the startup times less painful, for most modules used. Plotting, with the Plots module, is still a problem, though it has gotten dramatically better over time.
Basically, if you run your code more than once, with the modules being compiled into your cache, the nice part is that startup time is significantly better. Such that Python reverts to its lower performance than Julia. If the startup time on first run is important (think of it like a PyPy compilation step along with a run of the code), and you'll only ever run a code once, and for the less than 1/2 minute that it will take for this example, use whatever it is you are comfortable with.
FWIW, the author noted, with implied disdain, that Julia users are telling them that they are "holding the phone wrong." Looking over the code, specifically all the memory allocation bits, I could see that. Basically, I'm not sure how much, if any of that, is actually needed.
That said, very limited critique of the tests. I like to see "real world" examples of use. Kudos to the author for sharing!
[edited to "fix" table ... not sure how to do real tables here]
I think the real challenge for scientific computing (I'm a graduate student, so this is most of the programming I do) is that there is already a huge network effect around NumPy + SciPy + matplotlib and friends. Golang just doesn't quite have the library ecosystem yet, although gonum[0] shows some potential.
In my limited experience so far, I think Go is good in an environment where most people are going to have experience with C and/or Python. It also makes it much harder to write truly crappy code, and it's much easier to get a Go codebase to build on n different people's workstations than C.
Having written a lot of Python, and relatively little Go, I think I would prefer to write scientific code in Go if the libraries are available for whatever I'm trying to do.
It's also much easier to integrate Go and C code, compared to integrating C and Python.
Finally no other language except for maybe FORTRAN has seamless parallelisation support and first class low level numerical primitives developed by vendors. Sometimes you will get a massive performance increase by #pragma omp parallel for.
Even for visualization some python libraries will suddenly fall off a cliff (Altair) once you reach a moderately large number of datapoints.
For big numerical stuff and things that need to run on supercomputers, C/C++/FORTRAN are definitely very relevant and I don't see that changing. Likewise for edge stuff that has to run on bare metal or embedded, I think we're still going to be using C/C++ for a long time to come.
"Scientific computing" is a huge range of different use cases with very different levels of numerical intensity and amounts of data. I doubt very much that there would ever be a one-size-fits-all approach.
However in the context of the OP, I'm arguing that Go would be preferable to Python for the purpose of writing bioinformatics models, and certainly more suitable than Lua or JavaScript.
Of course Python can sometimes be very performant if you leverage NumPy/SciPy, since those are ultimately bindings into the FORTRAN numeric computing libraries of yore. But if we're talking about writing the inner loop, and the choices are Go, Python, Lua, and JavaScript, I think Go is going to win that on the performance and interoperability fronts handily (I omit Crystal, as I am not familiar with it).
I do still think it would interesting to see a comparative benchmark though. I know the Go compiler tries to use AVX and friends where available. I doubt it will ever beat a competent programmer using OpenMP to vectorize though (though Goroutnes might be competitive for plain multithreading).
A relevant consideration too -- OpenMP seems to be moving in the direction of supporting various kinds of accelerators in addition to CPUs*, so your C++ code has a better change of being performance-portable to an accelerator if you need it.
I think for rust, the barrier of entry is high in that it is a difficult language to learn. Admittedly, I don't know rust, but people I know who do have said as much.
I think Go strikes a good balance of being easy to learn and use, having a rich standard library, and also being performant.
I have sat down to learn Rust several times investing many hours and still wouldn’t say I’m in any way competent at it.
One cannot even write a generic `reduce` function...
For the uncompressed FASTQ benchmark, my relative timings differ from his, with Julia being ~30% faster. It's probably because he installed an old version of FASTX.jl, that would also explain his seemingly outdated comment about the FASTX.jl source code. Another 30% time can be shaven off by disabling bounds checks, which would put Julia as the fastest program on the list, around the speed of C. However, I don't think the speed increase is worth it, since FASTQ files are basically aæways compressed in real life.
For the compressed FASTQ file, it seems this is entirely explained by CodecZlib.jl being 2x slower than whatever his C solution is using. However, when profiling CodecZlib.jl, it just spends all its time calling `zlib` written in C (calling C has near-zero overhead in Julia). So I have no idea why his benchmark is slower there.
When using SBCL, for example, you get your application to be compiled to native instruction set. Moreover, you get control over optimization level for each piece of the code separately. Even more, you get complete control over the resulting native assembly. Something you don't get to do with other higher level languages.
One of the production applications I did in Common Lisp was parsing a stream of XDP messages (https://www.nyse.com/publicdocs/nyse/data/XDP_Common_Client_...) with a requirement for very low latency. I made the parser in Common Lisp so that it generates optimal binary code from XML specification of message types, fields, field types, etc. using a bunch of macros.
The goal of application was to proxy messages to the actual client of the stream. The proxy was there so that it was possible to introduce changes to the stream in real time without requiring to restart any components. Using REPL I was able to "deploy" any arbitrary transformation on the messages. The actual client of the messages was a black box application that we had no control over but we sometimes had problems with when it received something it did not like.
I liked Common Lisp in particular because it does not force you to make your performance decision upfront. You can develop your application using very high level constructs and then you get the option to focus on parts that are critical for your performance. Macros allow you to present DSL to your application but then have full control over code that is actually working beneath the DSL.
If everything fails, calling C code is a breeze in Common Lisp compared to other languages.
I would not use Common Lisp on the critical path because I would end up basically rewriting everything to ensure I have control over what is happening so that some kind of lazy logic does not suddenly interrupt the flow to do some lazy thing.
A large part of the application was basically about controlling the memory layout, cache usage, messaging between different cores, ensuring branch predictor is happy, etc which would be really awkward in Common Lisp (technically possible, practically you would have to redo almost everything). We have also experimented with Java with the end result being that the code looked like C but was much more awkward.
I have, however, successfully used Common Lisp to build the higher layer of the application that was orchestrating a bunch of compiled C code and also did things like optimizing and compiling decision trees to machine code or giving us REPL to interact with the application during trading session.
This to me is what the big appeal of Common Lisp for a trading system could be, particularly if it allowed live recovery from errors (dropping somebody into the debugger, rather than just core-dumping), which could save a lot of money by reducing downtime. But as you say it would require redoing everything to make the code fit latency constraints and be cache friendly, which would be a lot of work.
There's even a Rails-like [0] framework being developed!
Last weekend I spend a couple of hours getting Lucky up and running (it took some doing, I had to borrow a lot from people's Docker images to get it booted and working).
It's a big 'watch this space' situation. The macro system for metaprogramming is very easy to understand in Crystal [1] as well.
Can't wait for 1.0!
[0] https://luckyframework.org/ [1] https://crystal-lang.org/
The author might like to look at J for specific calculations (and Futhark for similar tasks).
I use Haskell, it's not quite on par with C though. Other MLs are wonderful too.
Perhaps a well-written bio informatics library would be a nice solution.
I would make a pull request, but their setup uses 1TB of RAM...
Anyway - good job!
I think you can also have VS Code just load up the packages you expect to use on a regular basis at startup.
but I agree. If you complain that compile takes 11 seconds when the program runs in 30, then I wonder if that's really the use-case where you need every last bit of performance.
Now, on a program that runs for two days straight in Python or Matlab, and Julia reduces that time by half, I can deal with a bit of compile time.
> I am equally new to Julia, Nim and Crystal.
He might be more familiar with a certain style or paradigm, but he wants to switch from Python to something significantly faster jumping through as little hoops as possible.
I don't think every article discussing benchmarks has to restate that the differences between programming languages are not just syntactical to be informative.
I wonder why BioPerl (https://bioperl.org/, https://en.wikipedia.org/wiki/BioPerl) has not been included, it would have been interesting. Perl is considered surprisingly fast for certain classes of tasks (for an interpreted language of course, no point in comparing to C for example).
My info could be outdated of course.
As for performance, it really helps if you write idiomatic Perl code. There may be more than one way to do it, but some ways are better. For example, an explicit for loop over a list has to be compiled to byte code that steps through the looping, but if you use map or grep the byte code calls a pre-written optimized function that loops as fast as code written in C would be. The more idiomatic your code is, the more optimized calls like that you get, and the faster your code will run.
Same is true for both CPython and certain sections of R code.
We need better profilers! Ones that don't require anything more complicated to use than passing a --profile-me parameter to the interpreter/compiler, and whose output can just be dragged into a pretty, user-friendly and fast application (included with the language runtime), and where reported results are both trustworthy and correspond to actual locations in the source code.
Jython does not have a GIL.
Like, one can certainly write a slow C... but writing a fast Python (real Python, I mean, with full reflection, run-time introspection, etc..) would be...near impossible.
Yes, I know there are projects that make large subsets of Python run fast, but it's that last 5% that kills you.
to see the problem with Python, take this loop as an example:
x = 0
for i in range(0,1000):
x = x + i
There is no way, looking only at this part of the code to know if the `+` operator does the same operation every time it is executed.If we're talking about custom classes and not ints, maybe it's a bigger problem. But if PyPy doesn't allow the required introspection to make this work, how does it run anything at all?
Most languages have a single implementation and definition of the language and especially its ecosystem is married to it. Some have more, but usually an ecosystem is popular for only one of them, because they are not entirely compatible.
Nowadays by a language people usually mean the whole package.
Types are child’s play for the crowd using Matlab and worrying about low-level performance like this though.
There is always a need for scripting languages. Not just for speed of development, but for interactive data manipulation and visualization. Static languages are a no-go in that regard.
My bet in on Julia. Although slower than C by itself, I think in practice, Julia code written by bioinformaticians will be faster than C code wrapped in Python, or C libraries used in inefficient workflows from the shell. That's certianly been my experience so far.
Why wasn't Go, Haskell, Java or OCaml even considered here? Seems much more appropriate than the likes of Javascript for scientific computations.
Modern Fortran would be a good option; 'right' is a different matter. Certainly, Fortran 2003/2008/2018 are expressive languages (way more than F77), has an extensive ecosystem with good integration with C, has compilers that generate very fast code, handles parallel computations and SIMD and first-class language features, is already well established in the STEM world, etc., etc.
The biggest downside of Fortran is it's called 'Fortran' and so many people are unwilling to believe it's changed since the 70s.
There are quite a number of books. I recently got this one that is pretty good: https://www.amazon.com/Modern-Fortran-Explained-Incorporatin...
One note, Fortran 2018 compliance in the compilers is still evolving, so not all features are in all compilers yet. Fortran 95/2003 support should be solid and most all have all of 2008 in.
brew install crystalOf course the build-everything-from-scratch gentoo linux crowd is going to have a harder time but isn't that part of the masochistic appeal?
And that's beside the fact that 1) outside of your bubble, more devs use Windows than any other OS, 2) the person who wrote this article isn't even a software engineer, and 3) the tests weren't even run on macOS.
C/C++/C#/D, etc are not high level by this criteria.
There are also of course degrees to all this. Rust is lower level than a lot of languages by virtue of having native pointer uses, even if it's frowned on.
There's also something to be said for language features, but I can't quite put my finger on it.
I think a good real-world example that exposes the problem in your definition is Go. Go has pointers, using them is normal. However, go does not pointer arithmetic and outside of unsafe (like Rust) they are memory safe. I consider golang to be higher level than C/C++ for this reason, and many others (GC, channels, defer, etc, etc) -- I'd also consider it lower-level because of its non-answer/cop-out to error-handling.
But what is special about pointers compared to references? If you have a language with pointers that are type-safe and memory-safe how is this distinctive?
The difference is one of model. Pointers are an exposure of the underlying computer architecture. Whereas references are more of a property of common language design. In theory, you could not have pointers but still have references.
Sorry, I just don't follow. How are the pointers in Go more exposing of the underlying architecture than a reference? (I'm using Go as an example to make it concrete, but any language with similar properties will do).
The syntax and some of the semantics of assignment and rebinding are different between say go pointers and python references, but that's the point I'm contesting, I don't see how one is necessarily higher level than the other. If you put automatic memory management, null pointer checks, removal of any "undefined behavior", pointers aren't necessarily low-level. It wasn't the pointer, it was the memory safety.
I personally think that once you tease it out that it becomes a semantics argument that unfortunately doesn't shed much light on what is "high-level".
Fwiw rust has completely replaces cpp for high performant bioinfx code in my workflows. Sooo much easier to write than cpp!
On the other hand, there are times when you really need to get into the nitty gritty of things for performance reasons, and it ends up feeling quite low level.
For example, here's a Rust program which does the same as fqcnt_py1_4l.py: https://gist.github.com/Measter/d31abe88b5e318ba98856bf6f047...
It's longer than the Python version[0], I'm going out of my way to not allocate for every string. But, and maybe I'm biased here, it doesn't really feel like it's low-level code.
[0] https://github.com/lh3/biofast/blob/master/fqcnt/fqcnt_py1_4...
Comparing scripting languages to compiled languages is apples and oranges. The advantage of scripting is supposed to be the convenience - not speed.
how times have changed
The other problem is that level doesn't really mean anything, or it means something different in each language. Is it feature count? Abstraction potential? Memory-addressed vs objects? Statically or dynamically typed? Closeness of fit to the machine it's running on (imagine a lisp on a lisp machine, or x86 on an emulator - is that high or low level?)? Etc.
That's how i'd have it if it were for me to decide, but this isn't how language evolves.
We don't need a new set of terms. Relative terms are fine, they're just context sensitive.
C (especially C99 and up) is far, far higher level than something like FORTH.
Rust is cool, but it does not have a horse in this race I'm afraid.
The choice of languages for this benchmark is laughable.
Results in clear code, easy implementation, and good performance 95% of the time.
Most biologists I met were far more comfortable in the pissing rain, up the mountain, in the middle of nowhere collecting animal shit than in front of a keyboard.
If there were a Venn diagram of biologists and computer and technology enthusiasts the overlap would need a micrometer to be read.
Disclaimer: Please take this extremely generalized and likely offensive to one of those small few in that tiny overlap I mentioned, statement with a grain of salt. Please don't take this too seriously, it's just from my own narrow sampling of people i've interacted with which is may or may not be representative of the overall population.
Disclaimer: my background has absolutely no overlap with biology/bioinformatics.
AFAIK a lot of chemistry equipment actually is built with Java and C#.
In fact you might encounter some biologists using VB.NET after they outgrown their Excel VBA macros.
I think that for most academic uses, ease of development is really important. Remember, most of these people don't have CS backgrounds, which makes python a good choice because it's easy to learn and doesn't require much setup. It's main drawback is that it's slow, but this doesn't matter in many cases. When I was in grad school, I'd use a small dataset to try to figure out the best algorithm and then if I had to leave my computer running overnight to run on a larger dataset, it usually wasn't a big deal.
On the other hand, sometimes you just need something really fast. In that case, C is your darling.
Java/C# make lots of tradeoffs on both sides, which works well for software developers but not necessarily researchers (who don't have to worry about code being maintainable).