Evaluation of C, Go, and Rust in the HPC environment [pdf]
octarineparrot.com
octarineparrot.com
He rewrites a distributed program as a threaded one, because Rust and Go didn't have a distributed computing library, but says Go and Rust have "similar performance" in HPC.
He says he was surprised at the C performance, but there's no investigation of a cause. Maybe just because C is slow...
I don't mean to be overly critical, but these results in the paper mean absolutely nothing. It's a waste of time.
I do agree that the advisor should have pushed back lots harder on the methods; it's a weak paper, even as an experience report.
I agree - this seems to be mostly the advisors fault, by the looks of it it's not that the student didn't put a lot of work into it but methodological problems are what advisors are supposed to be there to address; I saw a lot of this back in college, at least with technical papers it's easy to tell what's wrong, with the humanities it's just a complete disaster.
http://www.erlang.org/doc/tutorial/introduction.html http://www.erlang.org/doc/tutorial/nif.html#id64792
The advisors should have encouraged a smaller and more manageable scope with a more thorough methodology where the results would actually be worth something.
You mean it has nothing to do with learning correct methodology or actually doing real research?
That's exactly the attitude emblematic of the portion of academics producing worthless output.
Here is a recent reference: http://www.nature.com/news/psychology-journal-bans-p-values-...
Not a methodological problem?
The article submitted was not peer reviewed, and could very well be bad science, but I think accusations of non-science are not very productive and do not provide feedback in how to do better.
My point was rather that there was so much low quality research they decided to try to filter it by banning a statistical tool. Are you saying the reasons for so much low quality research had nothing to do with methodological problems? Maybe it's not low-quality, since you know it's "subjective"...
> The article submitted was not peer reviewed, and could very well be bad science, but I think accusations of non-science are not very productive and do not provide feedback in how to do better.
You are addressing a point I never made. What I said was(quote): The advisors should have encouraged a smaller and more manageable scope with a more thorough methodology
Because an important part of undergraduate education is in fact learning to be able to tell what good research is. The advisor's aren't really benefiting them by not pointing out methodological problems and letting them do research ">>>>> wherever that may lead them" - because that only encourages methodologically sloppy work in the future.
As briefly mentioned in the previous chapter this performance regression might have been caused by the two unoptimized libraries that were compiled on the development laptop and copied to the cluster.
You are correct this was definitely an oversight and I would have redone the measurements if time allowed for it. However I just want to state that the libaries were built with optimizations just not on the target platform which might not have been clear from the qoute.
Your thesis is in fact really impressive and very nicely written! Your conclusions and future work section is excellent. Really very well done. Congratulations on finishing it. I hope you don't let the irrationally negative feedback here dissuade you from continuing your work, you certainly have a bright future ahead of you.
Most interesting part of the thesis, for me, is page 42:
> In this case it was even an advantage that the Rust version was developed last since it revealed a critical error in the other implementations.
Null strikes again!
n1_idx, ok1 := g.nodeIdx[n1]
n2_idx, ok2 := g.nodeIdx[n2]
if !ok1 || !ok2 {
When you do not use the comma-ok form, you are specifically asking the map to return the zero value for non initialized elements (which is something fine to do in many cases).Also, it doesn't restrict the use of C libraries, so even if you've been writing C for HPC code all these years, there should be minimal technical difficulties in switching over, except for the time it takes to learn C++, which shouldn't be long if you know C well.
The programmer is by far the slowest part of HPC. It often makes sense to use the language you're fastest at instead of the language that will run the fastest.
It's very common for experiments to be written in languages not know for their performance (Java, Matlab, Perl, and Python for instance).
For reference, here are some of the packages commonly used in experiments[1]:
BLAS Fortran http://www.netlib.org/blas/#_software
NCBI BLAST C++ ftp://ftp.ncbi.nlm.nih.gov/blast/executables/blast+/LATEST/ncbi-blast-2.2.30+-src.zip
BFAST C http://sourceforge.net/projects/bfast/
BioPerl Perl https://github.com/bioperl/bioperl-live
Bowtie C++, C http://sourceforge.net/projects/bowtie-bio/files/bowtie/
Clustal C, C++ http://www.clustal.org/omega/#Download
cp2k Fortran https://github.com/cp2k/cp2k
Gromacs C https://github.com/gromacs/gromacs
HTSeq Python https://pypi.python.org/pypi/HTSeq
MUSCLE C++ http://www.drive5.com/muscle/downloads.htm
MrBayes C http://sourceforge.net/p/mrbayes/code/HEAD/tree/
OpenFOAM C++ https://github.com/OpenFOAM/OpenFOAM-2.3.x
SAMtools C https://github.com/samtools/samtools
SNAP C http://korflab.ucdavis.edu/software.html
fftw3 C https://github.com/FFTW/fftw3
[1]: I went down this list and picked some of the ones I remember using (been out of the HPC world for a couple years): https://portal.tacc.utexas.edu/software It'll be a bit skewed towards genomics.FFTW is C code, but nobody wrote the C code. The C code is generated by code written in OCaml.
As an example, I've seen physicists use fortran for parsing and transforming the output of a program. The cluster at our department has an awesomely updated ifort but no python2.7 (I think it has 2.2 or 2.3 or something). C++ isn't used directly either.
But yeah, there should have been fortran and some other languages there.
My experience is limited to one single field, but all weather prediction models and climate models that I know of, are in Fortran. And I would count them as big HPC projects.
Trilinos is another C++ example, while PETSc and HYPRE are written in C.
My tentative conclusion from this is that newer projects tend to use C++ and slightly older projects use C.
Yes, there are things like ODEPACK, QUADPACK, and FFTPACK, but those are not under development anymore (as far as I know). The only widely used Fortran library still under development I can think of is LAPACK.
I did not count the occasional 'bespoke' code. There are still some Fortran applications in development for particular purposes (like MOM), but those are harder to survey.
Most users are usually fine with whatever the GUI allows and a bit of Python, but for those of us developing new models, Fortran is probably the most useful language, followed by C++.
A quantum Monte Carlo code will of course include a model, but I think people don't want to call it "software" because it's so research-grade and janky. "Program" seems better, but I think that implies that it's a static thing (not in a constant state of development).
The plural "codes" is used because usually a research team has historically implemented a bunch of models into disparate codebases.
After reading the comments on C and the difficulties the author had in the compilation process.... All I can say is Garbage in garbage out. I have worked in the HPC enviornment for a number of years using both Fortrash and C. The errors the author encountered would have been avoided by someone with any experience developing large softwares. I learned these things after working on my first project and now they barely, if at all, factor into my development time. Make sure you build and keep a dependency graph while you are developing and this becomes a non issue.
Fortrash
Why the hate? One could misuse any tool, but that doesn't make the tool bad at what it does.I kind of wish I could find a good way to make Fortran and C play nicely together. (There are some compiler dependent peculiarities that I have encountered that make it difficult.)
Fortran's intrisics and file handling capabilities make it indispensable in the work I do.... So it is really a love hate relationship.
Fortran can call C and vice-versa, at least since Fortran 2003 there is a standard way to call C.
edit: Nevermind I found a few. When I looked for this a few years ago I just kept hitting dead ends.
It has some pretty basic flaws that are pervasive that the reader should be aware of. First off is the "productivity" measure: the last time I looked around, this was nearly impossible to quantify for software development. The author chose SLOC and development time as stand-ins for productivity. SLOC has a connection with code quality and time of development, but development time is well known to vary[1].
In particular, development time in this thesis is linked to results in a pretty classic "Psychologist's fallacy"[2]; the author generalizes their experiences as conclusive. This implies in particular that "development time" in this thesis should be thrown out for meaningful conclusions, as the sample size is 1. It is, however, an interesting experience report.
Another basic flaw is the connecting of 'modern' with 'good'. The author remarks the main disadvantage of C and Fortran are their age; this crops up here and there. Workflow tooling in Go and Rust are major focuses by the developers: C workflow tooling is usually locally brewed. This does not make C-the-language worse.
More on a meta level, my advisor in my Master's work drummed into me that I should NOT insert my opinion into the thesis until the conclusion. So I found the editorializing along the way very annoying.
And finally, and very unfortunately, the experience level of the author appears to be low in all three languages; this is significant when it comes to implementing high performance code.
---
Now, for the interesting / good parts of this experience report.
Standout interesting for me was the Rust speedup on the 48-core machine. I did not expect that, nor did I expect Go to make such a good showing here as well.
In general Go performance/memory made a surprisingly good showing (to me) for this work. I shall have to revise my opinion of it upward in the performance axis.
I am both surprised and vaguely annoyed by the Rust runtime eating so much memory (Would love to hear from a Rust contributor why that is and what's being scheduled to be done about it).
Of course the stronger type system of Rust catching an error that C and Go didn't pick up is both (a) humorous and (b) justifies the type community's work in these areas. I look forward to Rust 1.0!
One note by the author that is worth calling out stronger is the deployment story: Go has a great one with static linking, whereas C gets sketchy and Rust is... ??. I believe Rust has a static linker option, but I havn't perused the manual in that area for some time. For serious cloud-level deployments over time, static linking is very nice, and I'm not surprised Google went that route. It's something that would be very nice to put as a Rust emission option "--crate-type staticbin".
Anyway. I look forward to larger sample sizes and, one day, a better productiv
[1] http://www.amazon.com/Making-Software-Really-Works-Believe/d... The situation is actually far worse than just varying developer time.
I haven't spent time with the code yet, so I can't comment. It shouldn't be the runtime, though, as our runtime is about as big as C or C++'s. My first thought would be that Vec's growth factor may be poor in this instance, but without reading the code, who knows.
> Rust is... ??
Rust statically links everything but glibc by default. Experimental support for musl was added last week.
Rust's runtime is here: https://github.com/rust-lang/rust/tree/master/src/rt and https://github.com/rust-lang/rust/tree/master/src/libstd/rt
As you can see, it's very, very small. It mostly handles things like unwinding, at_exit handlers, and the like.