This reminds me I haven't written up stuff for chameneos-redux, which is probably the most fun and over-engineered of them all.
The fasta/Haskell thing I was mentioning are the parts at "The Haskell code does a quite clever alternative" and "the Haskell code just caches them all in an array".
The pre-compacting I mention in k_nucleotide was already done in the Scala implementation.
How parallelism is added also varies. The Scala code does each of the test cases in parallel whilst the Rust parallelizes each internally. (My new Rust code parallelizes even better.)
Several implementation make their own specialized hash table, and there isn't one particular hash table to implement. This is very important since hash table lookups take a substantial fraction of the time.
Note that I pick on Haskell and Scala here only because they have the best hacks, not because they're the only ones doing hacks. (I'd imagine listing all of the differences would take much longer than I have time for.)
thread-ring has lots of different implementations. Some are just scheduled coroutines restricted to a single process, then you have stuff like [1] which uses mutexes, real threads and sets thread affinity. I'm not totally sure what [2] is but I think it's something else entirely. I'm not being too specific here because honestly this benchmark confuses me.
[1] http://benchmarksgame.alioth.debian.org/u64q/program.php?tes...
[2] http://benchmarksgame.alioth.debian.org/u64/program.php?test...