Outperforming Rust DNA sequence parsing benchmarks by 50% with Mojo
modular.com
modular.com
e.g. these might be interesting:
https://www.modular.com/blog/mojo-llvm-2023 https://www.modular.com/blog/what-is-loop-unrolling-how-you-...
If you still have doubts, you could join the 20,000+ people in discord chatting about Mojo stuff: https://discord.com/invite/modular
-Chris
> Insightful Reddit comment https://old.reddit.com/r/rust/comments/1al8cuc/modular_commu...
> > The TL;DR is that the Mojo implementation is fast because it essentially memchrs four times per read to find a newline, without any kind of validation or further checking. The memchr is manually implemented by loading a SIMD vector, and comparing it to 0x0a, and continuing if the result is all zeros. This is not a serious FASTQ parser. It cuts so many corners that it doesn't really make it comparable to other parsers (although I'm not crazy about Needletails somewhat similar approach either).
> > I implemented the same algorithm in < 100 lines of Julia and were >60% faster than the provided needletail benchmark, beating Mojo. I'm confident it could be done in Rust, too.
https://github.com/biojava/biojava/tree/master/biojava-genom...
Disclaimer: MojoFastTrim is for demonstration purposes only and shouldn't be used as part of bioinformatic pipelines
> The TL;DR is that the Mojo implementation is fast because it essentially memchrs four times per read to find a newline, without any kind of validation or further checking. The memchr is manually implemented by loading a SIMD vector, and comparing it to 0x0a, and continuing if the result is all zeros. This is not a serious FASTQ parser. It cuts so many corners that it doesn't really make it comparable to other parsers (although I'm not crazy about Needletails somewhat similar approach either).
> I implemented the same algorithm in < 100 lines of Julia and were >60% faster than the provided needletail benchmark, beating Mojo. I'm confident it could be done in Rust, too.
In between we have something we're cooking that I think will be pretty interesting for GPU kernel authors, but it isn't public yet. :-)
The nice thing about this is that it is one system that scales, instead of a bunch of different/inconsistent tech built by different teams over many years, held together with duct tape. Simple and consistent makes it much easier to do the kinds of research and experimentation that power AI ecosystem.
Speaking as someone who has spent more than 20 years writing compilers (e.g. LLVM, MLIR, etc), my opinion is that autovectorizers are a class of tech that are best applied to get speedups on legacy code bases. If you care about performance a lot, you shouldn't use them IMO - they are unpredictable and have performance cliffs.
-Chris
Regardless, the code in this post seems to use the same input in both versions and the change looks loop local at first blush, but maybe I'm missing something.
I am looking for a language not called C++ that I can write some bottlenecked modules in and wrap it around in python.
I want the code base to feel like python with a few files offloaded to this language.
Any suggestions ?