What impresses me is that the Rust version didn't do any of that stuff, just wrote very boring, straightforward code -- and got the same speed anyway. Some impressive compilation there!
What impresses me is that the Rust version didn't do any of that stuff, just wrote very boring, straightforward code -- and got the same speed anyway. Some impressive compilation there!
It is a great example of how the Rust compiler can auto-vectorize code.
Are you sure the rust performance data isn't for one of the other implementations that use the same crufty tricks as the C version?
e.g.: https://benchmarksgame-team.pages.debian.net/benchmarksgame/..., https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
It looks like the autovectorizer did a really good job on this one.
As an aside: this made me notice this n-body simulation only has 5 bodies! This is a pretty strange case that makes these O(n^2) optimizations practical. The n-body simulations I've been familiar with in the past had thousands of bodies, where this approach probably isn't a good idea.
This blog post shows how to write simple idiomatic Rust code that will allow the compiler to auto-vectorize:
So maybe it's just the C benchmark being a cargo cult.