Is there any other downside? Electricity consumption / heat? Evicting other stuff from cache?
Is there any other downside? Electricity consumption / heat? Evicting other stuff from cache?
6 cycles for bounds checks that the branch predictor never had to rewind on is nothing in comparison to saving a couple trips to L3.
Dramatically better L1 and L2 cache behavior. It seems clear that the additional instruction load of the Rust driver is partially made up by the excellent cache utilization.
This "Rust vs C" document is just one part of a larger analysis of network driver implementations in many languages; C, Rust, Go, C#, Java, OCaml, Haskell, Swift, Javascript and Python. Have a look at the top level README.md of that GitHub repo.
Also, isn't it unintuitive that branch mispredictions go up with larger batch sizes? Wouldn't there be fewer branches per unit time?
[1] https://github.com/duneroadrunner/SaferCPlusPlus/blob/master...
It's mostly vmem until / unless the data actually gets used though, no?
The same way it does for every other bit of allocated memory: it allocates the physical page on a page fault in a valid mapping.
The effects of larger memory usage don't become obvious until other applications start contending for it and/or swapping happens, and it's conveniently also something that is not as easily blamed on one application "being slow", which is why it doesn't receive nearly as much attention as it should.