(Rust tends to use closures heavily, in places other languages don't. So they need all the performance they can get.)
(Rust tends to use closures heavily, in places other languages don't. So they need all the performance they can get.)
In most of those cases it's due to either wanting to run in very constrained space or distribution size, and then the problem is due to some blanket impl for a bunch of types that can safely be removed though. There are a few cargo submodules that will do executable size analysis with regard to monomorphization to show the outliers.
There's not really a good solution for this. Having your code be fast by default in exchange for more space seems like the trivially correct choice to make when most people have hundreds of gigabytes of free hard drive space and binaries don't get above a few megas, while CPU speed has stagnated. If they want runtime type matching instead, they can still do it.
However, in many cases, a well designed generic API will be able take a trait object in the place of `T`. That way, at the application layer, you can decide to pass trait objects around, reducing code size and - maybe more importantly to you - compile times. cargo does this for example in some places where the performance regression is insignificant but it saves noticeable compile time.
Sometimes libraries accidentally don't set things up so that you can pass either a concrete `T` or a dynamically dispatched trait object to their APIs; we're hoping to find a way of making it less likely for this to go wrong in the future.
Its also fairly common to have code like:
fn do_it(a: impl AsRef<i32>) {
let a = a.as_ref();
// do something with a
}
That could be instead monomorphised to be fn _do_it(a: &i32) {
// do something with a
}
fn do_it(a: impl AsRef<i32>) {
let a = a.as_ref();
_do_it(a)
}
saving the inlining of _do_it's code bloat/generation penalty.Perhaps Rust compiler could get feature flag for controlling optimization priorities as I suspect in webasm world one would more often prefer more reasonable network transfer sizes than the fastest running version that is possible.