Rust takes an approach with generic functions called monomorphization which means that we generate a new version of each function for each set of generics it's instantiated with. This means that a future of a String will generate entirely different code from a future of an integer. This allows generics to be a zero cost abstraction because code is optimized as if you had substituted all the generics by hand.
Putting all that together, highly generic programs will generally trend towards higher compile times. With all the generics in play, there tends to be a lot of monomorphization which causes quite a lot of LLVM IR to get generated.
As with many aspects of Rust, however, you have a choice! Rust supports what we call "trait objects" which is a way to take a future and put it behind an allocation with a vtable (virtual dispatch). This forces the compiler to generate code immediately when a trait object is created, rather than down the line when something is monomorphized.
Put another way, you've got control over compile times if you're using futures. If you're taking a future generically and that takes too long to compile, you can instead take a trait object (or quickly convert it to a trait object). This will help cut down on the amount of code getting monomorphized.
So in general futures shouldn't make compilation worse. You'll have a choice between performance (no boxes) and compile times (boxing) occasionally, but that's basically already the case of what happens in Rust today.
Unlike C++ templates, Rust enforces a single definition with uniform semantics, for a generic type/function (specialization going through the existing trait static dispatch mechanism), so we can take advantage of that to reduce compile times.
It's interesting that this is seen as significant here. Are we dealing with much shorter timescales, or just being eager to optimise everything?
Edit: spelling
Cost of a branch misprediction is 10s of cpu cycles. (1) Measured in gigahertz (10^9 cycles per second).
Time to turn around a web request is, if you're very lucky and have done the work, mainly about getting a value from an in-memory cache at multiple milliseconds (2). That's 1 / (10^3) seconds.
If you're not lucky, 10s or 100s of milliseconds to generate the response.
It seems that the second duration is best case around 10^6 times longer. I would not sweat the first one.
1) http://stackoverflow.com/a/289860/5599 2) http://synsem.com/MCD_Redis_EMS/
However, now that we have MIR we can optimize this on the Rust side first, before the type info has been lost. This was one of the reasons MIR was added :)
Maybe you mean LLVM is very low-level and thus it sees everything in the abstractions, which obviously takes time to sort through (and is especially wasteful to do multiple times for different monomorphisations). Thus, using MIR to optimise at a higher level can simplify code more efficiently before it hits LLVM.
Without knowing anything about LLVM's internals, I would assume it doesn't anticipate so much closure chaining, and therefore doesn't leverage the fact that they are so easy to inline.
Also, closures are just structs with functions that take them as arguments, so if the compiler can handle those, it can handle closures.