It’s a complex set of issues that we’re constantly working on.
(So yes, Go and C do not have generics, and so do not have this problem.)
It’s a complex set of issues that we’re constantly working on.
(So yes, Go and C do not have generics, and so do not have this problem.)
Would it make sense to have two types of generic implementation with a boxing unboxing mechanism. AKA a fast one that requires a rebuild when you monkey with something. And a slow one that doesn't but compiles fast.
I think it'd be acceptable if the 'dev' version used GC as long as the final build didn't. Basically anyway to get the development cycle to go fast.
fn<T: MyTrait> compile_time(thing: &T) {}
fn run_time(thing: &dyn MyTrait) {}
The reason that you can't just substitute them is that you can do more with compile-time generics than runtime ones -- runtime generics have unknown size (different concrete types have different size, so what size is the runtime generic version?) so they always have to be used through a pointer (and usually heap allocated). There are a variety of other restrictions on trait objects (runtime generics) for type-system reasons.It would be possible to find cases where compile time polymorphism is replaced with runtime polymorphism, but I'm not sure it would really gain much given the restrictions.
Now if you start changing the language semantics a lot more becomes possible, but I don't even know what changes you would have to make to the language to let that happen.
1. Maybe monomorphisation overduplicates code, so that different invocations produce the same machine code in the end?
2. If that's not happening, maybe idiomatic Rust code causes a lot of code to be generated where a C programmer would have written something differently, e.g. using a void* instead of passing by value?
3. If that's not happening either, maybe monomorphisation is more or less producing what a C programmer would have monomorphised manually, but it takes the compiler a long time to optimise away abstractions?
It seems that in this C++ example the problem is #3: it takes a long time to optimise away the range & lambda abstraction.
If so, is Rust worse than C++ when it comes to compile times of those? Heavily templated C++ code is known to be slow to compile due to this. Is it much worse in Rust? Are you aware whether C++ compilers do some special optimizations that Rust compiler doesn't have yet?
I don't see what can be done to make this better, the whole point of templates is to "pay" at compile time for some gains at runtime. (or in this case for gains in the amount of code a developer has to write)
There’s at least one thing that could be done, in Rust at least. But it’s only in some limited circumstances. See here: http://www.suspectsemantics.com/blog/2016/12/03/monomorphiza...
Monomorphization seems to be the source of 2 major problems: compile time, and binary size.
An optimization that could reduce the compile time by caching the generated code (at least when compiling over-and-over the same code base. i.e. "Incremental compilation") - and it seems that Rust already is doing something with it [1]. I wonder if something specific is done for template instantiations there.
Optimizations aimed at reducing binary size seems much more tricky, if not impossible, except as you pointed out in the limited cases described above. In the regular cases, templates are kind of working as intended: the developer has to think if he/she would have written the same amount of code N times if templates were non existent.
This is partly why there are hacks like unity builds (not related to the engine, where all source is bundled into a single translation unit). These have plenty of drawbacks too so it's not a clear win.
Adding to all of this, there are fancier mechanisms for template meta-programming like SFINAE rules + computed template values. Sure, it's "turing complete" but this is why we see such clever libraries with huge explosions in code generation size. I'm far from an expert in modern C++ features but it's clear that there is an entire interpreted programming language of templates bolted on the to the rest of C++. It reminds me of this post on a similar take on Haskell type level programming: https://aphyr.com/posts/342-typing-the-technical-interview (or similar feats by Oleg Kiselyov).
pub fn big_function<T: Into<i32>>(x: T) {
let y: i32 = x.into();
... code that uses y but not x ....
}
In the first line of the function you have the environment {T: Into<i32>, x: T}. After the let you have the environment {T: Into<i32>, x: T, y: i32}. If you keyed the cache on the full type environment you wouldn't solve the issue because the code that only uses y would still get specialised to the {T: Into<i32>, x: T} too. However, you could detect that that code doesn't actually use T and x, so that you can specialise it to the type environment {y: i32} only.That detection can happen as a side effect of specialising that code to some particular {T: Into<i32>, x: T, y: i32} for the first time. As you specialise that code you record which parts of the type environment actually got used, and you use only those parts as a key in the cache. The type environment object itself could take care of recording what the compiler looked up in it. Another advantage of doing it this way, rather than analysing the code ahead of time, is that it can handle cases where a particular type variable does or doesn't get used depending on what type some other type variable is instantiated to. You could also use the same system to avoid duplicating code that only relies on particular aspects of a type. For instance, a function that permutes the values in a &mut[T] might only care about the size of T and not about the precise type T, so that all its specialisations to T of 4 byte size can call into the same code.
Another thing you'd probably want to do is integrate a basic form of constant propagation & dead code elimination during monomorphisation, so that you don't spend a lot of time monomorphising code that ends up dead for particular type instantiations.
Yes, that's what I'm saying.
> Another thing you'd probably want to do
This is a very interesting idea!