How to speed up the Rust compiler some more
blog.mozilla.org
blog.mozilla.org
→ Systems programming newbie question: why are heap allocations bad for performance? Is it the additional level of indirection? The cost of calling your memory allocator? Something else?
My background, if that helps focusing answers: python/js programmer, did a tiny bit of C/C++, am ~approximately~ familiar with the stack (call frames, each with its context) vs. the heap (where to allocate memory for big/long-lived objects e.g. arrays and trees).
Sure, the allocator itself can be expensive, and that's certainly important if you're doing "too many" allocations, but in general I think it's worse caching that matters the most.
To other programmers not so familiar with CPU caching, resharing two blurbs that helped me get a start of a grasp on it:
- Why do CPUs have multiple cache levels? https://news.ycombinator.com/item?id=12245458 , https://fgiesen.wordpress.com/2016/08/07/why-do-cpus-have-mu...
- CppCon 2014 - Mike Acton: Data-oriented design and C++ , https://www.youtube.com/watch?v=rX0ItVEVjHc
It's really a shame more platforms don't show cache/icache misses in an easy to access manner.
why are heap allocations bad for performance?
1. You have to call to your allocator. Then do some type of search for free memory, of the approximate size. Then flag this memory as used. Then mark the remaining memory in that block as unused. Then mark the internal books to match the new layout.This is done efficiently, but modifying these collections take time. AND IT STILL FASTER THEN:
2. Calling the kernel to MMAP in new virtual memory, adding that to the pool, and well restarting this process all over again.
Allocation time is a big cost, and there is work to make allocation lazy by default in the Rust Compiler at the minute.
Also per thread pooling doesn't use more memory then not. In some cases it actually uses less. Citation: https://people.freebsd.org/~jasone/jemalloc/bsdcan2006/jemal...
http://www.dotnetcurry.com/csharp/1258/dotnet-platform-compi...
> Every time a developer changes a single character in any of the files, a new copy of all the data structures is created, leaving the previous version unchanged. This allows a high level of parallelism and concurrency in the Roslyn engine, as well as in its consumers, thereby preventing any race conditions to occur. Of course, in the interest of performance, these operations are highly optimized and reuse as much of the existing data structures as possible. Again, those being immutable makes this possible!
I first thought about this when working on UnderC, a hopelessly over-ambitious C++ interpreter. Functions could be recompiled, because there was an indirect reference to the actual code. And this is of course exactly how Lisp people used their compiler. (The 'image' reference of course is to Smalltalk)
[1] https://commandcenter.blogspot.co.za/2012/06/less-is-exponen...
The problem is, that usually this "keep state and just percolate changes" is easier said than done. But we're getting there.
See also:
https://www.youtube.com/watch?v=TS1lpKBMkgg#t=23m38s (scalac performance) https://www.youtube.com/watch?v=TS1lpKBMkgg#t=37m53s (what do we really need from "computing science" to do programming)
The symbol table, generated code etc. for every compilation unit is kept in memory and only discarded if the source is modified. All imports across compilation units are done via a double indirection, so the symbols can be unlinked and relinked more easily.
(Sub-second recompiles in Delphi are the norm, not the exception. In large projects, most of the time during dev builds is linking, and that's usually only a few seconds.)
Specially given Delphi features vs Go ones.
(I mention this in the context of my observation that the minimum bar to be taken seriously for a language is going steadily up. You certainly need a standard library that is powerful out of the gate, whether or not it is necessarily "part" of the language, and we're getting perilously close to the language being required to ship some heavy-duty HTTP stuff, possibly a server implemented in the language, before it stands a chance. Rust may have snuck in under the wire on that, though of course that stack is developing apace even so.)
Let the ecosystem provide, then, once there's a few clean winners, pick one as the official (while keeping the others there, obviously).
What is interesting is that I can't think of a lot languages that do it. It should be the default, as it's easy, safe, and gets good results. Somehow, it isn't.
And then there's the Haskell's way of: let the community informally choose the one best option, and when somebody uses any of the other ones, just get somebody near him telling "nobody goes there anymore, come to this other place". It works very well, but is a bit confusing for newbies.
Kind of, it supported distributed computing via CORBA and RMI.
A built in web server was released as part of Java 6, 2006.
> Can't speak to the other ones but I would imagine their first releases were also less useful than Go.
The fact is that Go, released in 2009, was a pretty bare bones compared to the state of those programming languages in 2009.
Also while Go might had an HTTP package on version 1.0, it surely still doesn't have a GUI framework on 1.8, which those languages had on 1.0.
EDIT: Rephrased a bit the answer.
Go isn't even that old yet. You're welcome to discuss what you like but I'm talking about what thing came with at first release.
"Also while Go might had an HTTP package on version 1.0, it surely still doesn't have a GUI framework on 1.8, which those languages had on 1.0."
Point. And even by 1.0 Java's was at least modestly capable by the standards of the time, as I recall. It didn't have everything, certainly couldn't compete with all the custom widgets you could buy for Windows, but usable.
Just in cases someone mentions NGEN, it is just intended to enable fast startups.
IDE power, the VCL, amazing documentation, rapid builds, clean OO language, it had it all.
It was a fantastic development environment even by modern standards but for its time it was well ahead of the curve.
I still try out Lazarus once in a while just for nostalgia :)
We are doing a project with a company that uses it for all their Windows applications.
Delphi conferences are still a thing in Germany.
Even C++ had such tools in the past via Energize C++ and VisualAge for C++ v.40.
Microsoft is now kind of following this path with /fastlink and improved database backend for code metadata.
EDIT: Typo
Seems like it might be worth the trouble/bootstrapping challenge if it yields another ~5%.
I mean the ARM stuff is getting more traction and maybe it will run around x86 in the next years.
- very large upfront costs in R&D and improved fabs
- The marginal cost for each additional CPU manufactured is, in fact, very low
Then, isn't this pretty much the textbook example of a natural monopoly?
I.e. in such a case, if the market was spread out over more suppliers, would the customer in the end pay more due to the vendors amortizing the R&D costs over fewer units sold? Vs. the current situation where customers are paying monopoly prices to the incumbent vendor.
Well... Seems the reason for the "multicore revolution" that started around 2005-ish(?) for desktop/x86 server CPU's was largely that CPU designers ran out of ways to make cpu's faster. Previously, there was always some micro-architectural feature waiting to be exploited (say, caches, pipelining, superscalar, OoOE, branch prediction, etc.) that enabled processor designers to utilize the transistors Moore's law gave us in order to increase single-threaded performance. However, it seems that at about the same time we ran into a triple whammy;
1) at about the same time that processor designers ran out of new tricks (with massive payoffs) to pull from their sleeves.
2) existing tricks ran far into diminishing returns
3) transistor scaling wasn't as good as before (the "power wall")
tl;dr: We seem to know a lot of tricks that can increase performance a bit, albeit at a heavy cost in power consumption.
Rust memory model around strong ownership and borrow-checking differs from the GC languages that can use generational strategies and lots of heap allocation re-use.
All languages are doing allocations to the OS, rust just doesn't have a GC like PHP does to intermediate.
Better than that is to not allocate at all, which is the focus of the effort here.
A step further down the line of optimisation may identify allocation hotspots and use custom allocation strategies rather than using the language default (the heap).