My experience goes contrary to that statement. How allocation is done makes all the difference. Between what the language implementations memory model is good or bad at, and what allocations/usage patterns the CPU is good or bad at, there are worlds of possibilities and trade-offs. Especially the “increased allocation equals decreased performance” myth doesn’t hold true. See http://mr.gy/blog/maxpc.html#section-3-2 for my own ragtag exploration of this phenomenon.
But you'll also get into the territory of what you're GCing. I take it from the fact that there are a lot of arenas that there are an awful lot of small allocations. Unless your GC lets you choose the level of allocation i.e. it doesn't necessarily allocate at the object level, your GC'd language is likely to be much, much slower. There may be GCs like this but I think they are relatively unusual (I'm not aware of them but I'm no expert).
It seems to me that the issue here is actually the same as you get when you try to improve performance of GC'd languages i.e. you want to try to avoid any allocation wherever possible.
I must say I was a little surprised by how much allocation is going on. Does it imply that idiomatic rust may lend itself to somewhat excessive memory use? Or maybe the rust code in the compiler is sufficiently old to not be considered idiomatic. Would be interested in the views of those more knowledgeable than me.
Rust used to be GCd, so the compiler uses a lot of things designed in a GC-esque way but tweaked so that they no longer need a GC.
In general the language has changed a lot, and the compiler source takes some time to catch up completely.
Yes, there are.
GC enabled system programming languages like Mesa/Cedar, Modula-3, Active Oberon, D, Sing#, System C#...
All of them allow you to control what goes into the GC heap, stack, global memory or is eventually managed manually in unsafe code sections/modules.
In any case my remark still goes, it doesn't matter how one acquires OS resources, there is always a performance penalty when it happens too often.
Deallocation is obviously a lot more involved, but it can be done in parallel on cores that may sit around not always fully utilised.
Compilers can normally max out all the cores you've got during a full build, so perhaps in this case it doesn't matter: enhancing the parallelism doesn't help you. On the other hand if you're rebuilding a single file due to incremental compilation, it's probably hard to parallelise that entirely by hand, so shoving some of the memory management work onto your spare cores can help.
It wouldn't reduce allocation count, it would reduce allocation cost to a simple pointer bump which is virtually guaranteed to already be in cache (unlike malloc tables).
The added cost is only tracing, but if, like the OP said, allocations conform to the generational hypothesis, then frequent minor collections are quite fast.
My guess is: Large compilation units and LLVM optimizations.
Since MIR has landed, IIRC work on doing MIR optimizations and LLVM IR simplifications is the next step here.
I would not be surprised if down the line, the Rust devs decided to write their own Rust backend rather than relying on LLVM.
Do any non statistical outliers care about non amd64 and ARM?
Even within the amd64 and arm families, there are enough variations (AArch vs qualcomm) and evolutions (SSE3 to SSE4) to make the task of developing a new and performant backend from scratch very resource consuming
Also while taken individually most of the non amd64/arm platforms might not be significant, combine together they still represent a big share of the pie, doubly more so in the embedded world which rust (tries?) to target.
For OpenPower, let IBM do it :)
1/3 or more of desktop Windows users are on 32-bit Windows, last I checked. So you definitely care at least about ix86 of some form.
What? No. It means that platforms with browsers are part of the domain Rust cares about supporting. Not the only domain. You asked for non-outliers, bz gave an example of browsers.
There's plenty of work to make Rust work on embedded platforms where no browsers operate, and some of this is being done by Rust contributors employed by Mozilla. No need to assume that Rust is only focused to the needs of browsers.
C and C++ optimisers weren't that good in the 80 and 90's even against junior Assembly developers, the 40 years of compiler research in optimizing C and C++ code that made them what they are nowadays.
Also not sure that reaching optimization parity with LLVM is an absolute necessity, Clang/LLVM is not as good at optimizing as GCC but that hasn't rendered it unpopular.
Looking at Go as a benchmark, the compiler seems to be improving the optimization at a good pace and supports X86, ARM, PPC, MIPS, S390 despite being relatively young language and also having recently rewritten the compiler.
For that you need to use gccgo.
The main problem with this kind of effort is mostly political, as many developers don't make a difference between language and implementation, thus making it harder for a language like Rust to win the hearts of its target audience, e.g. C developers that micro-optimize code as they write it.
Go usually wins in those benchmarks, where Java code is either interpreted or still compiled with the 1st level JIT.
And as mentioned, this is only relevant with the reference compiler, gccgo is much better.
It's not like the programs are run with -Xint :-)
It's not like a panacea -- https://arxiv.org/abs/1602.00602
Actually I did a benchmark run not that long ago between Gccgo and Go with the latter winning most of the benchmarks, I believe this mainly has to do with Gccgo lacking escape analysis.
Clang/LLVM is essentially as good as GCC. It depends on the benchmark you're looking at. Getting up to LLVM parity would take years and years.
> Looking at Go as a benchmark, the compiler seems to be improving the optimization at a good pace and supports X86, ARM, PPC, MIPS, S390 despite being relatively young language and also having recently rewritten the compiler.
It will take years and years for Go to achieve LLVM's level of optimization (assuming that they want to; I suspect that they won't when it starts regressing compile times).
Looking at Rust vs Go over at 'benchmarksgame' (granted these are micro-benchmarks), Go seems to compare very well for tests that are mainly computational and thus avoid involving the GC very much (even beating Rust on a couple).
http://benchmarksgame.alioth.debian.org/u64q/compare.php?lan...
And let's not forget that Go hasn't really had any focus on optimizations up until the last release where SSA was introduced.
Anyway, maybe I'm projecting as I would like to see a full Rust toolchain in Rust, and my (very possibly entirely wrong) impression is that Rust devs have been struggling with getting the full performance potential of Rust through LLVM.
It makes zero sense to write a new backend for Rust from scratch because Go does well on a few microbenchmarks. Go's compiler infrastructure is across the board worse than that of LLVM (in compile-time performance too; some of Go's algorithms have worse asymptotics than the more sophisticated ones in LLVM). LLVM doesn't do things like SROA and alias analysis just for fun.
> my (very possibly entirely wrong) impression is that Rust devs have been struggling with getting the full performance potential of Rust through LLVM.
I'm a Rust compiler dev and I have no idea what you're talking about. LLVM is working wonderfully for us. If we had done what Go did, we'd quite possibly have failed.
I have a different proposal. Instead of reinventing the world for no reason, let's treat bugs in LLVM/rustc as bugs, and fix them.
A copying-compacting GC has very fast allocation at the cost of very slow freeing (you have to walk memory).
DMD has very fast allocation and never frees memory.