Kotlin doesn’t need an LLVM backend (2016)
blog.plan99.net
blog.plan99.net
https://blog.jetbrains.com/kotlin/2017/04/kotlinnative-tech-...
It features some interesting language support for C interop, an LLVM backend, and a reference counting+cycle detecting garbage collector. I wish the team the best of luck with it!
Meanwhile, I've been working on the tool I suggested in the blog, that uses Avian to create small self-contained binaries easily. For people who like using Go to make command line utilities and the like, Avian+Kotlin could be a very nice alternative. For high throughput servers you are still best sticking with the JVM, in my view.
This is actually very false. Yes they use GC but they also drop 2-3 frames at each 30s interval because that's when the GC runs.
Go boot up Gears of War and let a character sit idle with no action, you can set your watch on 30s spans by the obvious frame hiccup. UnrealScript was great for prototyping and empowering designers but I wouldn't hold it up as an example of a high performance runtime.
(Yes, really. Adding atomic operations to every load and store of a pointer is a tremendous throughput loss. Latency is not everything!)
But I think a deeper reason might be that making reference counting reasonably fast depends on the compiler proving that certain objects aren't shared across threads. If RC itself were slow, nobody would use talloc().
1) Ownership model means RC only needs to be used rarely in Rust, and ARC even less, minimizing the throughput impact?
2) Simplicity benefits were overwhelming considering the targets you were aiming for?
Or something else?
Yes, they don't enforce you to properly specify and transfer ownership semantics at compile time with the borrow checker (although you can enforce some of these things with modern C++), but the cognitive overhead is there. And if you make a mistake you don't get a scary compiler error - you just get a hard to find memory leak, an even harder to find data race or a gaping security hole.
Even Go (which usually touted as the perfect alternative to Rust in these conversation), has a lot of cognitive overhead related to ownership. Every time you take a byte slice and pass that forward to another function, you have to know whether that function appends to that slice, in which case your original slice will get overwritten. Case in point:
arr := []byte("Hello World")
b1 := arr[:6]
b2 := append(b1, 'Z')
fmt.Println(string(b2))
fmt.Println(string(arr))
https://play.golang.org/p/QIN5SAIL5MMany byte manipulation functions in the go std library and ecosystem have append semantics, for instance: https://godoc.org/regexp#Regexp.Expand
Better yet, Go often makes it very hard to notice whether you're copying values or just assigning a reference. On the other hand, some languages like Java don't even let you copy values reliably without special support in the class and in the same time highly encourage doing everything with stateful mutable objects. You can't just go and pass all these neat POJOs and beans around to functions modifying their state if they get used around somewhere else, right? This is exactly what all these rust borrow checker warning are all about!
The only language I can think of which is completely devoid of ownership semantics overhead is Haskell. Everything is immutable, look ma no side effects! So if this is really the only benchmark for productivity, Haskell is the most productive language ever.
> ARC makes no guarantees in the presence of race conditions.
So, it seems they agree with you, to the extent that they agree that full atomic reference counting would be a bad idea for performance optimisation, which is why they don't do it.
The race conditions that quote is referring to other kinds of races, not races on the reference count itself.
Personally my go-to solution where I need the flexibility of GC is Lua. It's light-weight, stupid fast for what it is(thanks to LuaJIT). You can also constrain it really easy and running multiple instances is very well supported. If you partition your problem space right you can keep each domain small and thus your root set doesn't grow like it does in JS/UnrealScript/etc. There's actually a lot of parallels with Erlang's root-per-process GC and it.
Oh and it's also stupid-easy to drop into any C codebase.
So given Objective-C's semantics, having the compiler insert [... retain] and [... release] calls according to Cocoa patterns, makes more sense than the original GC idea.
Which only works for Objective-C classes that follow such patterns, not for plain structs or basic C like data types.
Swift also adopted ARC, because it needed anyway to interoperate with the Objective-C runtime.
RC is nothing new, it was the very first GC algorithm being implemented, there are many reasons why GC research moved beyond it.
Making RC algorithms achieve the same kind of performance in multi-core core using NUMA architectures (multi-level caches) is as complex as just having tracing GC algorithms, with the caveat of not handling cycles and possible stack overflows due to cascade deletions.
It's not about RC, it's about the A in ARC.
Removing redundant increment/decrements was a well known optimization long before Apple added that A to RC.
What is the technical difference at implementation level?
Because ARC is just shitty GC.
That is, for each RC system of a certain complexity / feature set / etc., there exists a corresponding GC system that requires the same amount of work, and as the RC system does more work, so does the GC system. The exact times at which you have GC pauses, or pool cleanups, or whatever you want to call them are different, but the total amount of overhead is the same.
Of course, it's still worth thinking in one of the two paradigms: for instance, if you need precise timings of a function's execution (but it's okay for that function to be slow), you may prefer the immediate overhead of refcounting instead of amortizing the cost into some later garbage collection. But if you can optimize your refcounting overhead based on some application-specific knowledge, loosely speaking, you could also optimize a corresponding GC system by using the same insights.
Really. There is no magic here. The way you get compilers with 100% or greater perf of gcc/llvm is not through silver bullets. It's through a large amount of hard work and tuning. IE hundreds of person years.
It's usually easy to get the first 60-70%, and then people say "see, we didn't need <heavyweight thing>, we'll beat them with <lightweight thing>". But if you have customers where no performance is ever "good enough", getting that 30%, and eventually beating other compilers, is the thing that takes 10+ years. You'll never do it with <lightweight thing>. That's why all these folks who try eventually end up with 3 layers of compilers, etc.
If it could have been done with <lightweight thing>, it would have been done that way.
Also, you can just about always transfer the metadata necessary to do high level optimization x at a lower level. it just may not be efficient to do so.
Conversely, compilers that focus on fairly high level optimizations often do a bad job at lower level ones :)
Again, tradeoffs, tradeoffs. (This is why, for example, swift/rust/etc have some higher level IR, do what they can, and hand the rest off to llvm).
Graal is "new" only in the sense that it only recently got good enough for people to care about it. It's actually been in development for years ... arguably close to a decade now given that it traces its heritage back to the Maxine VM.
Graal's generated code quality is similar to HotSpot C2 i.e. the fastest available Java JITC, at least for most ordinary benchmarks (I have a feeling it may be weaker at auto-vectorisation than C2 but I'm not sure). Where Graal really spanks the competition is with more dynamic languages like Scala or Ruby though. There it sees huge speedups compared to other state of the art compilers.
Graal can also JIT/AOT compile LLVM bitcode. I don't know what performance is like for that, but earlier prototypes that JITd C ASTs directly could compile C with performance about 7% worse than GCC/clang (best of for each benchmark). The last 7% will take some work I guess but it's not infeasible.
There's a long term plan to make Graal the default JIT compiler in HotSpot. The problem at the moment is it - no surprise - it uses more memory than C2, and shares the same Java heap as the application.
Adding to that the hairiness of having to deal with concurrent versions (often the system's versus what was included with the software package) and how even distro-official JRE packages usually follow the "postinst script executes installer, installer downloads a blob and covers the whole system in untracked goo" format, I'm usually very relieved to scrub my system of all traces of Java after I am done with whatever forced me to install it.
I don't understand them either. Perhaps they are focused on server software and so don't think about client/end user scenarios much.
In particular, I think the argument about memory consumption is that it's independent of compilation strategy, but rather it depends on the quality and tuning of the garbage collector. An AOT compiled program that uses the default JVM GC with default tunings will not see much better memory performance (unless the memory consumption is dominated by the JIT).
On the other hand, I think the "AOT improves startup times" argument is a good one, and the author's explanation was unsatisfactory: AOT Java programs don't start any faster than JVM programs. What is the methodology for testing? Is he using the later-mentioned Avian? Because that tool deliberately trades startup time for small binaries (via compression). Even if he isn't, it's likely a matter of tool maturity--there's no fundamental reason an AOT Java Hello World should be slower than a golang hello world (which clocks in at 50ms on my machine compared to 130ms for JVM). Also the author implies that startup times only matter for tiny unix tools--I'd like to counter with a practical example: Gradle is so slow to start up that it's been daemonized. I don't know if this satisfies the author's characterization of "tiny unix tools" or not, but clearly start up times matter.
Gradle is so slow to start up that it's been daemonized.
Gradle is usually run from the command line, so surely that supports my point? (perhaps emphasising the "tiny" was a mistake).
At any rate, Maven doesn't have the same problem, despite doing the same task. Gradle is slow to start for a lot of reasons.
You're right; I paraphrased poorly. I think my point stands--there's no reason an AOT compiled Java Hello World should be slower than a Go Hello World, which is nearly 3 times faster than the JVM version on my machine.
> Gradle is usually run from the command line, so surely that supports my point?
I assumed that Gradle was slow to start for JIT warmup and the like. It's hard to believe that daemonizing is the performance solution if the bottleneck is in the application (for example, if file parsing is the bottleneck, the natural solution would be to cache the parse result). Maybe my assumption is bad?
You can now configure Gradle (since v 3.0) using Kotlin instead of Apache Groovy. That might speed up Gradle start up times.
Luckily, RoboVM was forked after a while and it is being actively maintained by MobiDevelop [1], so give it a try if you're into Kotlin, that scene deserves a lot more love.
What is the cognitive load with the JVM? You install it and the type 'java my.jar'
BTW the JVM can reduce cognitive load hugely when it comes to profiling or debugging because of the tools.
Only knowing "java my.jar" means you have zero knowledge about the nature of your language's run-time and having a good understanding of your languages run-time is necessary to write good code. The more complex the run-time of your language, the higher the cognitive load of that language.
His discussion of memory use doesn't mention the pre-allocation of the heap for the JVM. The memory used by the application itself is irrelevant when the JVM is what is actually using the memory.
> pre-allocation of the heap for the JVM
Pre-allocation doesn't matter, it's not allocating physical memory?? The GHC Haskell runtime literally mallocs one terabyte when the program starts :D
However OpenJDK/HotSpot's baseline memory usage is indeed higher than many other VMs'. Even if you compare with alternative JVMs: I tested a basic Jetty app — it used ~60 MB on JamVM, ~70 MB on OpenJDK.
This assertion is too harsh. Eclipse does seem to be over-engineered(speaking from the point of view of someone who had to write a proprietary plugin for it) and Minecraft does seem to run slow on even the latest machines (compared to AAA games).
But Minecraft is not a trivial program to write, despite fooling people with low-resolution textures. it is, "algorithmically speaking", out of reach of many people, including those with a reasonable computer science background. I dare someone to do better (on the JVM!)
On a personal note, I cringe every time I have to go to a kid's computer and modify "-Xmx" settings. I'm running a game, I want it to have all available resources not otherwise allocated to the OS, just like almost every other game there is.
Minecraft's code is rotten throughout and nearly nobody in the community cares about code quality. Remember, this is a group that thought that TLS was easy to reinvent: http://wiki.vg/Protocol_Encryption
I would also consider, though, that the client side code for such games is more complicated to reason about. More iterations through the voxel data needed to compute lighting and meshing. Still not something that should necessarily be impacted by GC. From what I've read, Minecraft's more recent versions switched to a strategy that boxed more function parameters in objects, adding to the load.
I cringe every time I have to go to a kid's computer and modify "-Xmx" settings
I think if you switch to the G1 collector you can just take it out? Minecraft really shouldn't need that kind of tuning though. Again, they don't really care.
Tuning is required when you start to install mods. Or even texture packs.
>
> [Goes on to show how Hello World is just 1 MB, which is smaller than Go's Hello World]
> It’s not a trivial Hello World app at all, yet it’s a standalone self-contained binary that clocks in at only one megabyte. In contrast, “Hello World” in Go generates a binary that is 1.1mb in size, despite doing much less.
To show this to the extreme, here's a "hello world" that's less than 30 bytes:
#!/bin/sh
echo Hello world`go build -o /tmp/hello -ldflags='-s -w' /tmp/hello.go && upx --brute /tmp/hello`.
Also worth noting that the author is incorrect about the Avian example being self-contained--it dynamically links against 4 libraries (on OS X, anyway) while the Go version links against zero.
Further, the Avian example is broken on OS X (the GUI loads, but immediately freezes).
Of course, the applications aren't doing the same work, so comparing the binary sizes is silly. I just wanted to point out that avian's tricks are readily accessible for Go programs as well.
And how exactly does that work out on platforms that don't have stable syscalls?
On Windows, Go dynamically links to kernel32.dll and friends.
If BC concerns weren't a thing, a fresh Java stdlib would remove a ton of the size complaints.
const io = @import("std").io;
pub fn main() -> %void {
%%io.stdout.printf("Hello, world!\n");
}
It could be a lot smaller, but I'll have to fix an LLVM bug[2] or two.For comparison (these are all with full optimization and stripped):
* Go: 1,006 KB, static executable
* Rust: 411 KB, links to libdl.so, libpthread.so, libgcc.so, libc.so
* C: 6 KB, links to libc.so
* C++: 6 KB, links to libstdc++.so, libc.so
* Nim: 26 KB, links to libdl.so, libc.so
* D: 546 KB, links to libpthread.so, libm.so, librt.so, libdl.so, libgcc.so, libc.so
libc.so is 1.9MB. libstdc++.so is 1.5MB.
[1]: http://ziglang.org/
No I'm not going to install your app!
I agree with a lot of points in this article. The JVM does a lot of nice stuff for you. One of the big, non discussed, advantages of compiling to native would be removing the JVM dependency (and hence the dependency on openjdk or the terrible Oracle JVM license).
As far as LLVM, I'll agree it's a technologically impressive. Still, I feel a clear tech implementation was only partially the motivation, and the other was getting away from GPLv3/FSF licensing. That bothers me a bit. I like the GPLv3, but I feel much of the industry wants to distance itself from the licensing that essentially helped pull Linux into what it is today.
By "Linux doesn't use GPLv3?", I think the poster meant essentially "But...? Linux doesn't use GPLv3".
In any case, the idea is a valid point. The GPLv2 and the GPLv3 are substantially different and it isn't that surprising that some people who were okay with GPLv2 terms are not okay with GPLv3 terms.
"Larger code bloats downloads and uses more RAM, lowering cache utilisation". This depends on your definition of cache utilisation: The code gets compiled, regardless of whether it's AOT or JITed, therefore the cache gets used anyway?
I understood the JVM does an dynamic analysis of the code when running, so it can compile to better performance. For example it knows which branch of an if it is more likely to be executed, so it can compile that code to native in a way the CPU jump prediction executes the branch that will probably executed at the end. (it is hard to explain without more knowledge but CPUs do optimizations in branches and usually executed one branch before even know the result of the condition, so you sometimes save some cycles, with the dynamic analysis you know which branch is more often executed). I guess the author is talking about this type of things
Probably you know more than my about optimization, but I think the idea is clear here. It is easier to get better gains with high level code and profiling in real time, than with low level code and simulating the profiling.
This is pretty much false these days. Most folks who care about it use it.
You were right 10 years ago. At this point, there are enough companies where 10% perf is tens or hundreds of millions of dollars, and so the world changed.
"As a consequence compiler writers rarely focus on it and C++ is designed on the assumption that optimisation is always done statically."
Also false as a consequence of the above :)
"Java is designed on the assumption that JITC is available and JVM designers focus more on PGO as a result. " But they have less time to perform any of the expensive opts pgo enables. so ...
https://github.com/python/cpython/blob/master/Makefile.pre.i...
Not sure when it was added, but Python 2.7 has it.
It's not clear to me how to tell if say my Ubuntu binary is built with PGO, but it's something I'm looking into.
Yes it was. Maybe you didn't know, but it was.
"and Chrome is one of the most intensively optimised apps out there. " This isn't really true actually. So if that's your counter-argument, ...
"So I think it's fair to say that most C++ apps built by mere mortal companies probably still don't use it" No, it isn't :) Really. So far you've cited one example of someone being late to the party.
" But I'd be keen to see data showing I'm wrong." What form do you want it to take?
JIT PGO is adaptive to the workload that the process is experiencing. It is likely impossible to achieve this with AOT compilation.
Of course it's not :) You just generate different binaries for each use case.
Truthfully, the "million machines/million use cases" story is pretty non-realistic these days (except in the above case!)
Most people run mostly-specialized services for mostly-specialized use-cases. This is because it's necessary for performance in practice anyway.
The argument that JIT having more information and therefore should/could be able to produce better code never really materialized in most the of mainline JITs (JVM,.net etc...).
From my point of view, there is a lot of reasons for this :
- Most VM starts with higher level languages than native systems. And most of the "runtime informations" is not used to produce faster "native code" as used to efficiently deconstructs some of the higher abstractions (think devirtualization, unboxing, moving stuff from the heap to the stack etc...)
- Because Jitting happens at runtime, most jit shy away from complex compiler transformations : So you have more runtime information but less time to actually do something useful with that information.
- Tracking, recording, updating and all the booking involve in collecting any runtime information also limit "quality" of the profile in practice further more reducing their impact.
> I guess the author is talking about this type of things
I don't think so, the author is probably talking about optimization which lowers higher level constructs, like unboxing etc...
- deep inlining, by default the inlining depth of a JVM is around 10, you can not have this depth of inlining with a statically compiled language.
- prune non taken branches (and exception blocks), so you do not fill the instruction cache with code never executed.
- branch re-ordering, the code executed more often can be read from top to bottom, less executed code is offloaded to a side block.
- CPU instructions selection at runtime.
- lock profiling, generate different codes if the lock is contented or not.
BTW .Net do not do any code profiling, so can not use most of these optimizations.
Disabling JIT PGO in HotSpot loses you 15-20% performance, so I think it did materialise.
Because Jitting happens at runtime, most jit shy away from complex compiler transformations
I wonder what you classify as complex? Escape analysis driven scalar replacement, PICs, superword optimisations, lock elimination, automatic hardware transactional memory usage all seem very complex to me. HotSpot uses two JITs to trade off speed vs complexity, C1 does a very fast compile with only fast optimisations, C2 will revisit very hot methods and do the extensive optimisation of them.
Considering this was one of the main points of llvm in the first place, and it's had a jit since day 1 ..
http://llvm.org/pubs/2004-01-30-CGO-LLVM.html
"This paper describes LLVM (Low Level Virtual Machine), a compiler framework designed to support transparent, lifelong program analysis and transformation for arbitrary programs, by providing high-level information to compiler transformations at compile-time, link-time, run-time, and in idle time between runs."
Compiler IRs are normally not that difficult to implement (Designing them is a whole different thing), so it wouldn't be (If a group of devs decided do it) that difficult to define IRs with lowering between each of them. After transforming into a (Approximately C level) IR (Like Cmm in Haskell), you can emit LLVM IR to handle code generation, Register Allocation, LTO etc.