Why Aren’t More Users More Happy With Our VMs? Part 1 (2018)
tratt.net
tratt.net
I think that the use of small benchmarks obscures what’s going on. The VM is trying to win in the average. It’s like a professional gambler. Observing that the VM did something dumb for a program is like observing that a professional gambler lost a bet. That’s not interesting. In a game of chance, even a really great strategy will have its outliers.
I think that to understand the quality of a VM you have to throw millions of lines of code at it and see if the optimizing JIT can consistently produced speedups or at least produced speedups more often than not using some aggregate metric. As someone who studies the behavior of JSC on million line code bases, I can tell you that a pretty good outcome is if only a small number of functions experience an “upside down” effect from optimization and ends up running slower over time.
Finally, the whole search for a methodology to pinpoint warmup is broken. It’s pure brain damage. VMs need to be fast even for small programs that don’t have a chance to warmup. Startup time is absolutely important. So it’s a methodological antipattern to even try to find the warmup.
The questions worth asking are:
- for some program, how long does it take to run that program. Start to finish. No ignoring warmup.
- how long does it take to run some very long program or the average running time of a small program averaged over many iterations
- some percentile of behavior, like the 99th, to get an average of the janky behavior.
Ideally you measure all of those things and include both short running and long running programs.
This tells you how good a VM is.
If you’re doing math or methodology to identify the warmup point then you’re effectively biasing your experiment to forgive VMs for bad behavior so long as that bad behavior happens early. Nothing could be sillier. Users care about the perf of their VMs at startup not just in steady state.
Anyway, that’s the way I like to do optimizations in JSC.
This methodology likely comes from Java, which has long-running server applications. "How long does it run" is often "until someone hits ^C". Here, startup cost can be slow as long as the peak performance is fine. It's accepted that the first minute or two of the server are slow, but that's small compared to the month or so that the server will be running for.
> This tells you how good a VM is.
I think papers like this approach it from the wrong angle. I don't care about the VM's theoretical peak performance. I care about being able to measure and track performance in a reliable way. Put simply, I'm fine with bad codegen as long as I can consistently measure it. Feel free to improve it, but adding to sometimes give me good codegen, unreliably, is much more frustrating than bad codegen. But this seems to be the way the VMs are going, with things like probabilistic profiling.
If I refactor my code and replace for(let i = 0; i < L.length; i++) with for(const i of L), what's the cost? Will performance go up or down? We don't have tools or metrics to handle that right now. How can I ensure my codegen is good won't regress?
I work on a particularly demanding website in my free time ( https://noclip.website/#smg/AstroGalaxy , unfortunately won't run in WebKit due to missing WebGL 2 ), and performance varies drastically from Chrome release to release, and I do extensive testing with node.js to make sure that I'm getting good codegen.
I hear ya that having tools would be great - but the best speedups do come about from probabilistic methods so it would be weird to rely on whatever a profiler told you.
V8 used to have their own JS benchmark, Octane, but they retired it at about the same time as we beat them on it. So JSC is fast enough to make other people retire their benchmarks.
And by the way if you are interested in what we think of as good methodology you should read about JetStream 2: https://webkit.org/blog/8685/introducing-the-jetstream-2-ben...
Also a standalone build (kinda old) for various platforms: https://github.com/Lichtso/JSC-Standalone
JIT is very nice in theory. It's great in certain applications (eg; in very tightly scoped domains like accelerating linear algebra). Its proponents always talk about how it allows for optimizations that would be too costly or difficult when doing AOT compilation. But the operational complexity to get it to actually perform at that level on a production language VM (eg; oracle's JVM) is often its undoing.
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
Edit: looked at the source, which is easily accessible. The Java examples are reasonable (if slightly performance-minded, but nothing shocking).
Is ballast[1] still required for certain use cases?
[1]: https://blog.twitch.tv/en/2019/04/10/go-memory-ballast-how-i...
(Not a defense of Go's GC, necessarily. I'm not sure I'd call it "top-notch". A lot of its advantage over Java was that it got to look at Java and make different decisions, and by having a lot more values in the language with fewer references, part of the reason it tends to do a lot better than Java in terms of memory usage is just that it gave itself an easier memory management problem in the first place. Java GCs are frightfully good, yes, but to some extent they are that good because they have to be. Java at the very, very beginning was not designed for the sort of usage it has today (it was very originally a set-top box language, not a Big Iron language), so the language does some things that stress its GC. Go's doesn't have to be that intricate to be still quite good, so it isn't. And note difference between "quite good" and "top notch".)
In java you would simply set -Xms to increase the min heap size and then that additional memory would also be available to the application and not wasted as in the Go case.
It may be "nonsensical" but on the grand scale of "things done for memory management's sake" I find it unimpressive. (If 10GB were actually allocated and unavailable, I would find it impressive.)
To give a broad example, and to repeat the parent comment, I can avoid "ballast" by setting -Xms. But it goes further, if I have an existing JVM code base, and want to run it using a GC that exhibits similar low GC pause time to the Go GC, assuming I'm compatible with Java 12+ I can use the command line option -XX:+UseShenandoahGC without changing my code.
If Shenandoah isn't producing the behaviour I want, I can change to another algorithm with -XX:+UseG1GC for example.
I totally understand that Go's trying to avoid such complexity, but to quote Python's zen, complex is better than complicated. And I consider writing code to influence underlying runtime behaviour to be needlessly complicated.
No it doesn't. The Go GC is intentionally very simple and optimised for one specific metric, where as the set of JVM garbage collectors allow you to optimise for the metric that matters to you and are tuneable for the requirements of your application. The JVM has state of the art garbage collection; Go has My First GC Algorithm.
This is simply not true. The Go GC is an example of a sophisticated non-moving concurrent collector.
> optimised for one specific metric
This is only partly true. Low pause time is definitely the highest priority metric, but throughput still matters. The Go GC probably has the lowest pause times of any production GC these days, outside of perhaps nonpublic custom solutions, e.g. successors to IBM's Metronome sold to specific customers.
Citation needed? Java may not beat C++ in most use cases, but if you need a garbage collected language, it's likely the fastest you're going to find. But you get high memory usage in return.
The main trick is to interpret the short running code.
That said, some JITing makes page load faster. But only if the JIT only kicks in for functions that run more than some amount of time and the JIT is very cheap to run.
Like more resources have probably been put into the JVM than any other VM or compiler on earth and what has it given us, exactly? Performance is still worse than C and with homogenized operating systems, as well as the move towards the web, the portability guarantees don't feel very important. If I'm missing something, please tell me!
Meanwhile, C performance has not been a worthy goal in decades. Getting "as fast as C", if you could get there, would still leave you firmly in second or third place. Thr computers are not getting much faster anymore, but the problems are getting much bigger and the networks much faster, so performance matters more each year than the last.
True for each core on a CPU (broadly) but not true for the computer as a whole unit, we've seen an explosion of core counts on x86 desktops/laptops over the last 5 years (helped by the resurgent AMD forcing Intel to stop putting out incremental upgrades to 2C/4T on laptop and 4C/8T on desktop every year).
> People writing code that has to manage resources besides memory have found that, given the right core language facilities, managing memory too is no bother,
If we assume that it is 'no bother' as stated that doesn't address the 8C/16T elephant in the room in that for maximum throughput on a modern machine we need to orchestrate running across multiple cores properly and that's a tougher problem to crack in a general purpose popular language without an explosion in complexity to the programmer.
The only such Java GC I know of is part of the Azul Zing VM.
Regarding the move towards the web, I think portability still matters, because a lot of web developers write code and potentially test code on Windows and but deploy on Linux for production.
I don't think JITs are the only way to address that trade-off in compilation time vs runtime performance metrics, and whether their benefits are worth their other costs is an interesting, program-dependant question, but they have strictly more information available than a naive compiler and can potentially use that to make better decisions.
Here's a reason why this might not be a super great thing to do, if you're thinking of it from a language design standpoint:
- Profiling and rejitting costs you memory and makes all code pay a tax. The biggest tax is the safepoint/OSR tax. It costs significant memory and some small amount of time (maybe the time cost you add from comprehensive OSR support is like 1%-5% overall, but I'm not sure, because it's hard to isolate this cost and measure it). So you don't want to build a system that supports JITing unless it's going to give you big wins. That implies languages with lots of dynamic typing.
- The languages that most benefit from JITs (the win they get consistently overcomes the overhead of JITing) are the ones that have so much dynamic typing features that compiling them to native is a super hard problem.
Java is one of the few languages that is both dynamic enough to benefit from JITs but static enough to be possible to compile to native. And even for Java the native compilation is a lot of effort to get right.
Note that "possible to compile to native" to me means: compiling to native produces something that performs well so there exists some benefit to actually doing it. Like, I would expect compiling JavaScript to native to produce something that doesn't perform well at all.
So, basically, the issue is that most languages are either in the "benefit of JIT is smaller than cost of JIT" category because they have adequately static typing or in the "must have JIT and cannot compile statically" category because they don't have static typing. Not a lot of languages are in the sweet spot where doing both would help, but such languages exist (Java) and there's no reason why they can't do what you suggest and they may already do it (seems like a natural thing for Graal to do).
Every user should have it's own PGO and regularly refresh it to be competitive with the advantage of JIT profiling
A compiled language with a "micro-JIT" (e.g. for v-tables) seems like an interesting idea to me.
- gcc - clang - msvc - intel
In some sense we can always fully specialize a program to its current input and make a lot of radical changes to data layout and represesentation. But in languages like C/C++ where programs can "see" struct/array layout, member order, pointers as integer values that have fixed order, spacing, grouping based on the program specified data types, this is prohibitively complicated.
All of the warmup, and transitions from interpreter, to JIT, to optimised JIT , happen inside the first few micro or milliseconds of EVERY one of their thousands of process iteration. Their measurements are ALL of the system variation of the VM after warm up has taken place. The VM is optimizing within the first 1-1000 inner loops occuring at the start of EACH process iteration. For most working programmers, a variation of a few percent on a running system AFTER warm-up in "steady-state peak performance", and before any I/O takes place (because language benchmarks avoid I/O), would not be an issue. If it is an issue, then the article perhaps demonstrates that a compiled language would offer less variation.
The benchmarks listed range from a shortest of around 0.4s for fannkuch/hotspot/linux, up to 1.8s for n-body, pypy, linux. This 'long-running' benchmark code (of .4 to 1.8s ), by definition, has to include multiple inner loops/hot code, which is quickly optimized, otherwise benchmark code would have to be millions of lines long, in order to have a sufficient runtime length. Tests need to run for at least tenths of a second, for cross language comparisons, since JITted languages take some iterations to warm-up.
They’re trying to show that “warmed up steady state” isn’t something that reliably exists.
The final graph shows a binary trees program in C, with a 6% variation between "in process executions", and no steady state, it seems logical that most VMs will show the same or worse variation.
The "warmed-up steady state" does exist, but not if they define it so narrowly. All of their iterations and timings are running at x30 to x100 interpreted speed, the only 'cold' interpreted code is in a few microseconds of the first loops of an execution.
So you’d need a heap snapshot or some way to link the generated code to a different heap.
Maybe not impossible, just hard enough that it’s not widespread.
Still not impossible but I want to be clear on what exactly makes this hard. CPU model for example is not what makes it hard.
I assume this isn't literally using hardware watchpoints?
EDIT: worth noting that other JS VMs like JSC have these APIs for bytecode too, but i can’t remember off hand whether they can do generated code too.
Also, the charecteristics of your test suite may be very different than how it is run in production.
Outside of startup (which has gotten so much better in the last years I don't even see the splash screen anymore) and initial indexing upon project creation or library downloading, it's perfectly fast enough.
Not all the plugins respond in time while typing, but I can at least keep typing and it continues to function with basic editing (general worst case).
Virtual Machine Warmup Blows Hot and Cold
https://www.youtube.com/watch?v=LgCHAU8ZB00
Why Aren't More Users More Happy With Our VMs?
People are not using VM-based languages because they want to beat the numeric performance of hand-optimised Fortran. They use it because of the memory safety, compatibility and performance which is on par with or even better than other languages (at least when it comes to the JVM) for the tasks that they want to perform, which is for most programmers not going to be numeric simulation.
- Dependency hell and deployment problems: It's hard to make correct assumptions about which VM version is available on which platform. Pre-installed versions interfere with side-installed versions, and there is a ton of software that requires older Java VMs to work properly. It's a huge mess.
- Potential for losing future OS support: Apple, Microsoft, and others may at any time decide to block Java VM or no longer support it on their platform. That means you have to bundle your software with a Java VM, e.g. Crashplan has done this, making installation and deployment even more difficult.
- A thousand past problems on Linux: Various versions of OpenJDK and Oracle's java in combination with user software written in Java have caused massive problems on my Linux machines during the past 15 years, from causing extreme slowdowns to freezing the desktop until you hard reset.
Nothing else has given me as much troubles on Linux than Java, not even proprietary graphics card drivers and kernel extensions. Whatever the Java VM does, if it can freeze your whole system just because you run desktop software like Jabref, then there is something wrong with it.
Some programs don’t “run long enough” in the way the VM needs even when users run those programs.
I'd be surprised if the VM developers didn't run these.
More likely, the environment the VMs are run in has changed in the decades since their development: amount of RAM, cache size, latency of one subsystem over another....
"When we set out to look at how long VMs take to warm up, we didn’t expect to discover that they often don’t warm up. But, alas, the evidence that they frequently don’t warm up is hard to argue with."
By not warming up they refer to instances when early performance is higher than later performance or when the performance doesn't settle.
which are both cases where you would be disappointed were you to use the VM as a server.
I think their point is that VMs are more unpredictable than many realize and also intrinsically unpredictable in some cases.
The JVM has always had a slow startup path. Much slower-overall systems like Python don't. That's why people don't complain about Python performance and do complain about Java performance.