JVM implementation challenges: Why the future is hard but worth it [pdf]
cr.openjdk.java.net
cr.openjdk.java.net
There are only 512 L1D cache lines in a CPU core. After just 512 references to memory areas covered by different cache lines, L1D entries will start to drop. Or earlier. In Java this happens very easily. Like when just iterating a single, not very large, data structure.
Everything else (except maybe SIMD/GPU related items for compute purposes) are not anywhere near as important as fixing Java's current problems with CPU caches.
All system programming languages with GC, value types and AOT compilers (some implementations of Oberon did have JITs).
Apparently 20 years detour was needed to go back to those features.
There were Bartok, mono -aot, Cosmo OS, but it seems if it isn't part of the reference platform, developers tend to ignore them.
I have meet a few .NET devs unaware of NGEN.
I think this is explainable by Java's history. When it was created, the design docs said Java was a simple, interpreted language with distributed objects and the target hardware was set top boxes. When Java unexpectedly took off, I suspect Sun weren't quite sure which ingredients were key to that success but fixated on "simple". They unlocked a huge market that was unserved at the time: companies that needed a much simpler C++. So that wasn't unreasonable. They then spent many years trying to handle customer requests for ever more runtime speed, whilst simultaneously not making Java as complex as C++. Hence the focus on super whizzy optimisations inside HotSpot ... many of which are just undoing performance hits introduced by the simplicity of the Java language.
Value types increase the complexity of the language, no doubt about it. Heck, even having a difference between int and Integer increases complexity. When value types come out, I expect much rejoicing from people who need more runtime speed, and a whole lot of confused questions on Stack Overflow from people who aren't totally sure about the difference between a value type and a regular class.
I also expect to come across codebases where classes were turned into value types more or less at random, or where the style code forbids value types entirely, etc. Java sort of has this problem already where weaker programmers who are using threads just slap synchronized keywords on every method until the crashes stop. Then HotSpot goes in and tries to figure out if the locks are actually needed or not and removes them at runtime.
Don't get me wrong. I'm A++++ in favour of having value types on the JVM and in Java. But increasing complexity of the language is absolutely not cost free, especially not given Java's value proposition to big corporates.
https://www.google.com/search?client=safari&rls=en&q=value+t...
L1D latency with a pointer is mostly 4 cycles. Not sure, but I think having 1024 entries would increase that to 5 cycles.
Increasing cache line size from 64 bytes to 128 bytes would require more memory (DDRx SDRAM) bandwidth, because CPU needs to always fill a whole 64 byte - 128 byte with this change - cache line on a memory load. Even if you want just one byte. Wasted bandwidth for non-streaming workloads would be increased significantly. This change would force also L2 to have 128 byte cache lines, etc. LFBs and other buffers would of course also need to accommodate this.
Skylake will be able to process 64 bytes (512 bits) with a single SIMD instruction, with AVX-512. Perhaps Intel will need to soon increase cache line size, if 1024 bits wide SIMD instructions are desired.
SIMD gather instructions might be a hint that in the future memory loads will function in a different fashion. Maybe cache lines won't anymore be the smallest amount of data a CPU memory subsystem can deal with?
Is SRAM comparatively power hungry or something?
If you put DRAM on the CPU die I would also expect it to cost hundreds per GB, but that's not a very interesting comparison because we know from current DRAM best practices that you can put it on its own die at a much lower cost.
Also, I looked up RLDRAM [1] and it looks like their "low" latencies are still 10ns, which is good for DRAM but abysmal in comparison to SRAM.
[1] http://www.micron.com/products/dram/rldram-memory/1_15Gb#/
E.g instead of an array of 3-vectors, you'd have an x-array, a y-array and a z-array.
If you're unlucky, physical memory read pointer distance in bytes is modulo of 64 * <number of memory channels>, all 3 reads might hit the same physical memory module, but in a different DRAM page. All accesses would buffer on just one memory controller and one memory module, significantly reducing available bandwidth.
At physical memory level read pointer distance between different arrays could also cause cache aliasing and consequently extra cache invalidations.
See http://www.intel.com/content/dam/www/public/us/en/documents/... for details. Find "aliasing" from it in your favorite PDF viewer.
If you are just sequentially and continuously streaming from one array, none of this should normally happen. You can of course still get conflicts leading to CPU stalls when writing data, but that's a different matter.
Considering the above, you might want to store your vector x, y, z in a same array, interleaving the values:
x = array[objIndex * 3 + 0];
y = array[objIndex * 3 + 1];
z = array[objIndex * 3 + 2];
Not pretty! In real code, you might want to use 4 as multiplier. Then a JVM with vectorization support might have (or might not) better opportunities to load whole vector together with a single SIMD (SSE/AVX) instruction.Alas, like so many other things about the Java language (No operator/assignment overloading, no unsigned integer types, etc, etc), it seems to have been a deliberate, although potentially under-thought, design decision.
What makes it a bad idea?
Having to rewrite an entire struct to change one value within it seems... counterproductive, considering that the purpose of having structs is largely speed-based.
I can see it being useful to be able to flag a struct as immutable for certain cases, but for there not being a mutable version?
The practical reason is that structs help you with speed when their size is no more than the size of the machine's cache line, i.e. up to 64 bytes. Beyond that, you incur an extra cache miss when accessing them (even below, if they're not aligned, but the smaller they are the smaller the chance), so you might as well use a pointer, and your copying costs are no longer negligible, so a pointer might be a better idea anyway.
So conceptually mutable "values" are a bad idea, and there's no reason to have them because beyond the cache line size they might cost you in performance more than they buy you, and below it mutation doesn't help as copying costs are very low.
The only benefit mutable structs might have is when you want to lay out a large number of objects, which you'll be accessing sequentially, consecutively in RAM. This way, the CPU's prefetcher will help you stream them into the cache. But the way HotSpot's allocator works, you're almost certain to get the same behavior with objects anyway, and if whatever tiny performance difference this might buy is that important to you, you might be better off using C, anyway.
The goal of the JVM is to get you the best performance for the least amount of work; not to let you work hard to squeeze out a few more performance improvement percentages. It sometimes certainly allows for that, but not at the cost of such a fundamental change, which is bound to do more harm than good.
1. https://www.youtube.com/watch?v=DCdGlxBbKU4 2. https://lmax-exchange.github.io/disruptor/
That said, it's a little scary to contemplate the power that microsoft, steward of the clr, and oracle, steward of the jvm have over our industry. It's staggering to see the amount of engineering that's gone into bullet-proofing the jvm. What happens when microsoft or oracle decline to continue paying for hundreds of very expensive senior compiler/vm/language engineers?
Since the early 2000 the majority of Oracle GUI tools across their products are written in Java.
Long before acquiring Sun, Oracle did acquire Bea.
Oracle still sells Bea Weblogic JEE container and the J/Rockit JVM, which also has a real time version.
They have their own JSF framework and IDE.
So becoming Java steward is not so strange.
But what really happened is people who had been locked into Sun with legacy software wrote applications in Java and then switched from Sun to cheaper hardware running Linux or Windows.
Sun called it: Mankind vs Microsoft [1]
It drained their energy and stole their focus.
It was the apex of Embrace, Extend, Extinguish
[1] http://www.theregister.co.uk/2004/04/03/why_sun_threw/
Interesting from that 2004 piece
> Microsoft's biggest global competitors are exactly as they were on Thursday: Nokia and Sony.
How different things looked then!
No, but they have an awful lot of downside from not maintaining it.
Consider the history of IBM, their grand strategy in the early 90s was that Smalltalk would be the enterprise language of the next few decades. The Java juggernaut rode roughshod over IBM. Or Microsoft, they would have loved to have kept Visual Basic and Visual C++ cash cows going forever, Java completely blindsided them. Altho' it may be technically less powerful than Smalltalk or C++, in terms of mindshare (or hype if you prefer) Java is too powerful for Oracle to risk falling into anyone else's hands.
No one you know.
Suppose some large and wealthy entity has trouble running some money-making legacy application that happens to be running atop the JVM. Who do you think will they call to fix their issues asap?
.. Which is why there's a lot of interest in having the people working on the JVM on your payroll ..
As for the languages and compilers, I'm less worried. There are already plenty of open source languages targeting these platforms, many of them better than Java/C#.
People will move away for something else, while a few will keep on maintaining them.
I remember when the answer for any database related problem was Clipper. Now it is gone.
I remember when no one would question the use of Turbo Pascal for both system programming as application programming.
I remember when using a spreadsheet meant Lotus 1-2-3.
I remember ....
The 8 bit home computers were all about BASIC, Assembly and Forth.
Those powerful enough to run CP/M, had some C dialects available, but no one cared.
When 16 bit arrived in the home computers, BASIC and Assembly were still the way.
On *-DOS variants, the OS was developed in straight Assembly.
Turbo Pascal was widely used in my home country from systems programing all the way to business applications. Very few cared about C.
For CRUD applications there was DBase and Clipper.
We only started caring about C when the need to take code to Windows 3.x started to be a reality and Turbo Pascal for Windows started to get behind the times.
Then Borland went crazy in their business decisions and many moved away from Delphi into Visual Basic and Visual C++ (MFC).
In the Amiga world, the OS was coded in a mix of Assembly, BCPL and C.
For the coders it was all about Assembly and AMOS. Although I think there were some using C with the likes of MUI and similar.
MacOS was originally developed in Object Pascal, the dialect later added to Turbo Pascal, and Assembly.
Apple eventually added C and C++ support, and while trying to cater for developers went C and C++, re-writing the Object Pascal parts.
So it always saddens me to see young generations think C was the one and only systems programming language.
It only became that because UNIX based workstations succeed in the market and everyone wanted a piece of the pie. Which partially meant using C.
I used to be a Delphi programmer and worked on, amongst other things, an open source video game project :) Ah, good times.
I stopped at Turbo Pascal for Windows 1.5, because by Delphi 1.0 time-frame I was at the university and wanted something UNIX friendly so went C++.
I knew C and C++ still from MS-DOS, never liked C in regard to what Turbo Pascal could offer, hence C++.
Also I got eventually fed up of writing Pascal wrappers for all WIndows APIs not provided by Borland.
Also the desire to go meta-circular and reduce even more the amount of C++ code.
Most likely based on the Graal/SubstrateVM work.
Finally JNI being replaced by something developer friendly.
Maybe the JVM allows for more fine grained control on PC location, and allows to "go to" different places in the code without invoking a native thread. But I know that when you create a Quasar co-routine / green thread / user thread / continuation or whatever we want to call it, then it's not the same as Akka.
In Akka, if you have a "conversation" between a million actors in a circle, you can still get 1 thread to be reused, if there is no need to use more than one. But if you try to do 3 things in 100% parallel, you might find yourself with 3 threads (from the little that I know).
In quasar, you can do all that, without invoking an OS thread, I think this is the main difference.
sample code: http://docs.oracle.com/javase/tutorial/essential/concurrency...
http://docs.paralleluniverse.co/quasar/
They claim to have "true lightweight" threads. (I am sure pron if he is a around can jump in and expand on it).
It takes a lot from Erlang even pattern matching.
The downside is that libraries have to be integrated in order to block gracefully when called on fibers, but we already have a long and growing list of integration modules for popular libraries.
It seems this would move most of the heavy lifting into the kernel (and out of the already-very-complicated JVM). Anyone know what happened to that work?
> Rule #1: Cache lines should contain 50% of each bit (1/0) > – E.g., if cache lines are 75% zeroes, your D$ size is effectively halved
Can anyone explain this?
At least for data, a common win is to compress the data in ram and decompress once it is in-core; this is often essentially free as many machine learning algorithms are bandwidth starved but have plenty of compute available. Suppose you are eg storing small integer counts in ints; if it's a java int you are using 4B to store 1B, while if it's a java.lang.Integer it costs 16 bytes plus most likely an 8B pointer.
Another way to consider this is if you are using 8B pointers, you waste a lot of that as constant zeros -- 1TB is 2^40, so even on a 1TB machine 3B/24b are wasted. Particularly with (all? at least that I know of) jvms that have 8B alignment, another 4 bits are wasted. This is how the CompressedOops hack words to access 32G ram w/ 32b pointers.
I have problems imagining how it could be done. Would you mind elaborating and/or sharing some examples ?
That is more or less what the original Macintosh did with its handles. The hardware wrapped addresses around at 24 bits. That allowed Apple to store 3 bits of data in each handle ('data can be purged from memory if needed', 'data cannot be moved', 'data is read from a resource'). Third parties sometimes used the other 5 bits. With the introduction of Macs with more memory, that led to the "program has special memory requirements" system error, which actually meant "this is Excel version so and so. It cannot run on a machine that has 32 address bus lines because it mangles pointers"
Say you have 8bit integers; store them packed in ram, then upon reading, unpack. So instead of striping an array of int[], you have an internal array of long[] and you read them with a function. Your memory read will suck in 8 at once.
The same technique works for floats or ints with a small-ish range and limited precision; you can store a scale and offset, then pack on write / unpack on read. It's common to be able to quadruple your effective memory bandwidth, then the read operation -- ie
// instead of:
double[] _data;
// accessed as
_data[idx];
//instead you do
getd(idx);
// using the below
bytes[] _mem;
float _bias, _scale;
double getd(int idx){
long res = _mem[idx];
return (res + _bias) * _scale;
}
executes entirely from registers, and is essentially free. The price of all this is you have to process your data on ingestion, but if you run iterative algorithms -- like convex optimizers -- that repeatedly walk your entire dataset, this is often a big win. You can often lose some of the low precision bits on the float or double, but those probably don't matter much anyway.Like anything else, you'll have to measure.
Here are some specific instruction examples about how it's done.
Bit field operations; Bit field extract, etc.
http://docs.oracle.com/cd/E36784_01/html/E36859/gnydm.html#s...
Parallel bits extract / deposit.
http://docs.oracle.com/cd/E36784_01/html/E36859/gnyak.html#s...
SIMD (AVX2 in this example link); Pack, unpack, shuffle, permute, broadcast, etc.:
http://docs.oracle.com/cd/E36784_01/html/E36859/gntae.html#s...
>Suppose you are eg storing small integer counts in ints; if it's a java int you are using 4B to store 1B, while if it's a java.lang.Integer it costs 16 bytes plus most likely an 8B pointer.
This wastage does happens only in reference types? Tomorrow say if value types are been created, then these sort of issue will get resolved?
>Another way to consider this is if you are using 8B pointers, you waste a lot of that as constant zeros
Why of constant zeros? And in the line your saying "-- 1TB is 2^40", your referring 1TB of cache line?
That would solve most of the issue, yes.
> Why of constant zeros? And in the line your saying "-- 1TB is 2^40", your referring 1TB of cache line?
He was not referring to cache line or anything like that. He was referring to the fact if you can address just 40 bits, upper 24 bits in a 64-bit pointer are wasted in a pointer. I think on x86 real virtual address size is currently 48 bits, so upper 16 bits are "wasted". This wastage will occur with heaps larger than 32 GB when using compressed pointers JVM option. And of course 4 GB (or 2 GB?) with that JVM option disabled.
Similarly, if you use "bool" type, it still takes a byte in current JVM implementations. So you could say 7 bits are wasted there.
1. If OOP pointer is going to eat up our memory, there we loose / waste some bits of memory. 2. And as you have said even using bool type, is going to say waste by 7 bits!
Since we are dealing more with objects in java world, then I guess, 32 bit VMs will be much faster than 64 bit VMs. Or even 16bit will be more faster, from my understanding.
Still I have a final question over here, all the OOP pointers does goes the CPU caches and registers?
Example of generated code from some random web page that happened to be there:
http://cr.openjdk.java.net/~shade/8050147/rsp-minus-8.perfas...
[0x7f189918ae07:0x7f189918ae54] in org.openjdk.VolatileBarrierBench::testWith
# [sp+0x30] (sp of caller)
0x00007f189918ade0: mov 0x8(%rsi),%r10d
0x00007f189918ade4: shl $0x3,%r10
That "shl $0x3,%r10" multiplies the pointer in register r10 by 8 (2^3). So looks like JVM computes the base pointer to an object to a register (r10 in this case) and uses relative addressing from it.Here's example of using x86 indirect addressing mode to multiply the pointer by 8:
http://shipilev.net/blog/2014/safe-public-construction/stead...
0x00007fc7b4971644: mov 0xc(%r12,%r11,8),%r10d ;*getfield instance
In this case effective address for the load from heap is r12 + r11 * 8. Register r12 probably represents heap base pointer or similar in this case - I'm not sure.16-bit would be significantly slower, but for entirely other reasons. Also, you could have at most 512 kB heap with 8 byte address alignment...
The issue with 64-bit pointers is that Java (and JVM) is a very pointer happy language, so the cost is significantly more than for pretty much any other language I can think of. Memory usage and more cache misses is the issue.
https://wikis.oracle.com/display/HotSpotInternals/Compressed...