JEP draft: 64 bit object headers
openjdk.org
openjdk.org
About 500mb of that is netty SSL session cache...
It's true that C/C++/Rust don't have this problem because they are optimized for this. OTOH if the problem you are solving is more suitable to garbage collection than manual or static-analysis based memory management then Java's high-performance GC will shine.
These are not just memory usage gains but also execution speed gains since memory access is frequently the main bottleneck and cpu cache effectiveness depends on how many of your objects they can hold.
Java is actually one of the few languages that does at least something in this field already before this JEP: Compressed object references (https://www.baeldung.com/jvm-compressed-oops).
Doing these kinds of analysis and optimizations takes time (and might impact language design), most "popular" languages before Rust got traction seldomly took memory models into accounts (still in the mid/late 90s lookup tables was still faster than doing certain kinds of "primitive" calculations that CPUs/GPUs eats for breakfast now). Bakers's Linear Lisp (a kind of precursor to Rust) and some StandardML that had region allocations, but that research was still mostly about tradeoffs in memory liveness.
The "data-oriented-design" movement in games has led to ideas of allowing structure-of-arrays instead of arrays-of-structures to be used in identical ways in a couple of languages, sadly many people involved in these spheres are allergic to GC's so it's still a bit of a narrow design field whereas I'm kinda curious about where layered designs could be taken.
Game programmers hit this on consoles due to facing platforms with relatively slow/tricky memory a bit earlier than PC devs (Cell processor[1]). But HPC programmers had by then been worryign and programming around the problem for longer, since the extinction of SSI vector supercomputers and transition to message passing clusters. (Also the data oriented design movement is only partly about this, and much about other C++ / programming / measurement etc ideas)
I was searching for older memory wall discussions and found at least this from 1994: https://www.eecs.ucf.edu/~lboloni/Teaching/EEL5708_2006/slid...
[1] in fact if you browse some of Mike Actons slides you can see his username on slideshare is "cellperformance" - eg https://www.slideshare.net/cellperformance/data-oriented-des...
I feel like that's a crazily-specific worst-case. Could have made it HashMap<Integer, Optional<Integer>> to really push the boat out.
Both GNU Trove and fastutil offers primitive collections that are still Java, but much better memory optimized.
And that’s not one-time work. In the ideal world, if the size of your data grows (or shrinks), you’ll have to revisit those choices.
Also, as a library author, you can’t know what choice to make.
Yes, you’ll always have that problem, but this change would make ‘just use the provided collections’ a much safer choice.
Usually you just pick fastuil or kobloke and then use it every time.
Edit: as comments below corroborate.
julia> k = Int32.(1:1000_0000); v = rand(Int32,1000_0000);
julia> dict = Dict(zip(k,v));
julia> Base.summarysize(dict) / (2 * 4 * 1000_0000)
1.8874391
julia> Base.summarysize(dict)
150995128Checking with a million values, on my Julia v1.8.0, the dict takes up 18 MB, or about 130% more than the plain values (or an array of Pairs) would.
Since you didn't actually give any actual argument apart from "I disagree and I am smart" I don't think we have anything concrete to argue about, so let's stick to my original point: would you agree that in C# it is much easier to do "the right thing" when putting primitives in a datastructure such as a list, map, or set, because C# has value types and does not use type erasure for its containers, thereby preventing both a lot of overhead and the need for specialized container types for various (combinations of) primitives?
I don't think I took a stance pro/anti erasure other than to say I think it's way more complicated a topic than it might seem (maybe not to you, but likly to others!) and it's worth reading up on. AND I didn't want to dig up old HN comments/posts to confirm what I'd read about it, so I was necessarily vague.
Edit 1: For instance, I vaguely recall that Scala head trouble porting to C# because of erasure related stuff. But I'm not genius enough to remember the details so I didn't get into it.
Edit 2: So I dug one up: https://cr.openjdk.java.net/~briangoetz/valhalla/erasure.htm... And posted here https://news.ycombinator.com/item?id=33171832
Edit 3: Other good conversations on it: https://news.ycombinator.com/item?id=18679684 https://news.ycombinator.com/item?id=13051594
Again, my only stance is that it's a surprisingly deep topic — Not making any claims!
Also there is already compression for that pointer, making it 32 instead of 64 bits.
But who cares? 99% of the time you just need class identity, not values from it.
> 99% of the time you just need class identity, not values from it.
So no need to deference the class pointer, or the table index.
Well also you didn't actually need to decompress the pointer first - compressed pointers are unique as well of course.
[1] in HotSpot they are instances of InstanceKlass: https://github.com/openjdk/jdk/blob/master/src/hotspot/share...
The term "stack-lock" is new to me. What's that? Is that just when the lock record is stored on the stack? In which case, what is being locked by a monitor, because i would expect the lock record to always be on the stack. Is this about thin vs inflated lock states?
It looks like the HotSpot JVM implementers decided to use 5 bits for lock and GC, and 31 bits for identity hash code (not sure why they didn't use 32). That adds up to 36 bits. You'd want to round this up to a multiple of 32 bits so that both the beginning of the object and the class pointer are aligned to 32 or 64 bits.
- GC state bits.
- The class metadata, i.e. pointer to a Class<> object.
- Lock data, including possibly a pointer to a lock.
- Possibly, an identity hash code.
- Possibly, a GC forwarding pointer.
Several of these are pointers! If you're on a 64 bit machine, just one* of them implemented naively if immediately 64 bits. In other words Java object headers provide a lot of stuff.
To pack the whole thing down to just 64 bits requires a lot of clever implementation tricks, exploiting what exact combinations can appear, etc. Also note that this overhead isn't Java specific. In C++ you have malloc headers and then the C++ vtable pointer, in many cases. So just running "new SomeObj()" in C++ will immediately blow you past the heap overhead of Java after this JEP is merged, and if you need to add a lock and/or reference count to catch up with the Java feature set then you're at multiples of it.
2. “Malloc headers”? Malloc usually stores metadata in freed data.
3. C++ developers don’t call “new” for every single object, the way Java devs do. Idiomatic code defaults to allocating fixed-size objects inline (on stack or in a pre-existing bigger allocation), not in a separate allocation.
4. In C++ there are other ways to allocate dynamic memory than “new”. Arena allocators, directly calling mmap/VirtualAlloc/whatever makes sense in a given case.
1. a pointer to the class
2. the object's lock
3. some GC bits
4. the identity hash code
In a naive implementation, the class pointer is a machine word, the lock is multiple machine words, the GC bits are a few bits, which get rounded up to a word for alignment, and the hash is 32 bits.
Object headers used to be even bigger than 128 bits!
There have been improvements: compress the class pointer, use a lazily-inflated lock which takes up only one word in the header, make the identity hash smaller, and make the lock and the identity hash share space (when the lock is deflated, the identity hash lives in the shared slot; when the lock is inflated, the slot has a pointer to a lock structure which also contains the identity hash).
The change in this JEP is effectively to extend the inflation-dependent space sharing scheme to include the class pointer too. When the lock is deflated, the header contains the class pointer and the identity hash; when it is inflated, it is a pointer to the lock structure, which contains the class pointer and identity hash.
I suppose this means that virtual method calls through a locked object will involve an additional indirection - but this should be rather easy for a compiler to optimise, because it can safely hoist the loading of the class pointer.
My guess is it is padding so that the class pointer is on an alignment boundary.
But in any case this will be yet another awesome perf improvement bringing jvm closer to cpp speed.
It also had support for a special kind of value types (packed objects) with are seen as intrisics by its JIT.
https://www.ibm.com/docs/en/sdk-java-technology/7.1?topic=po...
I guess this was eventually dropped due to Valhala efforts.