JEP 450: Compact Object Headers
openjdk.org
openjdk.org
WebAssembly GC objects do not have identity hash codes nor monitors to strive for the lowest space overhead in the base object model.
I can understand the hashcode choice, but I never understood why they chose to add the ability to lock on any object. IMO it still is a bad choice now on server machines, and it certainly was in a language designed for embedded devices in the early 1990s.
Does anybody know what they were thinking? Were they afraid of having to support parallel class hierarchies with a LockableFoo alongside Foo for every class? If so, why? Or did they think most programs would have very few objects, and mostly use reference types?
Ergonomics. (Plus an inherent assumption of mutable data.)
You can write `synchronized (this)` or `synchronized` class method, etc. and it just works.
- made a "monitor slot" its own primitive type;
- made two overloads of `synchronized` — synchronized(MonitorSlot), and synchronized(Object); where for synchronized(Object), the compiler expects to find a MonitorSlot-typed field with a special system name (e.g. "__jvmMonitorSlot") on the passed-in Object
- taken the presence of `synchronized (this)` in body code of a class to implicitly define such a field on the class, meaning you can "just write" `synchronized (this)`, since it becomes `synchronized (this.__jvmMonitorSlot)` and also triggers the creation of the __jvmMonitorSlot field
- explicitly define the __jvmMonitorSlot field on Class and other low-object-count system types
The only change from today would be that you can't point at some arbitrary Object you didn't define the class for yourself, and say "synchronize on that." Which... why do you want to be doing that again?
synchronized class Foo {
//...
}The only reason I added the `synchronized (this)` allowance was because the parent said that they think that that's "good ergonomics" — and presumably, the Java1.0 authors also thought that — and I was trying to suggest an alternative that would preserve what they consider "good ergonomics."
But personally, if I was the sole dictator of Java1.0 language design with nobody else to please, I would just have `synchronize(MonitorSlot)` + explicitly-defined MonitorSlot members (that, if you declare one, must always be declared final, and cannot be assigned to in the constructor), and that's it. Just refer to them by name when you need one.
Java was driven by the admission that multithreading is hard, and the hybris that this hard problem would go away once and for all if only you added enough OOP state hiding and locks. All those strategies we now use to allow parallelism were either unknown back then, or unfashionable, because they were in conflict with that OOP dream world of "objects talking to each other". The ease of just writing synchronized everywhere and declaring done on the topic of multithreading was as much a selling point of early Java as memory safety.
> Plus an inherent assumption of mutable data.
Yet, Strings and the boxed value types (Integer, Long, etc.) are immutable, and you can synchronize on them (or does this matter less for those because of the granularity of the memory allocator?)
Side note: That is a very bad idea, because their object identity is iffy.
It wasn't just Java. For many years THE argument for functional languages was you'd better learn them because these languages will soon be automatically parallelized, and that's the only way you'll be able to master machines with thousands of cores.
With hindsight we know it didn't work out like that. We got multi-core CPUs but not that many cores. Most code is still single threaded. We got SIMD but very little code uses it. We got GPUs with many cores, but they are programmed with imperative fairly boring C-like languages that don't use locking or message passing or anything else, they're data parallel pure functions.
But at the time, people didn't know that. You can see how in that environment of uncertainty "everything will be massively multi-threaded so let's give everything a lock" might have made sense
Because programming for SIMD is wildly different than MIMD. (And MIMD is what `synchronized` is for.)
The theory was that there would be more machines like AMD Thread Ripper.
For Wasm GC, I think we need programmable metaobjects to be able to combine language-level metadata with the engine-level metadata. I have only a prototype in my head for Virgil and plan to explore this in Wizard soon.
> For identity hash code, there is no guarantee there are fields to compute the hash code from, and even if we have some, then it is unknown how stable those fields actually are. Consider java.lang.Object that does not have fields: what’s its hash code? Two allocated Object-s are pretty much the mirrors of each other: they have the same metadata, they have the same (that is, empty) contents. The only distinct thing about them is their allocated address, but even then there are two troubles. First, addresses have very low entropy, especially coming from a bump-ptr allocator like most Java GCs employ, so it is not well distributed. Second, GC moves the objects, so address is not idempotent. Returning a constant value is a no-go from performance standpoint.
How else could identity hashcodes be implemented? Is it just impossible to put a WebAssembly GC base-object as a key into a map?
Another strategy, used by sbcl, is to rehash identity hash tables after each gc cycle.
How could I implement either of these with wasm?
Under this scheme, if I allocate object A at address 0, then GC-move it to address 100 (such that it caches that it was "originally at" address 0); and then, with object A still alive, I allocate object B at address 0... then don't objects A and B both now have 0 as their identity hashcode?
(I'm guessing the answer here is "this only works with generational GC, and the GC generation sequence-number is an implicit part of the stateless-address hashcodes and an explicit part of the cached-in-member hashcodes")
Caveat: There are multiple .NET implementations so what I've said may not apply to all of them
https://openjdk.org/jeps/8277163
.NET has something like that already, but it would make a huge difference for things like Optional or if you want to make a type for, say, complex numbers where there is just 64-bits of data (for FP32) and even a 64-bit object header would be 100% overhead. Value objects could be possibly allocated on the stack and completely bypass the garbage collector in some cases.
Note that with that overhead, Java is an environment in which you can write multithreaded applications with good reliability and scaling, something that WebAssembly definitely isn't.
I do think some of it is kinda interesting: some people might say "if you never do X, we can guarantee Y". But it's nice, as a programmer, to be told "if you use feature A, you will never do X, which means we can guarantee Y". It's a lot more comforting that I can't slip up and accidentally do X.
Indeed, this is why value semantics (i.e. structural equality) for language constructs like ADTs is so wonderful. Because a program can never observe "identity", which is an implementation detail, the implementation doesn't have to use objects at all underneath. That opens a whole host of value representation options that aren't otherwise available.
Maybe the programming style at the time had something to do with it. Maybe they thought every class needs to be thread safe and mutable.
[0] https://en.wikipedia.org/wiki/Lock_(computer_science)#Lack_o...
There's also the question of which hash code. Do you use a fast but low quality hash function, or a slower but higher quality one? Does your hash function need to be secure against hash collision attacks? Does it have to be deterministic? The correct choice of hash function can depend on the data structure, and the same object might have to be hashed with different hash functions (or different hash function seeds) in the same program.
My experience with platform design has consistently been that handling version evolution in the presence of distant teams increases complexity by 10x, and it's not just about some mechanical notion of backwards compatibility. It's a particular constraint for Java because it supports separate compilation. This enables extremely fast edit/run cycles because you only have to recompile a minimal set of files, and means that downloading+installing a new library into your project can be done in a few seconds, but means you have to handle the case of a program in which different files were compiled at different times against different versions of each other.
[1] https://github.com/titzer/virgil/blob/master/lib/util/Map.v3
This is partly my style too; I try to avoid using maps for things unless they are really far flung, and the things that end up serving as keys in one place usually end up serving as keys in lots of other places too.
Java doesnt make this very composable
Rust’s approach to the Hash and Eq problem is to make them opt-in but provide a derive attribute that autoimplements them with minimal boilerplate for most types.
Also, Rust’s Hash::hash implementations don’t actually hash anything themselves, they just pass the relevant parts of the object to a Hasher passed as a parameter. This way types aren’t stuck with just a single hash implementation, and normal programmers don’t need to worry about primes and modular arithmetic.
Fully separating implementation of an interface from the data can create quite a lot of additional complexities. See my comment elsewhere about encapsulation and version stability.
Sometimes you have no choice. I had multiple times where I needed to store additional data for instances of class X but had not control over it, so I had to store it in a separate structure and keep track of things by object identity.
> You either should write proper hash code
And object identity fulfills all the requirements of a "proper" hash code.
Further: default implementations should be synthesized by the runtime. Prevents subtle bugs when the equals(...) and hashCode() implementations are incorrect or missing. eg contract for HashMap.
To backfill, perhaps Object could extend new base class NakedObject, and implement your new interfaces Equalable, Hashable. Then classes which don't need identity can extend NakedObject.
Maybe the "canonical Object" was motivated by prior experience with Self or some such. It definitely calmed down people new to Java & OOP. I think it was the right call at the time.
Am I misremembering that there was a time where the JRE lazily added locks to objects? I thought it was part of their long road to lock elision.
A better decision would have been to make them available via interfaces, like CLR ended up doing in .NET 2.0, although they still follow Java/Smalltalk as well, due to its heritage.
Depends what you mean by "database driver", but what disqualifies XTDB or Datomic, both of which have use outside of Clojure yet are implemented entirely in Clojure with the exception of the Java wrappers so you can use them naturally within Java or Kotlin?
Even then, talking about "pure Clojure" is a non-starter because Clojure is not a self-hosted language (yes, Java is self-hosted if you consider Jikes RVM, a JVM implementation in pure Java). Even if Java ecosystem were to dwindle, people can rely on the JavaScript, .NET, or Dart ecosystems thanks to ClojureScript, ClojureCLR, and ClojureDart existing.
I use ClojureScript daily and am not really affected by what happens in Java land. With that said, I always welcome improvements to the JVM, because I actually like the Java ecosystem; it's the only one that doesn't make me split hairs...despite the dreaded module system.
If you're referring to the storage engines, those are not Java, and even then, most storage engines of a database will resort to a lower level language because of the need to care about memory layout and things that are simply not exposed to you in a high-level language like Clojure (or Java, unless you are somehow able to use Jikes RVM in production, I guess, but I'm not aware of ANYONE using a Java storage engine atop Jikes)
Granted, not everyone is like that and there are many nice folks, but those vocal ones, oh boy.
While great for future-proofing, it's been hard to deny it's had overhead. Glad to see it's slowly getting undone: Compressed Oops, Compact Strings, now this.
C and C++ usually do that too.
Interestingly, I can't find any information on Google about this, but ChatGPT does support this. Unfortunately all the primary sources it gave me are for docs that no longer exist on the web. The Wayback machine was no help. The old web is dead :(
Or perhaps that never existed. It's well known that ChatGPT often hallucinates non-existing references (see for instance the discussion at https://news.ycombinator.com/item?id=33841672).
It's unclear whether it's actually a 64-bit build. Did Windows even have a 64-bit userspace on Alpha, or was it all ILP32?
Anyway, it's not really possible to implement the Java memory model on Alpha: https://www.cs.umd.edu/~pugh/java/memoryModel/AlphaReorderin... So it's not really a natural target for Java code.
Ouch!
The weird inconsistency is apparent as you fairly frequently stub your toe on the 2 billion size limits of Java arrays, which is an area where 64 bit pointers would have made much more size than in object references.
No they weren't, not in the enterprise market. At home? sure.
SUN Microsystems released the Sparc v9 architecture in 1993. SUN also made the Java language and JVM.
There a was HotJava browser, weren't servers just a small part of its target?
There's Java Card for smartcards yes. There was also Java Micro Edition (j2me) for phones.
The situation that an x86-64 build is faster than an i386 build of most applications (exception extremely pointer-heavy ones) is a bit of an exception because x86-64 added additional registers and uses a register-based calling convention everywhere. That happens to counteract the overhead of 64-bit pointers in most cases. Other 64-bit architectures with 32-bit userspace compatibility kept using 32-bit userspace for quite some time.
64 bit pointers was not "optimistic assumptions" about anything, it was just 64 bit pointers on 64 bits systems like most everyone else. And compressed oops were added more than a decade ago (https://wiki.openjdk.org/display/HotSpot/CompressedOops).
For reference that's about when x32 was added to the linux kernel, and unlike x32 compressed oops have not been on the chopping block for 5 years.
> enormous object headers.
Hardly? They're two words a class pointer, and a "mark words" for GC integration.
That is still recent on a Java timescale.
> Hardly? They're two words a class pointer, and a "mark words" for GC integration.
Well compare with for example C++, that can fly commando with no header, just (optional) alignment padding.
In many Java objects, the header is more than half the size of the object. That's just not good data locality. The speed-up from switching from an array of objects that has n fields to an object which has an n arrays of fields can be very significant.
Edit: mind explaining the downvote ?
The Jacobin JVM [0] does exactly what you suggest: 64-bit operand stack and local variables. Longs and doubles still occupy two slots on the operand stack, so as to avoid having to recompile Java classes that assume the two-slot allocation, but the design avoids having to smash together two 32-bit values every time a long or double is operated on.
In retrospect, making 8-byte constants take two constant pool entries was a poor choice.But this needs to replace stack locking with an alternate lightweight locking scheme, to avoid races. Unfortunately, that is opaque:
https://bugs.openjdk.org/browse/JDK-8291555
Does anyone have other pointers for the design or viability of the required alternative?Reduce the size of object headers in the HotSpot JVM from between 96 and 128 bits down to 64 bits on 64-bit architectures. This will reduce heap size, improve deployment density, and increase data locality.
ByteBuffers are low overhead, but native arrays are ~30% faster last I benchmarked, so if I can place an array off-heap, and never move it (granted, it's unsafe), I can use for example, the kernel's paging mechanism to swap data in and out data from disk, with low runtime overhead to access it.