JVM Anatomy Quarks
shipilev.net
shipilev.net
Thanks to much of his writings and the rest of the Java performance community (pretty large in the fintech sector), I can write faster Java than most people can with C++. It just takes some effort, but the control and performance you can get from Java is really impressive.
I've had to deep dive into Go a little more lately, and I really miss some of the Java support. I've found Go to be much slower when you have to do anything interesting. In high performance Java you often rewrite a lot of the base libraries in a very different style that gives you tight control over escape analysis, GC, call site inlining, etc. You actually have a decent amount of control for such a high-level language.
In Go, I haven't been able to find that control. The Go team seems to have taken an opposite approach and removed your control (I often joke about Go just being short for "Go Fuck Yourself" because of its attitude against developer control and the teams's "if we don't need it you don't need it" attitude).
It is resources like this that really make Java shine in its pro high performance developer attitude. (Current Go issue, getting select and channels to operate anywhere remotely efficiently and trying to find a way to keep high CPU goroutines on different OS threads - so far not much luck).
As usual, one platform kind of mirrors the other. :)
I'm assuming you're aware of the microbenchmarking framework built into Go's testing framework? If so, can you elaborate on where it falls short? I have my own gripes, but it would be nice to understand where others are coming from on this.
> In high performance Java you often rewrite a lot of the base libraries in a very different style that gives you tight control over escape analysis, GC, call site inlining, etc. You actually have a decent amount of control for such a high-level language.
In case you're not aware, the standard Go toolchain does expose some of the optimizations the compiler does, notably escape analysis and bounds check elimination. In the latest release you'll have the option to even export this information as JSON to inspect it programmatically.
Other things, like inlining, are not there (AFAIK), but that's partially because the toolchain is still maturing. For example mid-stack inlining (i.e. inlining any calls to non-leaf functions) is actually a relatively recent addition. Go will also let you choose to not inline something. There's still a lot of work to be done around inlining in general, and visibility into the process would be a nice improvement.
I'd also like to point out that part of the reason the GC has so few knobs has to do with maintainability. Every new knob means expanding the space of configurations significantly, and making sure they all continue to work is a big task. For example, V8 exposes a lot of knobs, but (IIUC) aside from a few default configurations shipped in Chrome, any deviation from those and you're considered "on your own", mainly because of this maintainability problem.
With that being said, I'm not really sure how OpenJDK deals with this issue; maybe there's just enough people out there and enough resources behind the project that it's fine?
> In Go, I haven't been able to find that control. The Go team seems to have taken an opposite approach and removed your control (I often joke about Go just being short for "Go Fuck Yourself" because of its attitude against developer control and the teams's "if we don't need it you don't need it" attitude).
I don't think this is the intended messaging from the Go team, and there's been efforts on their part to shed this image. I think part of it is maturity of the toolchain; Java has nearly 15 years on Go and in some cases there honestly isn't all that much to give visibility into or control over yet. Another part of it is an overall conservative approach toward evolution of the language and of APIs, primarily for long-term maintainability and compatibility. Expanding the API (including performance knobs) usually needs to show a clear net win (see SetMaxHeap, which gives you more control but never really made it in; it still exists as a patch).
> It is resources like this that really make Java shine in its pro high performance developer attitude. (Current Go issue, getting select and channels to operate anywhere remotely efficiently and trying to find a way to keep high CPU goroutines on different OS threads - so far not much luck).
You should definitely file a bug if you have the bandwidth to do so and haven't already. Channels and scheduling should be efficient and smart by default, and finding situations where they make poor decisions is how the runtime improves. The team is fairly responsive to such bugs and even if they don't get resolved immediately, having it on their radar will only help the team make better design decisions going forward.
It's a fairly fascinating window into the inner workings, and quite detailed.
Filip Pizlo’s site has a bunch of his presentations of how their JIT compiler works http://www.filpizlo.com/papers.html
even CHM uses synchronized nowadays.
Flip note: don't touch ReadWriteLocks
I am interested in references that back this statement
[0] https://blog.overops.com/java-8-stampedlocks-vs-readwriteloc...
- the lock has write a CAS on the =fast= read path, causing coherency traffic and a contention point between the readers. That's it the readers don't scale
- it's quite hard to use correctly, i.e. after read, determining the exclusive/write lock has to be acquired, the read lock has to be released 1st, the write lock acquired and the conditions that causes the grab to be rechecked
- Copy-On-Write should be a preferred solution for most cases, easy to understand and reason about. If not StampedLock is a better alternative.
> In Project Loom, there will be support for efficiently switching between fibers that use Java 5 locks (that is, the java.util.concurrent.lock package) but not native C locks. As a result, it is necessary to migrate all blocking code in the JDK over to Java 5 locks. So the legacy Socket API required reimplementation to achieve better compatibility with Project Loom.
See: https://blogs.oracle.com/javamagazine/inside-java-13s-switch...
If you really need this sort of dynamic locking, you can use a ConcurrentHashMap<String, Object> to achieve it (lock on the object). I'm not sure whether it's ever the _best_ design, but it avoids interning the string, and it keeps an anonymous lock object that you know won't be shared.
It's leak prone as it'd hold the references forever. It's a cheap and easy way to do it but far from ideal. I'd not recommend it.
Isn't that the point?
People use it to create symbolic locks in situations where they don't want to use any more formal link of linkage provided by the JVM. In your case you need some way to get a handle to that concurrent hash map, so some kind of formal JVM linkage. Sometimes that's hard.
Not saying how it's how I'd chose to design an application, but I presume people doing this have their own good reasons.
Now, if your string is sufficiently unique, the chances are relatively low, as long as you don't leak a reference to it (an object, or explicit Lock only has one purpose, so that's less likely).
Still, it's basically mysterious action at a distance. Whereas using the ConcurrentHashMap, it's very explicit action at a distance. Granted, it does require explicit JVM linkage, which is a cost.
It's hard to explain how terrible the idea to synchronize/lock objects you don't control.
As for interning (not for strings only and not guaranteed) I have a lock free table (not CHM) that keeps most used objects (with possible random eviction) to provide a good trade off between unnecessary memory waste, fast access and low memory footprint.
> what happens if someone changes that field?
The field is final. The static is final. How can the field be changed? Reflection?
The static reference KNOWN_M is final, and the only field of the pointed-to instance, x, is also final.