Tracking Java native memory with JDK flight recorder
morling.dev
morling.dev
* https://www.morling.dev/blog/finding-java-thread-leaks-with-...: Discusses how to find thread likes with JFR and JFR Analytics, a project I've created for querying recordings with SQL
* https://www.morling.dev/blog/towards-continuous-performance-...: Discusses how to use JFR for continuous performance testing, by means of asserting "proxy metrics" such as allocation rates and IO
* https://www.morling.dev/blog/rest-api-monitoring-with-custom...: Discusses how to create your own application-specific JFR events
If you're using glibc, then malloc does have information about that, and provides ways to read it, so it's a shame this isn't exposed. It would be quite helpful in the face of suspected native library memory leaks.
$ jfr view native-memory-reserved rec.jfr
and $ jfr view native-memory-committed rec.jfr
https://x.com/ErikGahlin/status/1736530559231201484A. it tracks native calls by default B. it can track wall time as well C. you can have neat interactive flamegraphs
It should all be in jfc/jmc
In the meantime you can get async profiler to output in JFR format and combine it with a separate JFR recording.
Memory mapped files and direct buffers do work well. I'd never consider it a bad idea; technically both would be freed if there are no strong references to them (not much different than allocating massive arrays).
Think of it like that the jdk standard library has to use native memory for any IO or even a trivial zip compression. ByteBuffers were introduced in 1.4 (around 20y back) to allow operations outside the managed system, the tools are available to anyone, e.g. implementing zstd with direct buffers is not hard.
I saw a pretty jawdropping speedup moving from bytebuffers to the new foreign memory stuff with in the marginalia search index code. Most operations run about 50-100% faster, even discounting the rigmarole needed to deal with mmapping >2 GB files. Haven't dug too deep into why this is so I'm not entirely sure why, may have something to do with being able to declare non-shared memory ranges in the arena allocator, saves a bunch of synchronization maybe?
I can't wait for this stuff to leave experimental. It's such a quality of life boon for dealing with off-heap memory, having explicit lifecycle control is.
I'd disagree. Heap based bytebuffers - I don't consider them interesting, they are just byte arrays. The direct ones suffer from unmapping (munmap on linux) - it's a slow process that has to flush the TLB. So the only sane way to use them allocate once/reuse. Anything else I'd consider a programmer error.
mmap on large files indeed does suck (a bit) w/ ByteBuffers as it'd require an array of them.
Personally I have been using direct buffer since 1.4, so I guess I am also quite used to them as well.
It should never be an issue in native code, of course. But yes - it's possible that it happens with Java code. In that case you'd need print assembly and looking at the generated code.
Some libraries have 'switched' to straight unsafe use (which has no bound checks, of course), so there is that.
That's just a sign of bad benchmark. Off-heap BBs and foreign memory are both native (non-JVM) heap. They are the same thing
However if this value ever does get exceeded, you need some way of tracking down what allocations happened prior to your OutOfMemory exception.
Though the above article implies the sampling rate is once a second (I guess there's some cost to increasing that rate). Usually you won't be allocating direct memory on the reg since it's expensive to allocate and deallocate relative to heap memory so you kind of want to capture ALL allocations and deallocations. As such a sample based approach is not ideal due to possibly missing some data between samples.
What made you decide to go this route, instead of pre-allocating a pool of buffers (possibly thread-local) and recycling them?
For the specific use cases I have worked on though, I'm not sure it would have had any additional performance benefits. The specific class of systems I've worked on usually have only a few sockets which open at the start of a day and stay connected until the end of a day. Pre-allocating would make the socket opening process faster but that is not usually the part of an application life cycle that needs optimising. If there were lots of sockets opening and closing or the speed of opening and closing needed optimisation, then what you suggested would be a good idea.
It's also worth noting that if your Xmx is larger than 32 GB, you can't use CompressedOOPs, which is a bad deal. Off-heap memory lets you skirt that limitation while still allocating >32GB.
I guess what you'd want to use this for, is when your application is directly allocating memory, e.g. via direct byte buffers. That's not something you'd do in an enterprise application, it's more something you'd need for high performance image processing, or maybe for some extremely high performance web server.