3 billion items in Java Map with 16 GB RAM
kotek.net
kotek.net
I've found JDBM3 a pleasure to use; it's fast and stable, with an excellent API. It's really quite easy to start using (it exposes Java collection interfaces), but you can configure it to do some powerful things under the hood. I'm not using the library in a really high-performance setting (just need a persistent key-value store), so I can't quite comment on the extreme scaling qualities.
I'm excited to switch to MapDB as it matures.
I'm on an old XP machine with very little memory and a slow disk, and I get quite good performance out of it, despite not being able to effectively use mapped or off-heap memory. I expect that moving to MapDB on JDK7 (with its concurrent collections and better memory allocation) will improve performance quite a bit.
"What you should know: MapDB relies on mapped memory heavily. NIO implementation in JDK6 seems to be failing randomly under heavy load. MapDB works best with JDK7"
The original JDBM2 apparently worked fine under Android; I wonder how the new implementation (JDBM4 renamed MapDB) fairs.
I do worry about Java as implemented on Android diverging from the official stuff in general. For example, Android does not have NIO.2, and Harmony NIO bugs would be different than OpenJDK NIO bugs.
> If this collection contains more than Integer.MAX_VALUE elements, returns Integer.MAX_VALUE.
Can someone explain why this happens?
EDIT: From rereading the original post it seems more likely that it was approaching the heap size limit set when running the JVM and not a physical memory limit. I would expect the same principle applies, though.
* http://www.drdobbs.com/jvm/g1-javas-garbage-first-garbage-co...
* http://www.oracle.com/technetwork/java/javase/tech/g1-intro-...
Note also that the JRE ships with multiple garbage collectors that you can select manually, and that some of the collectors tune themselves to match the application.
From the concurrent mark and sweep docs: "The concurrent mark sweep collector, also known as the concurrent collector or CMS, is targeted at applications that are sensitive to garbage collection pauses. It performs most garbage collection activity concurrently, i.e., while the application threads are running, to keep garbage collection-induced pauses short. The key performance enhancements made to the CMS collector in JDK 6 are outlined below. See the documents referenced below for more detailed information on these changes, the CMS collector, and garbage collection in HotSpot."
It does have to stop the world sometimes but those times should be quite rare!
CMS actually isn't the highest throughput GC, it's there for low-pause systems where there are extra cores available to perform GC in the background.
If you want stricter failure requirements use GCTimeLimit and GCHeapFreeLimit. By default the JVM with throw an OOM error if it spends 98% of execution time in the GC and frees only 2% of the memory, GCTimeLimit and GCHeapFreeLimit switches allow this these values to be changed. Also only the 'stop the world' portions of a collection apply to the execution time limit, concurrent phases do not. So if you managed to keep the collector in concurrent mode (by cranking up the duty cycle for example) then only the 2 mini-pauses will count even if the concurrent portion were now pegging a CPU core at 100%.
Before you say "well that's stupid", realise that the alternative would throwing an OutOfMemoryError, in which case the JVM would terminate. A slow JVM is better than no JVM.
Regardless, since Java 6 (I think), if the GC time becomes excessive (where "excessive" means x% of time is spent in GC and less than y% of memory is reclaimed in each GC cycle), an OutOfMemoryError will be thrown. This can be disabled.
This will cause OOM-exception to be thrown when at least 98% of program time is used in garbage collection and at most 2% of the total memory is reclaimed in such collections.
Indeed, you can do this; the JVM has simple facilities to report the memory use and heap status, which you can monitor and use to proactively scale around the bogged-down VM.
[1]http://www.oracle.com/technetwork/java/javase/tech/vmoptions...
Even when you reach the Xmx the VM will try to GC, and sometimes it can free a bit of memory so it doesn't throw an OOM, insert another 10 items and run the GC again, and so on - as long as you allocate small objects, it can take a while until you see the OOM.
It is basically testing constant rehashes which is obviously going to generate a ton of garbage
I suspect the default collection implementations would perform significantly better if they were used correctly
https://oss.sonatype.org/content/repositories/snapshots/org/...
Previous versions were called JDBM and it has been around since 2001, so it has some solid user base. Problem is that nobody tells me until there is bug, which is not happening that often :-)