SharedHashFile
github.com
github.com
This is limited to the size of RAM, LMDB is not - LMDB is only limited to the size of the processor address space. This uses locks for reads, LMDB does not - LMDB reads will scale to an arbitrary number of CPUs, perfectly linearly, this will not.
Hash tables are great for small data sets, horrible for larger data.
Read performance without any writing does seem excellent at 6.x million reads per second across the 16 processes. However, insert speed is very slow at only 0.1 million inserts per second. But then LMBD does not claim to be fast at writing. Unfortunately the mix 2% update, 98% read brings the read performance down from 6.x million ops per second to only 0.7 million ops per second. So LMDB seems like an excellent solution if one hardly ever wants to insert.
I would also be very happy if anybody can find ways to optimize the test since I could not find a tutorial on programming with the LMDB API. For example, is it necessary to always use a txn when putting and getting? I also couldn't figure out how to insert 100 million keys because in order to make the map size big enough then mdb_env_set_mapsize() always complained when giving it super large values. Is this a limitation of LMDB or how else to increase the map size so that 100 million keys (or more) can be mapped? And as a side question: Inserting the keys is so slow: Is there a faster way to initially insert all the keys in order to speed up the performance test?
I was also surprised at how big the LMDB data.mdb file gets with 70 million keys & values. The keys & values are 4 byte + 4 byte, so 8 bytes each. However, the data.mdb file ended up as 1.8GB which works out to about 27.6 bytes per key,value pair... which does not seem that good compared to a hash table, or?
Also, please read the LMDB docs. Note that it is a single-writer design, so spreading writes across 16 processes will be quite poor.
(edit: Also, just realized that you're the guy who wrote lmdb. Nice meeting you good sir.)