MDBM – High-speed database
yahooeng.tumblr.com
yahooeng.tumblr.com
mdbm performance is even better on FreeBSD than Linux because FreeBSD supports MAP_NOSYNC, which causes the kernel not to flush dirty pages to disk until the region is unmapped. Perhaps mdbm's release will finally get the Linux kernel team to provide support for that flag.
There's a mmap flag on Linux called MAP_LOCKED but I'm not sure how it behaves with MAP_SHARED, which mdbm uses (the man page isn't clear).
MDBM is pretty much an optimized persistent hash table. LMDB and WiredTiger aim to be full-fledged ACID compliant database storage engines with functionality similar to that of BerkeleyDB or InnoDB.
I agree that it's an apples to oranges comparison in any case.
It's an apples-to-oranges comparison only of MDBM wins significantly against LMDB. If they are comparable in timing, or e.g. MDBM is 20% faster, then it would be an apples-to-apples comparison, MDBM having 20% speed advantage, and LMDB having every other possible advantage (memory safety, ACIDity, ordered retrieval, multiple databases, etc.)
LMDB is truly, incredibly, really marvelous. On 64-bit it comes close to being the end-all-be-all local KV-store. If your databases are not more than a few tens of megs each, the same is true for 32-bit processors as well.
Why is that ? Shouldn't 32-bit processors give you enough space in the range of hundreds of MiB ?
I should note that in LMDB 1.0 we'll have dynamic unmapping and remapping, to allow 32-bit machines to work with larger DBs. (There's still a significant performance cost for this. It's only being done to allow folks to use the same code on 32 and 64.)
I haven't looked closely at Cassandra since it's in java, and after I didn't find a simple backend plugin API I didn't look any further.
https://github.com/jbooth/flotilla
Basically I'm just layering the raft consistency algorithm on top of LMDB. Both systems single-thread write transactions for consistency, so there's some mechanical sympathy. Doesn't mandate any specific data model or even a client-server network transport, it's basically a replicated embedded DB. Anyone could build replicated redis on top of it or a Cassandra clone if they want to get into managing shards/rings.
Sample app (still WIP) at https://github.com/jbooth/merchdb
From my totally biased perspective, MDBM is utter garbage. They use mmap but make absolutely zero effort to use it safely. This was the biggest obstacle to overcome in developing LMDB; I had a few lengthy conversations with the SleepyCat guys about it as well. It's the reason it took 2 years (from 2009 when we first started talking about it, to 2011 first code release) to get LMDB implemented. If you want to call something a "database" you have to do more than just mmap a file and start shoving data into it - you have to exert some kind of control over how and when the mapped data gets persisted to disk. Otherwise, if you just let the OS randomly flush things, you'll wind up with garbage. As Keith Bostic said to me (private email):
"The most significant problem with building an mmap'd back-end is implementing write-ahead-logging (WAL). (You probably know this, but just in case: the way databases usually guarantee consistency is by ensuring that log records describing each change are written to disk before their transaction commits, and before the database page that was changed. In other words, log record X must hit disk before the database page containing the change described by log record X.)
In Berkeley DB WAL is done by maintaining a relationship between the database pages and the log records. If a database page is being written to disk, there's a look-aside into the logging system to make sure the right log records have already been written. In a memory-mapped system, you would do this by locking modified pages into memory (mlock), and flushing them at specific times (msync), otherwise the VM might just push a database page with modifications to disk before its log record is written, and if you crash at that point it's all over but the screaming."
The harsh realities of working with mmap are what dictated LMDB's copy-on-write design - it's the only way to ensure consistency with an mmap without losing performance (due to multiple mlock/msync syscalls). None of these design considerations are evident in MDBM.
LMDB's mmap is read-only by default, because otherwise it's trivial to permanently corrupt a database by overwriting a record, writing past the end, etc. MDBM's mmap is read-write, and the only "protection" you get is a doc that tells you "be Vewwy vewwy careful!" Ridiculously sloppy.
LMDB's design and implementation are proven incorruptible. MDBM (and LevelDB and all its derivatives) are proven to be quite fragile. https://www.usenix.org/conference/osdi14/technical-sessions/...
Leaving reliability aside for a moment, there's also the issue of performance and efficiency. We used to use DBM-style hashes for the indexes in OpenLDAP, up to release 2.1. We abandoned them in favor of B-trees in OpenLDAP 2.2 because extensive benchmarking showed that BDB's B-trees were faster than its hash implementation at very large data sizes. The fundamental problem is that hash data structures are only fast when they are sparsely populated. When the number of data records you need to work with increases to fill the table, you start getting more and more hash collisions that result in lots of linear probes (or whatever other hash recovery strategy you're using). The other problem is that the very sparse/unordered nature of hashes makes them extremely cache unfriendly - you get zero locality-of-reference for groups of related queries. So as your data volumes increase, you get less and less benefit from the amount of RAM you have available. When the data exceeds the size of RAM, the number of disk seeks required for an arbitrary lookup is enormous, and every read is a random access. Using a hash for a large-scale data store is just horrible. (We tested this extensively a decade ago http://www.openldap.org/lists/openldap-devel/200401/msg00077... )
The benchmarks are obviously for our own benefit too - until someone does these comparisons, none of us knows where things truly stand.
Among other things, I like that LMDB has zero-copy reads and that's something I've taken care to preserve all the way through my layers.
Just wanted to say thanks for the great work. LMDB is a joy to work with.
Do you have any idea if a sqlite 4 release is imminent? Will lmdb work with it right out of the gate?
Thanks.
I'm going to guess that they will not ship an LMDB driver right out of the gate. The one we were working on was not completed (our contractor flaked), and while I know they did some work on their own, I have no idea how complete that was either.
Didn't bdb's linear hashing scheme extend the size of the hash table enough to keep it at the required loadfactor?
Our experience with it shows that resizing was itself a very expensive operation.
If you want I'll go shove a few GB into an mdbm, drop caches, and time a lookup.
2 seeks at the most, are you talking about a 32 bit address space? The only way that's possible in 64 bits is to direct map a hash into e.g. 2 32 bit chunks and use the hash as an actual disk block address for the first chunk, and an index into a block list for the 2nd chunk.
Not only that, we watched the bus on an SGI Challenge and counted cache misses and TBL misses. 2 TBL misses to get a key.
Saying that it isn't possible on a 64 bit VM system makes no sense to me. If I have a 2TB file and I seek to location A and read it, then seek to location B and read it, you are saying that's not possible? Same thing with mmap, I set a pointer to the mapping, read p, p += <number>, read p. Two seeks, two page faults, whatever you want to call it, it does 2 and only 2 I/O's to get a key/value (unless the pages are bigger than disk blocks but then those are going to be sequential I/O's, no extra seeks).
Anyway, I don't doubt that you can operate in 2 seeks in the normal case.
mdbm is certainly not without limitations, but is careful about its use of mmap to an extent that comparisons with MongoDB are laughable.
There are a number of use cases where in the event of a node failure it is better to rebuild from a replica or a log. Statistically, the RAM on another host is actually more reliable than local storage. Additionally, the database does have sync'ing primitives that allow for a variety of persistence strategies... just not the traditional ACID strategy.
In practice, there are lots of cases where the freedom to ignore transactional integrity is very handy, and yes, a cache would definitely be one of them.
So in other words, you're saying "MDBM is not a persistent database." Glad that's clear. Totally agree, there are probably lots of use cases for it. But persistent data store isn't one of them.
I'm guessing you are one of the people behind some other technology. Goody for you but do you really think you make your case by dissing anything else? If you have a solution that works for you, great. SGI, Yahoo, and other companies have found a use for MDBM. SGI was using it ~20 years ago and at the time there was nothing that came close to the same performance.
The only FUD here is advertising a piece of software as a high performance embedded database when in fact it's not suitable for such use on its own. The most viable use cases the authors have presented is when using MDBM as part of a larger distributed system such that the loss of a single DB instance isn't fatal. The above comment talks about restoring from a log, but the actual log mechanism isn't part of MDBM. I.e., MDBM is incomplete on its own and you must provide additional pieces in order to use it effectively.
I'm not solely interested in promoting my own DB. If you read the On-Disk microbenchmarks I linked you'll see that there's a broad range of use cases where LMDB gets trounced by LevelDB and other LSMs. I'm interested in facts.
From the original link:
"On clean shutdown of the machine, all of the MDBM data will be flushed to disk. However, in cases like power-failure and hardware problems, it’s possible for data to be lost, and the resulting DB to be corrupted. MDBM includes a tool to check DB consistency. However, you should always have contingencies. One way or another this is some form of redundancy…"
The fact is, this is a system that can lose data on a crash and it doesn't include its own recovery mechanism. Without such a mechanism you can't call it a persistent data store because MDBM by itself is not persistent.
Persistence != ACID
By your definition, ext2 is not persistent. Give it a break.
Sounds like you are shooting off without really understanding what you are criticizing.
You really started posting on this comment thread without reading the title of the article?
So you are spouting all your disdain based on a title? Just curious, have you even run any version of mdbm? Or are you just spitting out baseless opinions?
I'm sure your code is awesome and all but man, you are defensive. If your code is that great let it speak for itself. Bashing stuff that isn't even trying to compete with you makes you look pretty insecure. You remind me of me when I was younger; that's not a compliment, I was a raging asshole.
We used MDBM in OpenLDAP for a few years on SGI Irix. I haven't touched it myself in something like 9 or 10 years though. http://www.openldap.org/lists/openldap-devel/199903/msg00094...
http://www.openldap.org/lists/openldap-devel/200501/msg00053...
Using DBM-style DBs was an endless nightmare of corruption bugs. That technology belongs firmly in the distant past, we have better solutions today.
http://www.openldap.org/lists/openldap-software/200607/msg00...
I'm not bashing MDBM because it's a competitor; it obviously isn't a competitor. Heck the only reason I'm commenting in this thread is because you asked me to elaborate on its design flaws. Now you're hurt that I answered your question. Don't ask if you don't want to know the answer.
... seems like us folks in OpenLDAP weren't the only ones to have bad experiences with it. This guy in this same discussion seems to share the basic sentiment. https://news.ycombinator.com/item?id=8733819
I'm sorry the openldap couldn't figure out how to use MDBM in a way that worked for them. Yahoo clearly has. We have. Others have. It works extremely well for what it was designed for and the perf numbers show that it spanks the living crap out of your stuff. But your stuff does more, which is cool. Why don't you focus on that rather than trying to crap all over some useful code? It's clearly not a threat to you. I don't see how the world is well served by your comments. For certain problems MDBM is way the hell better than what you have. It's a narrow niche, doesn't compete with you, so why all the fuss?
As for me asking you to comment, yup, I did, but you were already well on your way of banging on technology you clearly didn't understand. I think I get why it didn't work for you, your comments have made it clear you have no idea how it works. I went and read all the links you provided above, not one mentioned mdbm having corruption bugs. It's entirely possible that the other DBM style dbs had bugs, or it is possible that people were inserting / deleting in a first / next loop, whatever. But pointing a finger at MDBM and claiming corruption bugs, how about you substantiate that claim? Seems somewhat flawed when multiple companies have used it for 10 years or more and it seems to work fine for them.
"Spanks the living crap" - must be some heretofore unknown definition of "spanks".
./db_bench_mdbm
MDBM: version 4.11.1
Date: Mon Dec 15 04:39:20 2014
CPU: 4 * Intel(R) Core(TM)2 Extreme CPU Q9300 @ 2.53GHz
CPUCache: 6144 KB
Keys: 16 bytes each
Values: 100 bytes each (50 bytes after compression)
Entries: 1000000
RawSize: 110.6 MB (estimated)
FileSize: 62.9 MB (estimated)
------------------------------------------------
fillrandsync : 40.627 micros/op 24614 ops/sec; 2.7 MB/s (1000 ops)
65604 /tmp/leveldbtest-1000
fillrandom : 16.356 micros/op 61137 ops/sec; 6.8 MB/s
122056 /tmp/leveldbtest-1000
fillrandbatch : 5.499 micros/op 181850 ops/sec; 20.1 MB/s
121936 /tmp/leveldbtest-1000
fillseqsync : 40.163 micros/op 24898 ops/sec; 2.8 MB/s (1000 ops)
65604 /tmp/leveldbtest-1000
fillseq : 16.424 micros/op 60886 ops/sec; 6.7 MB/s
175724 /tmp/leveldbtest-1000
fillseqbatch : 5.648 micros/op 177041 ops/sec; 19.6 MB/s
175724 /tmp/leveldbtest-1000
overwrite : 16.290 micros/op 61385 ops/sec; 6.8 MB/s
175724 /tmp/leveldbtest-1000
readrandom : 0.598 micros/op 1672444 ops/sec; (1000000 of 1000000 found)
readseq : 0.096 micros/op 10370216 ops/sec; 1147.2 MB/s
./db_bench_mdb
LMDB: version LMDB 0.9.14: (September 20, 2014)
Date: Mon Dec 15 04:41:37 2014
CPU: 4 * Intel(R) Core(TM)2 Extreme CPU Q9300 @ 2.53GHz
CPUCache: 6144 KB
Keys: 16 bytes each
Values: 100 bytes each (50 bytes after compression)
Entries: 1000000
RawSize: 110.6 MB (estimated)
FileSize: 62.9 MB (estimated)
------------------------------------------------
fillrandsync : 12.818 micros/op 78015 ops/sec; 8.6 MB/s (1000 ops)
224 /tmp/leveldbtest-1000/dbbench_mdb-1
224 /tmp/leveldbtest-1000
fillrandom : 4.275 micros/op 233923 ops/sec; 25.9 MB/s
116548 /tmp/leveldbtest-1000/dbbench_mdb-2
116548 /tmp/leveldbtest-1000
fillrandbatch : 3.490 micros/op 286502 ops/sec; 31.7 MB/s
126384 /tmp/leveldbtest-1000/dbbench_mdb-3
126384 /tmp/leveldbtest-1000
fillseqsync : 14.972 micros/op 66791 ops/sec; 7.4 MB/s (1000 ops)
172 /tmp/leveldbtest-1000/dbbench_mdb-4
172 /tmp/leveldbtest-1000
fillseq : 2.231 micros/op 448145 ops/sec; 49.6 MB/s
125872 /tmp/leveldbtest-1000/dbbench_mdb-5
125872 /tmp/leveldbtest-1000
fillseqbatch : 0.425 micros/op 2355457 ops/sec; 260.6 MB/s
125872 /tmp/leveldbtest-1000/dbbench_mdb-6
125872 /tmp/leveldbtest-1000
overwrite : 4.881 micros/op 204857 ops/sec; 22.7 MB/s
125872 /tmp/leveldbtest-1000/dbbench_mdb-6
125872 /tmp/leveldbtest-1000
readrandom : 1.166 micros/op 857624 ops/sec; (1000000 of 1000000 found)
readseq : 0.059 micros/op 17092556 ops/sec; 1890.9 MB/s
readreverse : 0.042 micros/op 23814626 ops/sec; 2634.5 MB/s
Feel free to submit a patch if I got anything wrong in that driver; it was a pretty hasty patch. This is running on tmpfs, so no I/O involved. MDBM is faster on random read, which is what you'd expect since it's a hash and doesn't have to navigate down a tree to locate a record. Aside from that, it's pretty pedestrian.Hmm... I think that might alter the results actually.
> which is what you'd expect since it's a hash and doesn't have to navigate down a tree to locate a record.
mdbm uses a hash to select a page, but it actually does store the keys within a page in order. It's kind of a funky mix.
Test MDBM LevelDB KyotoCabinet BerkeleyDB
Write Time 1.1 μs 4.5 μs 5.1 μs 14.0 μs
Read Time 0.45 μs 5.3 μs 4.9 μs 8.4 μs
Sequential Read 0.05 μs 0.53 μs 1.71 μs 39.1 μs
Sync Write 2625 μs 34944 μs 177169 μs 13001 μs
As to losing credibility, I'm semi retired, I stopped trying to impress people a decade ago. If i were you I'd be more worried about your own image, bad mouthing other people's tech when you demonstrate you don't understand it hasn't made you like good to at least a few people here.It wouldn't be that hard to say "MDBM is great when used as an index into a DB, it works just fine for that. But you are going to have to rebuild the index after a power failure unless you take care to flush the data. It's somewhat unfair that the OP compared against my database because mine is slower because it handles crashes."
Instead you come out with "MDBM is complete crap". Well, no, it's not. In the domain where it is useful it is actually quite useful, it's 10x faster for lookups than your DB. So the trade off is speed vs surviving reboots. For lots of people, speed is much more important. Machines don't crash every ten minutes. In fact, it's pretty common to see uptimes in 100's of days. Lets say that 100 days is average and lets say that it takes a full day to rebuild the MDBM. So MDBM is delivering 99/100 days of useful work. You deliver 100/100 days. Oh, wait, except that your useful work is running 10x slower if the DB is being used as an index. So you delivered 10/100 days. See why some people may prefer to use MDBM when it is put like that? Performance is a feature.
I think that statement is absolutely true. I just think you have a myopic view about how to achieve performance and reliability. Look at Yahoo's example use cases.
Look at the results I posted again, and look at the benchmark code. Tell me that I've made a mistake, that's fine. The thing about open source is there's no reason to BS, anyone can build and run it and see for themselves. LMDB is faster and its data is more compact than MDBM, so you can get more work done using less resources, and you don't have to worry about losing your work for unexpected downtime.
Hashes suck for large volumes of data. That's just the reality of it, plain and simple: Low storage efficiency, memory-intensive, and cache-unfriendly. Whether people like or dislike me personally for saying so doesn't change the facts.
As for uptime - sure, and my PCs have uptimes for hundreds of days too. But I'd be a fool to just take that for granted and not take regular backups. The problem with your 99/100 days math is that you can't actually account for the cost of a crash that way. It might only set you back to 0/100; if you're unlucky it will set you back to -100 or more.
Any idea why it is that much better than LMDB?
I actually agree with you in that hashes sort of suck, just look at Git when the repo gets big, a hash is a miserable way to traverse all that data, very cache unfriendly.
But the use cases for MDBM are the same, you have lots of keys and you want to get to any key very quickly. It appears to me that it (still, 20 years later) wins that race.
Yes, for exactly the reasons you'd expect:
mdb_stat /tmp/leveldbtest-1000/dbbench_mdb-6/
Status of Main DB
Tree depth: 4
Branch pages: 204
Leaf pages: 31250
Overflow pages: 0
Entries: 1000000
As you said, MDBM can find any record in 2 seeks; for this database the LMDB tree height is 4 so any random access takes 4 seeks. 2x perf difference.On a larger DB we would expect MDBM's random read perf advantage to get larger as well, until the DB exceeds the size of RAM. I've tried to duplicate my http://symas.com/mdb/ondisk/ tests with MDBM but it makes XFS lose its mind by the time the DB gets to 2x the size of RAM. First I had to increase the MDBM page size from my default of 4KB to 128KB, otherwise I'd see a lot of this in the output:
2014/12/16-13:13:29 ... thread 0: (200000,12200000) ops and (3436.0,8072.2) ops/second in (58.206750,1511.358831) seconds
3:54903002:dc057:00864 mdbm.c:1809 MDBM cannot grow to 33554432 pages, max=16777216
but on this VM with 32GB RAM, every time the DB hit 60GB in size the kernel log would start getting spammed with Dec 16 22:19:19 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
Dec 16 22:19:21 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
Dec 16 22:19:23 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
So thus far I've been unable to load the 160GB DB to reproduce the test. But this underscores the basically uncacheable nature of hashes - once the directory gets too big to fit in RAM, your "2 seeks per access" goes out the window because the kernel is thrashing itself trying to keep the whole directory in-memory while finding the requested data pages.B+tree performance degrades gracefully as data volumes increase. Hash performance falls off a cliff once you cross the in-memory threshold.
MDBM: version 4.11.1
Date: Tue Dec 16 20:34:36 2014
CPU: 16 * Intel(R) Xeon(R) CPU E5-4650 0 @ 2.70GHz
CPUCache: 20480 KB
Keys: 16 bytes each
Values: 2000 bytes each (1000 bytes after compression)
Entries: 76800000
RawSize: 147656.2 MB (estimated)
FileSize: 74414.1 MB (estimated)
------------------------------------------------
2014/12/16-20:34:38 ... thread 0: (200000,200000) ops and (128923.6,128923.6) ops/second in (1.551306,1.551306) seconds
2014/12/16-20:34:39 ... thread 0: (200000,400000) ops and (110571.0,119044.1) ops/second in (1.808792,3.360098) seconds
2014/12/16-20:34:42 ... thread 0: (200000,600000) ops and (87921.9,106480.3) ops/second in (2.274746,5.634844) seconds
2014/12/16-20:34:44 ... thread 0: (200000,800000) ops and (107459.0,106723.3) ops/second in (1.861175,7.496019) seconds
2014/12/16-20:34:46 ... thread 0: (200000,1000000) ops and (90923.1,103138.7) ops/second in (2.199660,9.695679) seconds
2014/12/16-20:34:50 ... thread 0: (200000,1200000) ops and (51852.2,88542.6) ops/second in (3.857118,13.552797) seconds
2014/12/16-20:34:53 ... thread 0: (200000,1400000) ops and (57516.7,82207.6) ops/second in (3.477253,17.030050) seconds
...
by the end 2014/12/16-21:18:19 ... thread 0: (200000,22400000) ops and (858.0,8541.6) ops/second in (233.094661,2622.447456) seconds
2014/12/16-21:19:19 ... thread 0: (200000,22600000) ops and (3321.7,8424.5) ops/second in (60.210362,2682.657818) seconds
2014/12/16-21:20:28 ... thread 0: (200000,22800000) ops and (2880.3,8284.6) ops/second in (69.437524,2752.095342) seconds
2014/12/16-21:21:31 ... thread 0: (200000,23000000) ops and (3203.6,8171.9) ops/second in (62.429181,2814.524523) seconds
then a stream of these start showing up in dmesg Dec 16 21:14:09 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
Dec 16 21:14:30 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
Dec 16 21:14:32 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
Dec 16 21:14:34 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
Dec 16 21:14:36 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
and on and on until I kill the job. Dec 16 22:19:01 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
Dec 16 22:19:03 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
Dec 16 22:19:05 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
Dec 16 22:19:07 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
Dec 16 22:19:09 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
Dec 16 22:19:11 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
Dec 16 22:19:13 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
Dec 16 22:19:15 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
Dec 16 22:19:17 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
Dec 16 22:19:19 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
Dec 16 22:19:21 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
Dec 16 22:19:23 localhost kernel: XFS: possible memory allocation deadlock in kmem_alloc (mode:0x250)
Dec 17 00:14:16 localhost systemd: Got automount request for /proc/sys/fs/binfmt_misc, triggered by 1372 (vmtoolsd)
LMDB does this load in 8 minutes at an effective data rate of ~283MB/sec. Pretty much as fast as the hardware will stream (peak throughput of 300MB/sec in this VM). I don't see how you can possibly use MDBM as an unreliable data store if it could potentially take hours to reload the data.At least back when I was using it, there is a bit of a science to tuning mdbm correctly for your data. In general, it is a tool designed to allow you to tweak with whatever application level knowledge you have of your data's structure, and as a consequence it can be terribly suboptimal out of the box. Depending on circumstances, even if you don't need disk level transactional integrity, LMDB may indeed be a be better choice.
That said, the test case you are describing is definitely for a use case for which mdbm is suboptimal. The results you are getting are if anything surprisingly good under the circumstances, and frankly I'd have never even considered using mdbm for that kind of work load (LMDB would definitely be one of the first choices I'd consider for that kind of work load).
As you've mentioned, a hashtable based key-value store is fairly suboptimal for disk based storage (though depending on circumstnace and how you tweak it, a hash table with mdbm can work surprisingly well with an SSD based store), and the numbers you are presenting seem if anything better than I'd expect.
The reports you are seeing with XFS seem... odd, and almost feel like either a simple issue with a bug in how mdbm is talking to mmap/vfs layer, or more likely within XFS's implementation, but it doesn't seem like they are slowing you down.
In general, mdbm is most useful for storing and accessing compact rows in memory across potentially many processes with a random access pattern, which IMHO is not the problem space that you are testing and tackling with LMDB, LevelDB, or many of the others. One can (rightly) argue that that is a fairly narrow and simple problem space, but as with almost anything in computer science, doing even fairly narrow and simple things efficiently (and mdbm has a number of clever design choices that help it be efficient) and in an error free fashion is enough trouble that having a standard tool for solving that problem is terribly useful.
MDBM isn't really the only tool for solving that problem out there. Pretty much any shared memory hash table based solution may be a good fit for it, and there are alternative data structures that have desirable advantages over hash tables (critbit based structures are one of my favourite pets for such problems). Heck, in C++ I've used the Boost.IPC library's unordered maps for the job with reasonable results.
I bet there are probably some implementations that perform better than mdbm in certain cases (I'd actually be interested in benchmarks comparing that kind of workload with other tools designed for that problem space). Still, the mdbm codebase is battle hardened and really does perform well as long as your data set size & access patterns don't cause thrashing of the page store.
As another note, I also tried to use mdbm_pre_split() at the beginning of the job. That churned for 10 minutes at the beginning of the job before adding any records, then started processing at its normal speed, and then quickly degraded into the same state.
The "2 seeks per access" absolutely does NOT go out the window. The whole point was that this works for any size DB. It's a function of the size of the pages, the keys, and the values, for a run with key+val @ 20 bytes and 8K pages, 25 million entries had a directory of 16KB. You hash the key, walk however much of the 16KB you need to find the page, and then you go to that page and only that page.
2 seeks for any lookup for any sized DB. The directory could be better but you walk sequentially (hell, map that and lock it in, then it is 1 seek for any lookup).
The reason you are thrashing XFS so much is we're growing the directory and the number of pages. Each time you split a page you have to copy to the new pages.
Even if they do, it is often faster and more reliable to recover from the RAM of a surviving node than to try to recover from the disk on the crashed system.
> lets say that it takes a full day to rebuild the MDBM
I can't imagine a circumstance where it'd take even an hour to rebuild an MDBM from raw data.
Who said anything about them not being atomically visible? Not being ACID doesn't mean "none of the above", and atomicity in an embedded database is practically a footnote in a distributed system anyway, because you're looking for atomicity on a much grander scale. You can impose whatever locking strategy you want around the memory to ensure updates are atomically visible. It's an embedded DB so consistency is in the hands of the application logic. Traditional ACID terms for isolation and durability aren't there, but for say a distributed system you'd do that in the network layer.
If you have multiple replicas, if one node dies you don't try to recover its data (oddly, storage corruption and node failure have a high co-occurrence rate ;-), you just wipe and pull from the other nodes. If you can't get a clean, consistent state from a node that hasn't failed, you've got bigger problems than your embedded database. As I pointed out, pulling from other nodes is often faster and more reliable than pulling from disk anyway.
In general, with a high performance network application, it is often preferable address these issues somewhere other than in your embedded database.
So what is the corruption you are envisioning?
Could there be a comparison between these datastores and the traditional ACID compliant databases when it comes to retrieving actual data in a useful format? E.g. perhaps doing a join or an ordering of some sort? I don't expect databases (e.g. Oracle, MS SQL Server, DB2) to be faster in raw performance, but I do expect them to be faster in terms of total development time and bug fixing since the application developer wouldn't have to do the locking, page pinning/unpinning, etc. manually.
ln -s -f -r /tmp/install/lib64/libmdbm.so.4 /tmp/install/lib64/libmdbm.so ln: invalid option -- 'r' Try `ln --help' for more informatio
just use a std::unordered_map, or better yet a tbb::concurrent_unordered_map or whatever the equivalent is for your language
Another reason to cache to disk is that you want to store more data than you have ram.
Practically speaking, Boost.Interprocess includes a shared memory hash table implementation. Boost Multi Index, which is a further generalisation of containers to allow the construction of database-like indexes, is also Interprocess compatible.
http://www.boost.org/doc/libs/1_57_0/doc/html/interprocess/a...
Good examples: http://duktape.org/ (it might seem silly but that right column makes people want to try it!), http://redis.io (i bet this page wins many folks http://redis.io/topics/twitter-clone)
It seems over the last year technology has been growing more rapidly than any other period.
Fun times but so hard to keep track of everything!
I tend to agree with this statement. The entire stack appears to be going through a revolution.
The data layer in particular is seeing very rapid change after being largely (not entirely) static for decades.
While I'm sure someone out there will see this and say "wow, that's exactly what I need!" chances are that if you have these sorts of scale issues you're going to have to figure it out on your own.
I'd rather see a write-up of how they arrived at this particular conclusion than another non-database.
At any given time, you either have a need/problem, or you don't. If you DO, you evaluate the current tech available, and hopefully select something that fits your needs. You build out around said tech, and if your choice was correct, that means it's either solving your problem, or on it's way to.
If something comes along while you're implementing with your chosen solution, that looks similar, but better, it's only noise - because hey, you found a solution.
Just as we don't all re-write all of our code whenever a new language comes along (unless the thing in question was desperately in need of a re-write anyway) even if newer languages are nicer, we needn't switch DBs or frameworks for the same reasons.
1) it's non-stop
2) there seldom sems to be anything truly novel in a broadly meaningful way (i.e. esoteric, if anything)
3) there is rarely an objective improvement on existing options
I no longer feel compelled to replace or adopt though, precisely for those reasons.