Berkeley DB Architecture - NoSQL Before NoSQL Was Cool
highscalability.com
highscalability.com
However I think it was not a NoSQL database for an important reason, it was only embedded in other programs, and not in the form of a networked server. Not only this raised the barrier to entry, it never made Berkeley DB an "object" in your infrastructure that you could use in different ways, and in a competitive way with other databases.
So the Berkeley DB creators did not started the NoSQL era long time before because they missed the importance of what they were doing, thinking that their DB was limited to just something you could bind to programs not needing all the power of an SQL database. At least this is what emerges form their choices. And the implication is that they were also considering relational databases as the only "real DBs" in my opinion.
So there was no real competition with traditional DBs, and it was not a NoSQL DB.
EDIT: I understand this is a mostly personal point of view, and not an objective critique, but I wanted to share it with HN nevertheless.
Besides it's ability to be embedded, BerkeleyDB doesn't really provide any advantages over an RDBMS, other than raw simplicity. It's unlikely anyone looking at data persistence solutions would consider BDB alongside an RDBMS; they solve different problems.
I think BDB is a useful and well-designed tool, but it's not NoSQL.
I'm using LDB in my most-recent project. I wrote a trivial HTTP wrapper around it using Mongoose to make it a "server". :)
I like the snapshotting feature, the support for transactions, the built-in support for accessing the db from multiple threads, and that it maintains keys in sorted order (this great for timestamped keys like I have).
I looked at BDB a while back, and IIRC it doesn't have snapshotting and sorted keys. Unsure about threads.
I use the snapshotting feature dump a backup copy of the database.
The BDB API seemed significantly more complex than Leveldb's. IIRC, the native API is C.
I like C++ and like that Leveldb's native API is C++. It uses the standard library string (which allows embedding null characters).
I'm not familiar enough with Berkeley DB to comment on the rest, but it definitely supports B-tree indexes which store keys in sorted order.
It is easy to programmatically configure it incorrectly. E.g. not use transactions when you should, fail to deal with deadlocks or not clean up locks of failed threads or processes.
Some of the mature features it has are:
Page-level locking for concurrency via threading and multi-process. Snapshot isolation for MVCC - you can get a high isolation level without read locks. Nested transactions. B-tree indexes. Two-phase commit. You can also throw all this out the window and use an in-memory database or non-durable
The newer versions even have replication.
BDB is a hard beast to tame properly, but is definitely fully featured, and mature.
The Oracle paper on the BDB backed SQLite makes an interesting read.
With all the NoSQL Hype Redis and MongoDB get I think these are too often overseen. What's really nice is that you can use them in embedded (cabinet) or server (tyrant) "mode". What's also great about them is that they provide a lot of flexibility, since they are not just one database. Also great that they have official bindings that work. They have tons of features and sometimes an embedded database is really interesting when you don't want (need) to care about the protocol you send it over.
Also, I already did have one Tokyo Cabinet Hashtable go corrupt on me--which caused certain operations to lock up the CPU indefinitely. Hmm, that didn't ease my conscience. However, it was completely recoverable with "tchmgr optimize". It'd be nice to google for more war stories, but like I said... breadcrumbs so far.
I've not found any such limitations with LevelDB
Running a single threaded access pattern can easily get you 20K plus reads/sec, but if you try to run more threads the throughput per thread just goes down, up to a point where more threads actually slow you down.
Make sure to run extensive benchmarks if you consider using it in a multi-threaded application.
I've never tried running BDB in a multi-process architecture, so I have no idea how it'd behave when used that way.
I don't know all that much about it, but given the background of those involved, should be worth keeping an eye on.
[1] http://en.wikipedia.org/wiki/Linda_(coordination_language)
Dunno about performance numbers vs BDB etc.
This builds on the ideas from BDB and Tokyo Cabinent with a newer codebase: http://fallabs.com/kyotocabinet/
Both have network server implementations available (see links from their pages).
I used Subversion since pre-1.0. The only problems I've ever had with Subversion were caused by BerkeleyDB failing to be sufficiently robust. Since Subversion eliminated BDB use, I've never had a problem with it.
BerkeleyDB is dead to me.