SenseiDB: Open-source, distributed, realtime, semi-structured database
senseidb.com
senseidb.com
Sensei (先生) means teacher or professor in Japanese(http://en.wikipedia.org/wiki/Sensei).
It shares the same pronunciation and writing with the Chinese word that has the same meaning. This name indicates that the system can be used in place of Oracle database in many applications.
Okay, so on the name alone I can replace my Oracle database! Great!Seriously though it goes on...
They claim the database is ACID: http://javasoze.github.com/sensei/data-guarantee.html
But they built the entire thing around "eventual consistency."
And statements like:
"Sensei provides a high-level of durability by maintaining N replicas of each shard to guarantee a level of availability and fault-tolerance"
Don't seem to make sense when talking about ACID given that a write operation will happen at some point. Looks like the data event producers will shard the data across N replicas without quorum... so there's no guarantee that there will be N replicas available... is that right (and that the transaction won't be lost mid-stream either)?Skimming through the source it doesn't seem to be doing anything terribly revolutionary... and I can see the usefulness of the trade-offs they made in this database for certain scenarios. However I don't think the claims of ACID guarantees and "real time" are particularly representative of what this DB will actually do. They just don't seem to jive with "eventual consistency" models.
I'm not a hardcore database guru though so maybe I'm missing something?
Sensei needs an event stream to process. We've open-sourced and apache-fied Kafka which is a great candidate for an event stream. For Atomicity and Isolation, the event stream must provide these guarantees.
Consistency is handled with a routing parameter. Requests partitioned around an id will always go to the same searcher, so they won't go backwards in the stream except in failure scenarios. This is eventually consistent, but tries to keep things sane.
Durability: The event stream helps with this. We don't immediately flush while indexing in Lucene, so if there's a crash we can replay the persistent event stream.
Does this make sense? SenseiDB is not intended for purely transactional processing. For some applications, sensei would make a good candidate for replacing your DB. For others, not so much.
I'd suggest rewriting the copy as something along these lines:
# Data guarantees: How we manage your data
## Not quite ACID (and that's not a bad thing)
Although the principles of ACID are important, strict conformance has its costs. Sometimes it's the right tradeoff — but not always. By relaxing those requirements, Sensei can offer superior performance and durability. Here's how we approach the four ACID principles, and why you might prefer to do things Sensei's way:
And they're all well and good features! I can tell SenseiDB isn't transactional. Like I said, I skimmed the source and understand at a high level what it's does. I could see it being very useful in certain conditions as LinkedIn currently does and I'm sure others will.
However, I think the copy is confusing (at least it was for me). On the guarantees page there's an "ACID-ity" headline. For each aspect of ACID, as you say, the page describes what Sensei offers. The confusing part was that I was mentally comparing each aspect against what I understand to be the common semantics of ACID. I think it would be more clear if there was some distinction under the main headline that acknowledges this difference.
hth!
(Nitpick - none of the urls change. Some js error where you're not doing a pushState? You should report this.)
edit - this works -http://senseidb.github.com/sensei/index.html
Looking through the documentation a couple of things come to mind:
* Does SenseiDB support nested data structures?
* What happens when you modify the schema? Can you add/delete/modify attributes?
1) is straightforward, but can be restrictive if you want to in-nest faceting.
We are still debating on the route to take.
As for schema changes, we should document that better: certain schema changes simply requires you bouncing the node, but if there are data integrity changes, you would need to reindex. We will work on more documentation on the specifics.
Each system serves a need. We also use off the shelf databases, where appropriate: it should be noted Sensei uses Lucene (a well known search library), Voldemort has a pluggable storage engine (where we mostly use BerkeleyDB and a custom read-only storage engine), and Espresso uses MySQL.
We don't build databases for reasons of NIH: we focus on building features (faceting, real-time indexing, partitioning, fail over, etc...) that enable us to build fast, usable, feature rich, scalable, and reliable applications. We readily use open source components in many places within both our infrastructure systems (search, various databases, Kafka) and in the applications, and contribute to existing open source projects.
One thing is for sure -- NoSQL isn't so mysterious anymore :)
Would I have to publish messages to Kafka in my language of choice to avoid writing any Java?
We love Kafka!
Edit: You can get a pretty quick idea of what's supported the query language at http://senseidb.github.com/sensei/bql.html. Since Sensei was designed with document and text uses cases in mind, we have plans to support advanced queries where a relevance model is part of the query, allowing you to perform a custom sort on the server.
Thanks for sharing this with the world. :)