RethinkDB (YC S09): MySQL Storage Engine Built From The Ground Up For SSD
techcrunch.com
techcrunch.com
Since your primary market will probably be developers, describing RethinkDB will be necessary, but not sufficient. Also, demonstrated performance will be necessary, but again, probably not sufficient. I, for one, want to understand what goes on under the hood (and be able to describe it to my customers).
The "Time to Insert 2 Million Records" graph was impressive. How does the "Time to Retrieve, Sort, & Present 2 Million Records" graph look? How about the "Time to Modify 14% of the 2 Million Records" graph?
Your append-only approach sounds great for adds. How is it for changes and deletes? How will garbage collection affect performance?
"No more locks" is a great claim, but how will it work in a real world enterprise-quality app? User A takes 30 seconds to change Zip Code, Phone Number, and increase Credit Limit while User B takes 10 seconds in the middle of that to change City and decrease Credit Limit for the same customer. Who wins? This is a difficult scenario in both optimistic and pessimistic environments. I can only imagine how it's handled in a "no lock" environment.
(I'm not looking for answers here, just spouting off what's on the top of my head, but I will be looking to better understand on your website. Your white papers oughta be interesting.)
You make ambitious claims. I look forward to seeing you fulfill them. Best wishes!
Hope you enjoy the papers.
Locks in this context refers to readers and writers at the DB engine level.
The former is much harder -- it's a significantly complex task just to figure out which data records are still live.
> All of our interesting work is MySQL API-independent, so a Postgres port is not out of the question. We’ve also been entertaining the idea of porting to SQLite, as many embedded devices use that, and have SSDs already.
- Allow others to vet your codebase for stability and security
- Give customers some recourse if your startup folds
- Make you comparable to MyISAM/InnoDB/PostgreSQL, unless you want to be compared to Oracle or Microsoft SQL
- Nothing much I can say about stability and security, but then again, we're not saying it's stable and secure yet. Besides, people trust Oracle without seeing their sources, don't they?
- If our startup folds, I doubt we'll drag our source to the grave. That said, maybe it gives our users incentive to make sure we don't fold? :P
- There are closed-source MySQL storage engines out there with whom we'd rather be compared (TokuDB, Falcon (is Falcon open-source?)).
We just don't want to close any doors yet. If it makes good business sense, we'd be glad to open the source.
Interesting that there are other closed-source MySQL storage engines. I wonder how big a piece of the pie (among paying or willing-to-pay users) they have compared to MyISAM/InnoDB.
I wouldn't bother commenting on it, but surprisingly this is not the first time I've seen someone lifting Apple icons for their startup website, and it's something that needs to be addressed: artists are expensive, but you can't take other people's artwork, and worse yet, it's blazingly obvious when you use Apple's.
[edit]
The Xcode icon you linked to comes from a user-uploaded KDE Icon Theme:
http://www.kde-look.org/content/show.php/Dark-Glass+reviewed...
If you download the actual theme set, you'll find that the copyright ownership is unknown: "99% of this set is GPL now and what's not is most likely creative commons (a tiny number of the mime types may be proprietary). PLEASE abide by the licence rules, if you use icons from this set please research and credit the appropriate people. I have been given permission to release other peoples art work under the GPL so respect the licence." (from the README)
The original icon theme may be found here: http://www.mentalrey.it/project.html
As noted by mikejs below, Apple actually uses the same icon you're using. Digging a bit, it appears it's the icon Apple used for Xcode in Mac OS X 10.4 Tiger (Xcode 2.5): http://www.command-tab.com/images/photoshop/tiger_icons/prev...
Xcode 3.0 actually introduced the right-facing hammer:
http://developer.apple.com/DOCUMENTATION/Cocoa/Conceptual/Ob...
Rather than figuring out who has provenance, we decided to just change the icons (which are now all GPL). If you notice any other issues, please let us know (founders@rethinkdb.com).
I run Iconfinder and I'm trying hard not to add proprietary icons, but sometimes icons slip through in the large icon sets with 1,000s of icons. The icons you mention above are removed from the site.
Best regards, Martin Leblanc
You may not use locking for concurrency control (plenty of "traditional DBs" don't, either), but you still need some sort of concurrency control scheme -- just using append-only/log-structured storage doesn't make CC free. I'd be curious to hear how you guys are doing this.
I don't think I can explain the technology in a comment, but we'll definitely be releasing more information in the coming weeks. We had to get this release out for now. Stay tuned!
Eventually, we will allow multiple writers, and use some form of STM for that; we just haven't gotten to it yet.
If you'd like to speak interactively on this, shoot me an email.
The "traditional" (and I don't know how you could call the latest version of any of the major databases "traditional", this is a brutally competitive market) databases don't lock just for the fun of it, but to enable features that users want. Anyone can come up with a product that doesn't do Y if it can't do X either. So what're we missing here?
Oracle has done MVCC for many years, as has Postgres. The canonical paper on optimistic concurrency control for DBs is from 1981 (http://www.seas.upenn.edu/~zives/cis650/papers/opt-cc.pdf).
I didn't follow the rest of your comment, I'm afraid -- I was just saying that I didn't see how using append-only storage immediately makes concurrency control a non-issue. The comments from the RethinkDB guys upthread support that: not supporting concurrent writers makes your concurrency control much more straightforward.
I would like to read the benchmark script.
Also, I'm afraid to read their source code after reading the license agreement. I can't sell support for any product that can communicate with RethinkDB? That sounds unenforceable, but it is scary enough to prevent me from even looking at the code.
The engine isn't open source. We're considering open sourcing it in the future, but we want to understand all business implications of this decision before we proceed - it's a decision you can't easily retract. The license is a bit draconian, but this is because we've only released a developer pre-alpha. We don't want people to use the engine in production yet - it's not ready. AFAIK, the license says you can't sell RethinkDB support, not that you can't sell support for software that uses RethinkDB.
From the license:
Prohibited activities include but are not limited to:
Selling support for products which incorporate RethinkDB.
This is the problem with rolling your own software licenses.
I just had one of those "why didn't I think of that" moments as I read through your wiki thinking back to this paper I read a few months ago: http://publications.csail.mit.edu/lcs/specpub.php?id=773
Cheap commodity disks however have an annualized failure rate of 4% in Google's datacenter (according to their disk analysis paper.)
Stealing thunder = releasing the same exact product, a month before you finish.
It's also nice to see a lot of interesting database related companies coming out of Stony Brook -- I think one of the founders of tokutek is from Stony Brook as well.
They are doing some _really_ cool stuff!