Answers to common questions about RethinkDB
rethinkdb.com
rethinkdb.com
We ran into a little bug w/ the ajax requests using the wrong port when run behind an nginx reverse proxy, and the team was very helpful and responsive in getting it fixed: https://github.com/rethinkdb/rethinkdb/issues/63
Seems like a very promising project
No good deed goes unpunished, etc.
We'll fix this ASAP (though this is somewhat of a low priority).
It's awesome that you're hosting this you should just be sure it's on a machine you're not depending on for other services.
Only try this at home.
Anyone know of any projects to create a Rails ODM for RethinkDB? (a la Mongoid/MongoMapper?)
Seems to me a natural fit given the (relative) popularity of MongoDB in the Rails world and the comparisons/improvements of RethinkDB over MongoDB.
Might be a good place to start.
For Golang, mgo http://labix.org/mgo is a quality implementation driver for MongoDB. Hopefully there is something similar for RethinkDB.
As far as ORMs go, we do plan to do that, but it will take a little bit of time. It shouldn't be very difficult, so we hope to get to it soon.
What would be the best way to set up a one-to-many relationship in RethinkDB? For example, a User has a Store, a Store has Categories, and a Category has Products.
Along the same lines, is there some functionality similar to MongoDB's compound indexes?
Use the former approach if your array will be relatively small (say, under ~2000 items). Use the latter if the number of relationships is much bigger than that.
As far as compound indexes, we don't have support for that yet, but it will happen.
Also, in your API documentation [1], I see there's an `append` function to add a value to an array. Is there an equivalent function to remove a value from an array?
I can then say: table("A-to-B").eq_join("a_id", table("A")).zip().eq_join("b_id", table("B")).zip()
And you'll get a document for each corresponding pair. If the ids are in an array then it's going to be a bit painful to write that query. (Although it can be done I believe).
A future feature will probably be to make eq_join work with arrays of foreign keys.
Nice!!!
This approach could work well when:
- you don't access child documents directly
- you don't update child documents too often
- child documents are usually unique entities
As you can see from the rules above, this would not be a good solution for the Store -> Categories -> Products (contradicts 1. and 3.) But it could work with models like Content -> Votes, Blog -> Comments, etc.
I have to build an analytics db for commercial launch in February, it'll be a single machine cluster for the first few months. Would rethink be a risky choice?
That being said, if you choose to build analytics on top of RethinkDB, we'll support you all the way.
By the way, for the benefit of the commenters here, I checked RethinkDB two days ago for http://www.instahero.com (the analytics app I'm building), and I ran into a bug where RethinkDB's count() would only return half my data.
The team was very responsive and fixed the bug in a few hours (and pushed the fix so I could test it), so I'm very impressed by the response time. I can't hold the bug against them, because they do say it's not production-ready.
Also, I pointed out that the DB listening to interfaces other than localhost by default was a potential security risk and that, too, was fixed in around a day. Kudos.
It would be nice to mention that RethinkDB only compiles on x86-64. If you follow the instructions on the linked "building from source" page, it works fine on 32 bit machines ... until you actually run "make".
It wasn't that much time, but I wouldn't have tried if I'd known ahead of time.
Thanks for the effort!
That being said, the way we're approaching this issue is by overhauling the test infrastructure to stress the server in as many ways as we possibly can. I suspect it will take about three months for us to feel satisfied with how extensive the test infrastructure is. Once that's in place, we'll be able to recommend rethink for production. Of course if the test infrastructure exposes hard-to-kill bugs, it will take a little longer.
How would you re-architect it if you knew ahead of time it would only sit on SSDs - never on spinning media?
SSDs are, unlike rotational drives, inherently parallel devices. Each drive has a number of flash memory units which are capable of concurrently doing writes. To split the load evenly we divide files up into zones which we call "extents" the number of extents is configurable and varying the number of concurrent extents will give varying performance for a given drive. Normally there's a sweet spot where you have as many extents as it has independent chips. Extents are written to in a very specific manner. They are append only which means that you start at the top of the extent and fill it up with blocks of a predetermined size. The block size can be tuned to optimize performance as well although I'm not sure off the top of my head which properties of the drive determine which block size is best. Once an extent has been filled up top to bottom we stop writing to it and move on to another extent.
Garbage collection: Each block that we write in an extent is actually a version of a logical block in a btree. When blocks are changes they are rewritten and the old block is marked as garbage. This leads to extents which are filled almost entirely with garbage. When we hit a threshold we look for the extent with the most garbage copy all of the non garbage blocks to an active extent and then reclaim the extent. SSDs have their own garbage collection mechanisms that can seriously degrade the performance of the drive if when they misbehave. If we get the extent size right the drive will have a much easier time of garbage collection (I actually think the drive won't have to do anything at all for gc in many cases.