RethinkDB 1.6 is out: regex matching, new array operations, random sampling
rethinkdb.com
rethinkdb.com
Also, a nice +1 for RethinkDB is when I was importing way too much data for the (accidentally) way too small EC2 instance, the OOM killer kept knocking off RethinkDB but not a single doc had been corrupted. I've done similar stupid things with other DB's and... well, data was lost.
We're still working through it for analytics, I hope to do a write up sometime soon.
As far as a JVM-based client driver, kclay in this thread mentioned that he's been building a Scala driver that is compatible with 1.5, and will update it for 1.6. There's also a thread on Google Groups (https://groups.google.com/forum/?fromgroups#!topic/rethinkdb...) detailing some efforts on a Java driver to date.
If you're interested in contributing, shoot me an email at mike [at] rethinkdb.com. I'd love to see a Java driver happen, and we'll do everything we can to support such a project.
I would encourage everyone working on drivers for RethinkDB to publish them regardless of how complete they are, and advertise them as partial implementations. I know of a lot of people interested in working on a Java driver, and it would be great to see collaboration, even early on!
This has actually been one of my biggest frustrations -- there is an enormous amount of really cool stuff in Rethink that's relatively poorly documented, so people can't find out about it. I'm really looking forward to fixing this soon.
It looks like an open source project with an enormous amount of amazing talent behind it, and YC backing; but considering that it's FOSS now; how do you plan on generating revenue?
RethinkDB already has paying customers, and we'll soon be offering publicly available commercial support (not unlike JBoss, MySQL, 10Gen, etc.) There are also other revenue streams I can't talk about yet.
For anyone who's interested in the pilot program before commercial support becomes publicly available, shoot me an email -- slava@rethinkdb.com.
I'll see what we can do to accelerate this -- possibly by enticing the community to pick this up.
If you post what needs done in detail, someone will surely pick it up. I'd certainly take a look and see if it was within my capabilities.
It seems like it would be hard to guarantee that the document and its index entries are always in the same shard, and since RethinkDB doesn't seem to support general multi-document transactions, I'm curious how you go about updating them simultaneously.
There are tradeoffs to this approach -- to read a secondary index the db has to go to every master/shard for the table, but there are lots of tricks we use to minimize the impact of this tradeoff.
For example the "insertAt" operation is really helpful when you need ordered arrays and you want to swap or change some elements. I'll give it a try definitely.
* A biased/big picture one -- http://rethinkdb.com/blog/mongodb-biased-comparison/
* An unbiased/technical one -- http://rethinkdb.com/docs/comparisons/mongodb/
Rethink has this air of being rock-solid about it too. The team worked really hard on making the core great at storing data, then built features around that. Mongo, after using it for many years both in production and personal projects, seems like it's all slapped together. The early-on technical decisions they made to get it out the door are still biting them in the ass to this day.
That's not to say there aren't some things Mongo is probably better at (especially considering how early-stage Rethink is) but overall I prefer Rethink hands down.
Reading over your guys' documentation and history, it seems that you have succeeded in every way Mongo fails: MVCC for high-write, homogeneous clustering (no config servers or route services), actual auto-sharding that doesn't make you want to gauge your eyes out, joins, a useful management interface baked in to the DB...those are the main ones I can think of.
I'm currently working with a lot of crypto data, so a binary storage format would be really nice in Rethink, I guess Mongo wins there (but I'm still using Rethink and just Base64-encoding everything).
There are one or two things I really do like about Mongo. For instance, the find-and-modify command on top of the built-in data structures makes it dead simple to build something like a queuing system. It can't actually handle the writes a real queuing system needs, but for lower-write stuff it's useful. For instance of you wanted to have a set of servers run shared cron jobs...they could all "check in" and see if another server is already running that job atomically (and if not grab the job in the same operation that checks it, ensuring that no two servers are running the same job at the same time). I'm not sure if Rethink supports anything like this at the moment (please correct me if I'm wrong).
Mongo also has GridFS (admittedly, I've never used it), but if I was going to do some sort of clustered filesystem, I'd use Riak, definitely not Mongo. Also, I believe Mongo now has full-text search, but once again, I've given up on primary databases having real/useful full-text search and would much rather use ElasticSearch to supplement the primary DB.
Once Rethink has a query optimizer (I know you are in the process of figuring this out and to what extent it will work), I don't think Mongo will have anything on Rethink. I'm certainly not ever going to choose Mongo over Rethink for another project. I may choose Postgres/Mysql for strictly relational data, but Mongo is pretty much dead to me at this point.
I'd rather suffer the consequences and growing pains of a newer DB that wins at just about everything it does than use something slightly more stable that has burned me a number of times.
Plus, I like your team better. I know none of the Mongo devs, but I've talked to many members of the Rethink team on many occasions and you've all been really helpful, supportive, and quick to fix any issues. I've said this before, but you guys aren't running around screaming about how great your DB is and how it will solve ALL OF YOUR PROBLEMS, get you a promotion, improve your sex life, etc etc. I feel like a lot of the reasons I've been burned by Mongo was because 10gen just wasn't straight with me or the teams I worked with when using Mongo. They'd much rather sell a support contract than actually see you succeed, it felt like.
Sorry for the book, TL;DR:
- Atomic "find and modify" is very useful in Mongo. Rethink may already support it via the QL.
- Binary storage type would be nice.
- Query optimizer (I know you're in the process of figuring this out).
- Packaged linux/mac Rethink binaries so people can download the latest version without a recompile (I do like this about Mongo)
I think that's about it. There may be many technical things I'm missing out of sheer ignorance of the internal workings of Rethink and Mongo. Maybe there are things Mongo is great at but Rethink isn't, but it would be news to me.
> the find-and-modify command
This has been baked into ReQL from day one and doesn't require a special command. Here's an example:
# 'jobs' table contains documents of the form
# { type: 'type', jobs: ['job1', 'job2'] }
r.table('jobs').filter({type: 'printer'}).update(function(row) {
return r.branch(
row('jobs').contains('job1'), // if there is job1 in the array
{jobs: row('jobs').difference(['job1'])}, // atomically remove job1
null // else don't modify the object
);
})
This isn't limited to arrays -- you can do all sorts of atomic find and modify this way within a single document. We clearly need to do a better job documenting this. I'll make sure that happens.> Packaged linux/mac Rethink binaries
That's been available for a while -- http://rethinkdb.com/docs/install/, you can download a binary pkg for OS X if you don't want to use brew, and we support apt-get on ubuntu/debian via a private PPA. We'll be adding binary packages for more distros soon.
> binary storage format
Mongo wins on this one, but it's definitely on the horizon -- https://github.com/rethinkdb/rethinkdb/issues/137. The storage engine bits for this have been done for a long time, we just have to figure out a sane ReQL API (which is harder than it seems).
> Query optimizer
This is a somewhat longer-term project (i.e. we probably won't get to it by mid-fall), but we're definitely thinking about it. So far though, there hasn't been a real use case yet where it's a problem.
Also: All the ruby driver doc seems wrong. 1.6 changed a lot of the APIs?
When I loop and keep generating these documents:
to_insert = {}
10.times {|i| to_insert["key#{i}"] = rand(33333).to_s * rand(6) }
I get around 140/s. Without the insert() call, I reach 50.000+, so it doesn't seem to be the overhead.With simple documents (3 keys, int values) I reach 350-400.
p.s. script @ https://gist.github.com/rb2k/5777997
In the meantime, you could try running a multithreaded/multiprocess script -- that would significantly increase throughput.
Sorry you ran into this -- it'll be fixed ASAP.
The old version basically looked exactly like the python code.
i could not find a good article about rethinkdb in production.
Can maybe someone share some thoughts about using rethinkdb on a database heavy web e commerce shop?
thx!
We also have a Google Group for driver developers (https://groups.google.com/forum/?fromgroups#!forum/rethinkdb...) that you can always ask questions on.
Thanks for working to build a Scala driver, I'm excited to see it!
https://github.com/rethinkdb/rethinkdb/wiki/Community-contri...
Every time I see an announcement from RethinkDB I get excited to try it, and then I see there is no driver for Java/Scala.
This is great news and excellent progress. Keep up the great work, Rethink team.
Expect an email shortly briefing you on the protobuf changes for 1.6!
Our major usecase for mongo is returning entries in ordered distance from a location and / or within some radius of a location.
Support for such queries would be seriously useful.
I'd like to get geospacial support in, but it will take a bit of time before we can get to that. We'll likely get it in within a year, so it's a fairly long-term feature for now.
Porting to Windows is another level of complexity altogether.