RethinkDB 1.16: cluster management API, realtime push
rethinkdb.com
rethinkdb.com
If you're interested in joining, you can RSVP here: http://www.meetup.com/RethinkDB-Bay-Area-Meetup-Group/events...
The last thing we need is another MongoDB that can not be trusted with anything valuable.
It's rotten from the core, poorly designed on every layer, and has caused many companies great grief.
I wouldn't touch it or any of it's descendants with a 10 foot pole.
We're systems people. You won't get another Mongo.
The biggest difference in RethinkDB's approach has been data sanity first, niceties a close second (nice devops and dev interfaces)... It's a very well thought out database, and as long as people are moving a little farther away from SQL thinking and interfaces it makes more sense.
I'm not sure where automatic failover stands wrt RethinkDB, but barring that, I would probably reach for RethinkDB, MongoDB or ElasticSearch for most small-medium db tasks, depending on what the application's needs are. By medium I mean less than 10-15 server cluster. Anything bigger, Cassandra would probably be my first choice.
One need I have in nearly every project is record versioning. Think implementing a modern wiki, in a CRM (what did X alter? Which phone numbers have been assigned to this person?), versioning updates to data records etc. Realtime web with multiple users editing records simultaneously, it's even more important.
Versioning is a pain to implement in every system I've used, and I'm looking for something to lower that pain.
CouchDB et al have built-in versioning, but cause pain in other places. I've looked at the RethinkDB docs and found no mention - so how would RethinkDB users handle versioning, and are there any helpers on the way?
Datomic on the other hand does claim to provide "rewindability".
Yes, Datomic advertises that it keeps track of a documents history.
If you set returnChanges, the result will contain two properties, old_val and new_val. You could take the contents of old_val, give it a unique identifier, and then store it in another table with r.create().
[1] More about r.update(): http://rethinkdb.com/api/javascript/update/
I would have to be convinced that the DB is an appropriate place to handle versioning of the kind you are speaking of.
For instance, see this article: http://www.xaprb.com/blog/2013/12/28/immutability-mvcc-and-g... Upshot is that RethinkDB started out append-only with rewindability etc. and then ran into real-world problems.
Every system I've ever seen that did versioning does it at the software level and stores multiple records in the database.
Rails gem "paper trail" is an example.
My comment can be rephrased that it's still painful to do this at the application level: for such a common task, the frameworks/libraries/software I work with don't have anything baked in to handle it.
I've started trying to use Git as an additional data backend to hold the versioning, but changing from a JSON/Database storage structure to one Git will happily diff and reconcile is in itself not pretty.
In this case, you can do r.table("data").get(1) // get the current version of the document with id 1 r.table("history").getAll(1, {index: "id"}) // get all the previous versions of the document with id 1. And you can order them by `last_updated` if you want too.
We're implementing real time features in our product in the next couple of months so this could not have come at a better time.
I missed the Q&A this afternoon, but I see lots of RethinkDB engineers on here... so, would there be severe performance implications of holding open that many cursors?
If you go over a couple 1,000 active changefeeds in 1.16, I recommend setting the `maxBatchSeconds` optional argument to `run` to something like 10 (the default is 0.5).
That change should significantly reduce the CPU overhead on the server if you have lots of idle changefeeds. Note that this does not affect how quickly changes get delivered - you will still get changes instantly.
The exact performance will of course depend on what queries you're going to run exactly, the rate of writes etc.
(Related: the RethinkDB team is very responsive in the IRC channel. It's great to see how interactive they are with the community!)
What is the relation between write availability and A(tomicity)?
I think this sentence is just phrased poorly. It means to convey that there is no atomicity across multiple documents, which is unrelated to the previous sentence. I've opened an issue to fix that: https://github.com/rethinkdb/docs/issues/633. Thanks for noticing!
I could not find enough info in the documents and the github issues are not clear about the status and the roadmap :(
Let me know if you want to know more about a specific aspect.
Sounds like you need to implement raft. Do you have any plans in that direction?
From the website also I cannot understand the advantages compared to a relational database.
The idea behind the architecture was that performance should be significantly better than rolling your own infrastructure, because the database has a lot of information that userland (from the database perspective) software doesn't.
The performance of inserts might slow down slightly (matter of microseconds in insert latency) if you create many feeds. The database has to look at each insert and figure out if it applies to any of the feeds. This code is written in optimized C++ and is very fast. We're still running benchmarks, but we're shooting for performance levels where you (as a user) might not even be able to measure the difference.
The same applies for inserts that aren't affecting feeds (on a per table basis).
Same goes for throughput -- it might slow down slightly, but we're shooting for making the slowdown barely measurable if at all.
EDIT: in clustered environments, if you're subscribed to 1000 changefeeds on machine A and a write happens on machine B, we do constant work on B to send the changes to A and then A does all the work to figure out which changefeeds need to see it. TL;DR: We don't block out other writes for time proportional to the number of feeds.
Can you elaborate?
_but we're shooting for performance levels where you (as a user) might not even be able to measure the difference._
Benchmarks are almost always skewed towards the preferred workload of the DB (they're like science experiments, the result is heavily biased). How are you ensuring this isn't the case with your benchmarks?
Also what are you benchmarking against? I'd like to see one against Aerospike.
Let's say I have a table containing data for many users, while each subscription only needs data for a single user. Instead of scanning through all the changefeeds, you could put subscriptions in a hashmap and figure out which ones to update in O(1) time rather than O(N) time in number ob subscriptions, per update.
Obviously this is much harder in the general case, but do you do anything along these lines?
PS Absolutely loving RethinkDB, I've been using it as my main database since April and its a joy to work with!
We are actively developing the foundations for integrated auto-failover at the moment, so auto-failover is going to come soon.
The solution is to re-architect RethinkDB so that it can reconfigure a table even if there's a disconnected server. This is a pretty big project, but we're working on it, and it will probably ship around April or May. We'll also include server-side auto failover at that time, because it's easy once this problem is solved.
I guess my question is why would this be a preferred solution? It seems to run afoul to the 'one job' design goal. What am I missing?
This is a preferred solution because the only efficient alternative is to inspect every insert or update, manually determining which clients care about those changes.
- You don't have to write logic to figure out which clients need to be updated with some data. You just run the queries that you need to to generate the data for the client, and then RethinkDB sends you an update only if the result of that specific query changes.
- Having the thing that modifies the data send the updates to the connected clients becomes increasingly difficult if you have multiple application servers. Then you would need to set up some separate message passing / broadcasting infrastructure if you also want to update clients connected to other servers. RethinkDB takes care of "passing the news around", even in a distributed environment.
- RethinkDB supports changefeeds not just on the raw data, but also on transformations of the data. Not all transformations are supported yet (for example map reduce queries are not), but our goal is definitely to support changefeeds on pretty much any query. Just knowing that the underlying data has changed isn't enough. In a traditional setting, you would still need to either recompute the whole query, or implement your own logic for incrementally updating their results for every type of query you want to run. RethinkDB updates query results incrementally and efficiently.
(I work for RethinkDB)
In the case of an exchange I never split a single order book across multiple servers, but I can imagine a lot of applications where this could be an issue. How do you handle data consistency across nodes? Ultimately you have to solve the same issue...
Again, this isn't a big limitation for my use case. That said your answer has certainly given me a greater understanding of other circumstances where it would be very useful. Thank you.
If you shard a RethinkDB table to split it across multiple servers, and then create a changefeed on the table, the database will automatically send changes from both servers. Basically, server management/sharding in RethinkDB is visible to ops people, but is completely abstracted from the application developer. All writes are immediately consistent.
Rethink doesn't provide ACID guarantees, though. If you want to make a change to multiple documents in a table and have ACID guarantees, I'd stick with traditional RDBMSes.
I suppose I could just ask if you have an architecture document floating around :)
We feel like for many queries (especially the ones you find in web applications) that's not a big deal, since you can efficiently re-run them after reconnecting. In other cases it definitely matters, and we are going to add what you describe in a future release. You can follow the progress (or chime in if you like) at https://github.com/rethinkdb/rethinkdb/issues/3471 .
- https://github.com/rethinkdb/rethinkdb/issues/3579
- https://github.com/rethinkdb/rethinkdb/issues/3471
EDIT: danielmewes beat me to it.
" An alerting application is used to notify users when new content is available that matches a predefined (and usually stored) query. MarkLogic Server includes several infrastructure components that you can use to create alerting applications that have very flexible features and perform and scale to very large numbers of stored queries. "
I am not sure how commits on multiple documents work in RethinkDB, but MarkLogic the reverse query will be executed only after the commit (multiple or single-doc) will have happened (since MarkLogic is ACID).
Another system that supports similar thing is EllasticSearch with query percolation feature
http://www.elasticsearch.org/blog/percolator-redesign-blog-p...
(I am asking if there are any technical limitation...)
The hard part is building out a change feed that lets you synchronize on a subset of a table in a way that is performant, and to make this work with joins (i.e. when you're using the database as more than just a document store).
Looks awesome really want to use this on my next couple of projects.
I'm not sure about Swift. If anyone wants to give building a Swift driver a try, check out the "Contribute a driver" section at http://rethinkdb.com/docs/install-drivers/. One is pretty easy to build, and is a lot of fun!
Is there a way to get an entire table as the initial result set before getting update diffs? Something like:
r.table("users").between(-Infinity, Infinity).changes().run() // Not actually valid