RethinkDB: An open-source distributed database built with love over three years
rethinkdb.com
rethinkdb.com
Slava is a deep thinker, which makes me very excited to take a look at RethinkDB.
He helped get me into functional programming, which got me a contract job [1], which is how I met one of my current co-founders.
[1] http://martin.kleppmann.com/2009/09/18/the-python-paradox-is...
Kind of makes me wonder why C++ was chosen...
In a previous incarnation RethinkDB was a highly optimized storage engine for SSDs implemented in C++ to be able to take full advantage of both low level SSD and kernel access.
The current distributed engine was built on top on this storage engine and I think it only made sense to continue with C++.
They pivoted away from MySQL after my short stint in the beginning so I can't speak to why the storage code was kept (though I can't imagine it's because my code was so great they couldn't bear to throw it away).
Storage people tend to stick close to the metal, in general. This means C or C++ in most cases, for better or for worse.
A yc company hired me. I showed up at their mountain view office. The founder said "This is the former office of RethinkDB! I hope we are as successful as them."
I didn't know who/what RethinkDB was, so I said ok, sure.
3 days later he asked me to clear my desk and leave. He said "You are the sort of person who should work in RethinkDB".
So I asked "What does that mean ?"
He said "RethinkDB is trying to solve very deep algorithm problems. They want somebody with CS knowledge to do deep research. That is what you are good at. But here we are just trying to run a business. You are not a good fit for that!".
So I left.
I am going to go out on a limb here and suggest you try to work on being a bit more practical. Don't complicate things for the sake of solving difficult problems. Don't try to shower people with your engineering knowledge when it's not necessary, and don't expect everyone to know everything you do. And don't be an asshole about it either.
But sometimes someone with very limited knowledge of something asks me a detailed question about X.
What they ask is too difficult and complex to be described in a simple way. Either I have to overly simplify it which may insult them and will do no good or I have to go on and step by step give them digestible chunks of explanation which will inevitably be a bit technical even though I try to minimise that.
I think it's a two-way thing. The person should also consider their own level of knowledge before asking for an explanation of something and adjust their question based on that.
I don't ask my doctor to tell me why my heart does this and that because I simply don't have the knowledge to be able to understand his answer. I ask, my heart is doing ok? cool!
Reading, writing, and any sort of public speaking (especially debate if there's an opportunity) are TREMENDOUS for teaching you how to get into other peoples' heads, figure out how they're perceiving what you're saying, and adjusting what you're saying. I will sit in most of my CS classes and be PAINED when the class spends 10 minutes on a question because there is a disconnect between the misassumptions the student is making and the instructor thinks something else is the culprit.
These skills are also incredibly helpful in being charming and getting what you want without being manipulative.
[1]: Which is why some, like mine, created programs to blend CS and business to start to solve that problem. Students have to interact with paying customers, have to be accountable for their own code releases, are responsible for ALL of the requirement soliciting and fulfillment, etc.
The point is, if you don't ask questions in areas you are not familiar with, you will never become familiar to these areas. Well, unless you learn everything from books and Wikipedia.
I'm not sure how many people see me as an annoyance, but at least I'm consistent, in that if other people ask me questions in my expertise, I'm happy to try my best and explain.
Oh, and if it's just impossible to reasonably answer my question in a way that makes any sense to me, I expect you to just say it, and I'm happy with this.
I do what you described too, in parties or whenever I have an opportunity for a discussion with someone and I enjoy it and they ask me questions too and we try our best to teach each other something which is perfectly ok and fun.
What I was referring to was mostly employee/boss situations where the boss asks the employee about the details or internals of system X and then gets pissed when the engineer can't explain it to him and blames it on them because they were incapable of explaining complex things to non-technical people.
I mean they have to appreciate that there's a limit to how much you can explain to non-technical people in simple terms. At some point it just doesn't work, and either you have to use the big words, and concepts and assume knowledge, or drop back to dead simple insulting analogies. You see that server boss? That's like a train! choo-choo!
As an example, I like how Basho provides some comparisons[1] of Riak vs other popular options.
http://docs.basho.com/riak/latest/references/appendices/comp...
alex @ rethinkdb
Can you talk a bit about how RethinkDB compares?
I was looking at the github comments about a home brew recipe in which it was stated that aside from a recipe creating a VM, the Mac OS X port would take a bit longer.
Is that a full port from one language to another? Or just an issue of the different flavors of *nix that need dealing with and probably some of the dependency tree issues that come with it?
I'm curious what needs be done to get it building on Mac OS X — perhaps I could assist somehow.
I see a few dependencies that don't immediately sound familiar. You may have better luck with MacPorts, which uses tcl as the language for their portfiles.
Portfiles are just like homebrews recipes, but MacPorts always builds new, including the entire dependency tree ( and dependencies of dependencies etc., etc. ), for which they have thousands of working portfiles. Since those are completed and working, you wouldn't have to worry about those until you wanted to be able to make a binary outside of any package manager.
MacPorts can build binaries now ( new feature ), so you could just as easily instruct it to create a standard Mac OS X installer .pkg which makes sure everything goes in the right place, on the right platform, for the right architecture.
They are an exceedingly friendly and helpful group, I'm sure they would live to see this software in their package/portfiles list.
I am not a Mac user, but a designer using MacBook joined our team last week, and we struggled for half a day with Homebrew. The next day, we installed MacPorts instead, and with just:
$ sudo port install python27 py27-virtualenv gcc46
we were able to proceed and get the whole stack up and running. Not to mention everything from MacPorts is installed nicely under /opt/local.
MacPorts is just way ahead of Homebrew. OTOH, Portage is way ahead of MacPorts ;)
Maybe. Maybe not. But I don't think that's the reason homebrew doesn't have gcc. The OP is pointing out homebrew isn't extensive, and misses some commonly used utilities.
brew install rbenv ruby-build; /* rc file shenanigans */; rbenv install 1.9.3-p327; rbenv global 1.9.3-p327; ruby --version
I had to fight for days to get MacPorts to install anything properly. It gives me flashbacks to the horrors from 3-4 years ago of compiling open source software on Linux.
Homebrew has been fuzzy kittens in comparison.
$ brew search gcc
apple-gcc42 gcc
homebrew/versions/gcc45 homebrew/versions/llvm-gcc28
I'm a bit confused since I thought that these are gcc. Can someone tell me what those results mean? :S
$ brew tap homebrew/dupes
$ brew install gcc --enable-all-languages
To use your new gcc-4.7.2 when installing new packages just add '--use-gcc' at the end of the command.MacPorts seems better to me, after years of fink in the past. I need to build for universal (386/x86_64) for testing purposes, so it fits well for me.
I actually rebuild stuff later myself, since I can't really package stuff and require people to have that in /opt/local/bin or anywhere else, but a local folder to the main app.
(I use the same way cygwin on windows, like macports - I love the tools, the stuff, I test a lot of things, but afterall for things I want to distribute I compile myself, and post binaries).
I ran a fairly convincing "Linux-like" alternative desktop using Awesome in XQuartz for about a year before I switched to Linux full time. That was before Homebrew but I'm quite certain that it would have been impossible with it.
jedberg already asked for a compare/contrast, but let me provide some specifics I care about that you might be able to answer.
1. Is it fair to say that thanks to MVCC, running an aggregation or map-reduce job isn't going to lock the whole damn thing up like it does on MongoDB?
2. You've got a distributed system that is seemingly CP, do the availability/consistency semantics compare with HBase? Master-slave? Replication? Sharding?
3. Latency is a big one for us and is a large part of why we use ElasticSearch. How does the read-latency on RethinkDB compare with Mongo/MySQL/Redis/et al ?
2. Short answer: we favor consistency (via master/slave under the hood). It allows for much easier API, much fewer issues in production, etc. The user experience is just better. If you're ok with out of date results, you can do that too without paying the price of consistency guarantees. The downsite of our design is that you might lose write availability in case of netsplits (if the client is on the wrong side of the split). Longer answer: checkout the FAQ at http://www.rethinkdb.com/docs/advanced-faq/
3. Read latency should be equivalent to other comparable master/slave systems. We don't do quorums, so latency will be much better than quorum/dynamo-based designs.
In reality, most transactional database deployments are heavily skewed towards read workload, so reading from hot slaves is basically a requirement for master/slave databases. So, in most real world applications at scale, apps already deal with inconsistencies between slaves and the master and are making the "difficult" choice of dealing with CAP trade-offs. Asynchronous replication also creates a potential for difficult or impossible to recover from data loss in the sense that masters & slaves always have a continuous possibility for split-brain.
RethinkDB does not provide multi-shard transaction atomicity and/or isolation, which in my experience is the biggest difficulty thrown up in front of developers coming from single-node databases. I feel like the difficulty of dealing with inconsistencies across multiple versions of a single object is far more familiar as most developers have at least dealt with cache invalidation in some form. It's really having to ensure and deal with potentially out of order operations (inconsistency in the ACID sense) across a "graph" of data that's more insidious.
We set out to make these things be really easy (whether we succeed or not remains to be seen). We want the users not to have to deal with these issues at all whenever possible. You should be able to set up a cluster, add shards, and run cross-shard joins and aggregation in five minutes.
Of course once that problem is solved, there are tougher problems like high-performance cross-document distributed ACID, but I think the industry as a whole is relatively far away from that right now. (there are some solutions to this - e.g. Clustrix, but they require specialized hardware which makes it out of reach for most developers)
Regarding your last comment on high-performance distributed ACID, that's what we've built at FoundationDB, although FoundationDB is a key value store so transactions are multi/cross-key instead of cross-document.
Megastore and Spanner solve that problem, with varying tradeoffs:
This isn't "the future", this is now. People are doing it, and have been for awhile. If you're going to "rethink the database", distributed global consistency should be at the top of your list today. RethinkDB seems like its merely "rethinking" Mongo.
The main benefit of global consistency, of course, is ease of use. Global consistency is so much easier to reason about and write code for!
Currently if you write an infinite loop in js, or write code in a way where it starts eating up memory we don't do anything to restart the js process, but it would be relatively easy to implement.
* A far more advanced query language -- distributed joins, subqueries, etc. -- almost anything you can do in SQL you can do in RethinkDB
* MVCC -- which means you can run analytics on your realtime system without locking up
* All queries are fully parallelized -- the compiler takes the query, breaks it up, distributes it, runs it in parallel, and gives you the results
But beyond that, details matter. Database system differ on what they make easy, not what they make possible. We spent an enormous amount of time on building the low-level architecture and working on a seamless user experience. If you play with the product, I think you'll see these differences right away.
Note: rethink is a new product, so it'll inevitably have quirks. We'll fix all the bugs as quickly as we can, but it'll take a few months to iron things out that didn't come up in testing.
Also, I am excited to try this out. I always enjoyed your writings and I am sure you + team have made something awesome.
Does it means that every query touches all servers ? Or does it sends queries to only a subset of servers when possible ? (e.g. range queries on PK)
> Or does it sends queries to only a subset of servers when possible ? (e.g. range queries on PK)
The query planner distributes the query between the nodes that actually contain the relevant data. Here are a few examples:
In your example, a range get on the primary key, the query would touch one copy of each shard of the table. *
A more interesting example is a map reduce query. That query will also only touch one copy of each shard of the table but the mapping and reduction phases will also happen on those shards which makes the whole process a lot faster.
But shouldn't it be fewer than "each shard"?
Let's say the range is 3 < PK < 7. If all PKs in that range only lives in 2 shards (out of a total of say 10 shards) then the query should only be run in those 2 shards, no? Or will all 10 shards still be touched by the query?
While it's true that on a single node MongoDB map reduce is single threaded, it is parallelized when running on a sharded cluster.
A few questions:
1. Will secondary indices be ever supported? Range scan with a different order than the primary key is very welcomed. E.g. date range query.
2. Do you support conditional update? Or any kind of optimistic locking or versioning to coordinate concurrent updates from different clients?
3. Related to 2. How can loosely-sequential Id be generated using a table?
4. Will some transaction support be added? Don't need full ACID, just grouping updates (intra-table and/or inter-tables) in one shot would be nice. Should be feasible with MVCC already in place.
5. Do all the clients hit a central server to initiate queries which then farms out the requests to different shards? Or the client library knows how to get to different shards directly? First case has a single-point-of-failure, and bottleneck in scaling.
6. Do you support automatically re-balancing of shard data (data migration) when new shards are added or old ones retired?
7. How are authentication and authorization done? Or any clients can come in?
8. Internal detail. For out-of-date distributed query on the slave replicas, is there a cost-based (or load-based) decision process to pick the most idle replica to do the sub-query?
9. Internal detail. Do you use Bloom Filter to optimize distributed joins?
Secondary indices are one of the most asked for features so they'll probably be added in the next release. No promises though secondary indices are tough to do right and we won't ship them if they're not great.
> 2. Do you support conditional update? Or any kind of optimistic locking or versioning to coordinate concurrent updates from different clients?
Updates can be done with conditions on the row. For example: table.filter(lambda x: x['age'] > 25).update(lambda x: {"salary" : x["salary"] + 25)
> 3. Related to 2. How can loosely-sequential Id be generated using a table?
Loosely-sequential IDs would have to be generated client side for now.
> 4. Will some transaction support be added? Don't need full ACID, just grouping updates (intra-table and/or inter-tables) in one shot would be nice. Should be feasible with MVCC already in place.
Eventually. No concrete timeline for this right now though.
> 5. Do all the clients hit a central server to initiate queries which then farms out the requests to different shards? Or the client library knows how to get to different shards directly? First case has a single-point-of-failure, and bottleneck in scaling.
A client makes a connection to a specific server and all queries go through that server. However every server can file this role so connections can be distributed and there's no single point of failure. An even better option is to run a proxy on the same machine as the client. For more info run:
rethinkdb --help proxy
> 6. Do you support automatically re-balancing of shard data (data migration) when new shards are added or old ones retired?
Right now sharding is a manual process. You tell the server how many shards you want and it handles figuring out how to evenly split the data, picking machines to host them and getting the data where it needs to go. What it doesn't do is readjust the split points when the data distribution changes. This will be a feature in RethinkDB 1.3.
> 7. How are authentication and authorization done? Or any clients can come in?
RethinkDB has no authentication built in to it. You should not allow people you don't trust to have access to it.
8. Internal detail. For out-of-date distributed query on the slave replicas, is there a cost-based (or load-based) decision process to pick the most idle replica to do the sub-query?
Right now we just select randomly. This is slated as a potential upgrade for 1.3. Especially if it proves to be a problem for people. Thus far it hasn't been for us in profiling runs but this is the type of problem that's more likely to show up in real world workloads.
9. Internal detail. Do you use Bloom Filter to optimize distributed joins?
We do not currently use bloom filters to optimize this.
2. Yes. There is no special command, you just combine update and branch (http://www.rethinkdb.com/api/#py:control_structures-branch) Here's an example in Python:
r.table('foo').get(5).update({ 'bar': r.branch(r['baz'] == 0, 1, 2)})
This will set attribute bar to 1 if baz is 0, or to two 2 otherwise. Everything is atomic on that document.3. Currently the server doesn't support a sequential (or even loosely sequential) id autogeneration. You'd have to do that on the clients, but using a timestamp for example.
4. I don't know yet how to do this really efficiently. It's relatively easy to do on a single shard, but cross-shard boundaries make this really hard.
5. Any client can connect to any server. The server will then parse and route the query. There is no central server, everything is peer-to-peer. The client library doesn't know about multiple servers now, so responsibility is on the user to hit a random server. Alternatively you can run "rethinkdb proxy" on localhost and connect the client to that. The proxy will then route queries to proper nodes in the cluster.
6. In the web UI, if you click on the table and reshard, everything will be rebalanced. You don't even have to add or remove shards, it'll just rebalance data for the number of shards you have. The UI has a bar graph with shard distribution, so you can see how balanced things are.
7. Currently there is no authentication support - we expect users to use proper firewall/ssh tunneling precautions.
8. Yes, that's how queries get routed. Currently this isn't very smart, but it will get much better over time. If something breaks for you performance-wise, just reach out and we'll fix it.
9. No, not yet. If you run eq_join on a small subset of the data (99% of OLTP workloads) it will be very fast. Other joins work ok, but there's A LOT of room for optimization.
Phew!
For 2 and 3, I think I didn't make it clear. Let me clarify. A common db problem with multiple clients is dealing with concurrent update on the same piece of data. E.g both client1 and client2 read D as D=15 at the same time. Client1 adds 1 to D as 16 and saves it. Then client2 adds 1 to D as 16 and save it as 16, which is wrong. It should be 17.
Conditional update is one feature db usually provides to let clients deal with this problem, i.e. the update would only go through if certain condition is met otherwise abort. Update D=16 if D==15. Client1 would succeed while client2 would fail, where it can retry the whole read-increment-update cycle again with the new read value.
The litmus test to see if a db system can handle this problem is to try to implement a sequential Id generation feature run by multiple clients at the same time.
For 8, if the query is parsed into a query execution plan, you can ship the plan to all equivalent replicas to ask them to estimate the execution cost based on their current load. After they reply, pick the lowest cost one and send the execute command. Even a simple approach of asking for machine load of all replicas and picking the lowest one could have adaptive utilization of all the servers.
For 9, Bloomer Filter is a relative simple technique that can dramatically reduce the amount of data to ship across peers to do join. You basically filter out the vast majority of the non-matching data before shipping.
It's a good start. Good luck going forward!
r.table('tv_shows') .filter({ name: 'Star Trek TNG' }) .update({ episodes: r('episodes').add(1) }) .run()
The scenario I described has to do with read-consistency, where the value read by a client should not be changed during the time of the read and the time of the update. The usual way of handling it was to take a write lock for the duration to prevent update from others but that degrades concurrency. The other way is to do optimistic lock (or conditional update) to allow the client to detect change during the time and retry with the new value.
E.g. the client reads a value, displays to the user, gets input from the user which is based on the old value, and stores the updated value. If another user doing the same thing has already changed it, the client would like to know that and let the user retry, with the new current value.
r.table('foo').get(5).update({ 'bar': r.branch(r['baz'] == 0, "foo", r.error("invalid baz!"))})
have not tested it, but this is how I understand it... r.table('foo')
.get(5)
.update({
'_rev': r.branch(r['_rev'] == 5,
r('_rev').add(1),
r.error("invalid revision")
),
'name': "awesome name"
})
the basic idea is that `name` should be update to "awesome name" and `_rev` should be incremented by 1, but only if `_rev` is 5, otherwise an "invalid revision" error should be thrown.* How does rethinkdb compare to MySQL Cluster? Both are distributed, replicated databases with a sql-like query language.
* Any plan to offer a java client?
* RethinkDB has flexible schemas and a query language that integrates straight into the host programming language and doesn't require string interpolation. As far as clustering goes, RethinkDB is a) really really really easy to use, and b) does a lot of query parallelization and distribution that MySQL cluster doesn't do. The product feels totally different, I think in a good way. The downside, of course, is that rethink is new and it will take some time to work out all the kinks.
* I can't commit to a timeline yet, but yes, absolutely.
This feels new and refreshing, I hope things turn out positively for you.
Sure, you can put that stuff in strings, but then you'll run into limitation with queries where you want to, e.g., aggregate a total, or do timestamp arithmetic.
I could do everything with strings, custom map-reduce, etc., if you're inclined to suggest that as a workaround. Still doesn't mean JSON's a good idea.
A similar problem exists if you use JSON numbers (aka doubles) for timestamps –- the numbers just aren't big enough to do it accurately.
(But what financial software is running on a NoSQL database?)
I didn't pose the question "what financial app needs >53 bits of precision?"
Clustered databases are essentially a solved problem, and have been for years. What's needed today are databases solving the problem that Google Spanner addresses – global consistency across distributed clusters in separate data centers. If you want a challenge in the DB world, that's where it is.
But another clustered, schema-less JSON database? Might as well open up Intro to Algorithms and run through the exercises -- it's no longer a challenge, algorithmically or otherwise.
Sorry to be a downer on this, and it does still take a strong coder to implement one, so well done on that front. :)
That said, with a name like RethinkDB, I guess I expect more than a feature list I could have reasonably put together three years ago and gone, yeah, that's straightforward to do.
I've written my own database (and continue to improve it), so I'm pretty familiar with the issues involved. You're absolutely right that many of these JSON database have serious problems under load with their clustering abilities (and it's always under load, they tend to work fine on simple workloads).
Perhaps RethinkDB can carve out a niche for reliability-under-load among the existing JSON DB field. That's got to be worth something.
Reminds me of Freud's story about the peasant who says to another, "Hey, you broke that kettle I lent you", and the other says, "It was fine when I gave it back to you, it was already broken when you lent it to me, and I never borrowed it."
Why? Analytics are CPU hogs, tend to access tons of data in random fashion (blowing caches and hogging the SSD drive), and given that RethinkDB has no secondary indexing, are likely to be especially slow.
That's why people have separate machines dedicated to analytics. What I think a team would actually do with RethinkDB is the same thing people do with Cassandra: include a separate cluster (in the same or a remote datacenter) and replicate data to it from the transactional cluster(s). They would then run analytics on the analytics cluster.
This approach won't impact transactional latency, and also allows you to have different hardware altogether for running analytics (e.g. tons of cores and RAM that might go wasted on the transactional DB machines).
This is all Big Data 101; it's not controversial.
alex @ rethinkdb
1. How much data can you put in one instance before seeing performance degradation? I know that you still working on good benchmarks – but do you have any ballpark figures?
2. How does replication work? Is it closer to row/document or statement based (or something completely different)? How fast is the replication?
3. What is your envisioned used of the replication? Are replicas supposed to serve read traffic, or their goal is to keep the data safe in case of a catastrophe?
4. Can you tell me something more about cluster configuration propagation? The Advanced FAQ answer doesn't get into much detail.
5. Am I correct to assume that you are using protocol buffers? What motivated your choice?
2. We do do block-level replication. On each node of the btree we store replication timestamps. When a node asks for new data, we can cull away parts of the tree the node has almost instantly. So replication is very very efficient for most OLTP workloads. We don't have statement-level replication yet, so if you do a range update on a large table, we'll have to replicate data block by block. It'll take a while to add statement-based replication - we'd have to do a pretty significant refactoring to make it happen.
3. Either. Replicas are great for failover -- if the master dies, you just failover and a replica picks up where the master left off. If you're ok with out-of-date reads, you can also hit replicas directly (e.g. for reports, etc.) and spread out the read load across the cluster.
4. This is a really complex question - we didn't document this because doing it properly would take a lot of time. I'll ask jdoliner to chime in -- he designed the architecture and wrote most of the code, perhaps he can describe it succinctly while we write deeper docs on this :)
5. We use protocol buffers between the client drivers and the server. We picked that because there were libraries for the initial three languages we picked (Ruby, JS, Python), they were really easy to use, and very efficient. We could also have a single spec for the client/server API. Internally we use our own serialization scheme which allows us to dump arbitrary C++ objects on the network. It doesn't support other languages (which we didn't need), but is much more versatile for writing complex cross-machine code.
Short answer: Our configuration data is most similar to git. Any machine can be used as an administrative node via the WebUI or the CLI. It will make changes to the metadata which then get pushed to the other nodes. If 2 nodes make conflicting changes you get a conflict which the system will help you to merge.
Long Answer Cluster configuration is stored in semilattices which are a neat mathematical structure with a few very desirable properties. Semilattices support have a join operator. For our cluster metadata joining is the means by which metadata is updated. When one server connects to another the two swap metadata and each joins the other's metadata into his own. In essence learning what the other knows.
There are two properties in particular of the joining that are nice. First off joining is commutative. This means machines can exchange data in whatever order they want and get the same result at the end. Secondly they're indempotent. That means machines can resend their data without fear. The value doesn't change if the same value is joined in twice. These help us with a lot of the worries of distributed systems.
If I understand correctly, the client can connect to any instance and its request will get routed appropriately. Let's assume that you take a master offline and promote one of the replicas to be a new master. Won't that lead to a window in which (from the point of view of different instances) there are two masters at the same time and some writes are sent to the wrong instance?
EDIT:
One solution for such things is to use something like Zookeeper (or some other system whose documentation mentions "Paxos" ;)). Have you considered that? How does what you are doing compare with that?
We basically have something very similar to zookeper baked into rethinkdb. We wrote it internally from scratch to better suit the needs of our architecture.
Does anyone know if something like that exist?
The only docs I found in the company website that goes deep into the internals are Advanced FAQ (http://www.rethinkdb.com/docs/advanced-faq/). It is more of an architecture view, though.
The reason I ask is that with a good understanding on the internals, the engineers who understand database internals and distributed systems will have an "more" accurate idea on the capabilities and the limits of the features. Thus, if they decide to adopt RethinkDB, the understanding will help them design their applications to take advantages of the benefits and avoid the potential issues (or surprises!). MongoDB was not very good at documentation. It claims this or that feature works smoothly. Then, people found out many potential issues and limitations. That is one reason it leaves a bad tastes to many engineers.
-harryh
1. Yes, I know there are other ways to do this besides hashing the shard key, but this is often the best way.
We'll be addressing this at some point, we have to sort through the list of feature requests first. It's a long list :)
Is this just a hipster marketing term to tell us that it's small and cute and made by people who play ukuleles and ride unicycles in their spare time, and not by evil corporate people who commute to work and have mortgages?
I find a lot of advertising eyeroll inducing, and the current trend of more-hipster-than-thou posturing is right at the top.
Get it now?
When you have a vision of something great that ought to exist and set about bringing it into the world, you are in an isolated position: other people don't yet see what you see. This leads to a lot of doubt by others and by yourself too. The longer it takes, the more exposed you are. To make it through that you are going to need a deeper source of motivation – an underground spring. Love is a fine word for this, and it makes me happy that Slava put it in his title: it's a clue to this experience that rarely gets mentioned, especially in the land of pivots and MVPs and weekend hacks.
Rule of thumb is if you build something this nice and with that order of magnitude in complexity you can put My Little Poney stickers on your homepage and still get respect. Who cares about the "attitude" and the "language" for Christ's sake, they BUILT stuff with their own hands and are offering it to the world, they can do whatever they damn please.
I particularly like the perspective of an easy onramp to get started, knowing that I will never have to leave because of scale or reliability.
Please, please give me a SQL adapter! My marketing team needs SQL. My business app developers need SQL. Give them an adapter and I will get them to use RethinkDB - knowing that 1) my data is safe and I'm not 6 months away from a painful re-architecture and migration, and 2) as my developers hit the limits of SQL they can gradually (gradually!) peel the paint off and start using your more powerful query language.
Schemaless is clearly a convenience win over SQL because SQL's way of modeling nested/repeated data doesn't map as easily onto programming languages. But for all the people who are using JSON-based databases these days, I'm curious how many of them couldn't easily write a JSON schema or a .proto file that describes their de facto schema.
I ask because a lot of things become easier to reason about (and optimize) if you know that a field won't be a string in one record and a number in another. And writing a .proto file (or equivalent JSON schema) would give you an authoritative place to document what all the fields actually mean.
I don't have any actual experience with JSON-based databases, so I was interested to hear the opinions of people who do.
That doesn't mean I want to deal with the implementation detail of columns, but I definitely wouldn't mind some type safety.
When a filter is written in the client language it gets compiled into a protocol buffer which is sent to the cluster. This gets compiled into a query which is sent to each of the relevant shards for the table. This query has the filter baked right into it. The shards then go through their local copy of the data and filter out the rows which do not meet the query predicate. This data gets returned to the coordinating node and eventually to the user. Thus only the data the will actually be returned is ever transferred over the network.
Furthermore this process is done lazily. On the client side rather than getting back a huge array with the results of your filter you get back an iterator. This iterator stores a buffer of data which will be refilled as it is incremented.
It's rather difficult to integrate into a host language like that smoothly from the driver implementation perspective, but once the driver is written the user experience is amazing because you can write queries that look exactly like Python, but they're executed entirely on the server.
r.table('foo').update(lambda row: {'bar' : r.table('bar').get(row["bar_id"]))
This still works but gets evaluated in a different way to make sure every secondary winds up with the same value.
The use of r in filter is getting the attribute bar of the row.
Have you found that there are useful expressions that are awkward to express without the lambda trick?
Lambda syntax is really nice too, I actually prefer it for writing queries.
Again taking an example from SQLAlchemy, you can explicitly make a subquery, and then reference it instead of the original Table. A binding more like SQLAlchemy can probably written for RethinkDB.
Also, have you talked to the Meteor folks about swapping Mongo out for this? Or would this be 'newness overload'?
Sure they can take the shortcut and just produce some single case and claim 'faster than a.n. other db' and that will probably satisfy the 'my db is faster than your db' brigade.
However, based on the effort that appears to have gone into this product to get it to this stage, they probably wouldn't be happy with that.
More power to them!
In the meantime, you can always chat with us and we'll help you work through any issues you might run into.
I feel being based on JSON is a big con though. While it's popular, it was never meant to be a rich serialization format, just simple. How to implement more complex fields like dates, and query efficiently on RethinkDB?
We had to start from somewhere and we also wanted to get a feel of what are the most requested libraries so we can focus our energy in those directions.
alex @ rethinkdb
2. We're putting together some comparisons, hope to add them to the site soon.
3. There are already a couple of answers to this question on this thread:
https://news.ycombinator.com/item?id=4764137
https://news.ycombinator.com/item?id=4763939
alex @ rethinkdb
Congrats on shipping!
joe@alchemist~$ rethinkdb
joe@clockwerk~$ rethinkdb -j alchemist:29015
Dota player?I play DOTA 2, by the way, hopefully you do too.
With some simple formatting functions (its just json after all) it's sixes to me.
EDIT: also, you can run queries with an out_of_date_ok flag, which will give you what you want. This only works for read queries though, the architecture is pretty much set up in a way where this would be very very difficult to do for write queries.
And JSON has the huge advantage of supporting hierarchical data -- arrays with objects inside, etc. It seems a like a huge step forward.
Also, you seem to be confusing primitive data types with complex data types. Yes, JSON doesn't have a 'color' data type. But guess what? Neither does C, nor Java. If you want a 'color' type, you'll have to create one yourself! Mind blowing, I know.
So here, let me suggest a possible solution:
dates: string
timestamp: integer
colors: string (RGB,BGR,RGBA,...), integer, object
Part of 'data modeling' is to model your data out of basic types. Shocking! If JSON had types for every type of object under the Sun (like you seem to want), JSON parsers would be a lot more complicated and little thinking would be required in the process of modeling your data.
GH issue https://github.com/rethinkdb/rethinkdb/issues/2
edit: fixed already!
Will doing a query like "age > 25" perform something equivalent to a full table scan?
All of that is only good for relatively small amounts of data though, so we'll be adding secondary indexes soon.
-Joe Doliner, engineer at RethinkDB
Disclosure: Ok. I'm biased. :) I've designed a similar DSL-style query API in another project. https://github.com/williamw520/jsoda
The Github graphs are really interesting too, that's a lot of love/work right there!
edit: Indeed, my architecture i386 doesn't match the only available amd64 binaries. Thanks
EDIT: the main thing missing from earlier ubuntu versions is TCP_USER_TIMEOUT. We can work around it in the server, but we haven't done it yet.
Did you rethinkAuth or am I just too stupid to RTFM?
We're also looking into launching services, which is a great revenue stream for people who prefer to pay for convenience of not having to deal with operations at all.
One note, there's a typo in the code in the tutorial on top
r.table('users).insert({'name': 'Slava', 'age': 29 }).run()
users needs a closing quote.
A C driver would definitely be possible, but would be a little bit clunky.
In all seriousness while we'd eventually like to have support for every language Arc is farther down our list than others such as PHP and Java. We will get to it eventually though.
If Arc could be made to use protocol buffers it wouldn't be too hard for a contributor to write the driver themselves though.