RethinkDB 2.0 is amazing
rob.conery.io
rob.conery.io
Just my 2 cents of Mongo hate :-)
No software system is perfect, but there are definitely practical balances to be made. Especially when you are beyond what a single database/server can offer in terms of write throughput. The fact is, when your traffic needs exceed what a single database can keep up with in terms of writes, you have to give up some level of reliability.
Granted that you can only prove that the system is vulnerable and not the reverse, but if there is a vulnerability it is much harder to trigger it.
Postgres is a good DB, but since it's not distributed, it's not very useful to compare it to distributed databases. Yes it's consistent, but it's only as reliable as the single node where it is installed.
[1] http://www.datastax.com/wp-content/themes/datastax-2014-08/f...
All of that said, you have to take a research paper funded by a database company (Datastax is backing Cassandra) with a grain of salt. Not to mention, that most people reach for MongoDB because it has some flexibility, and is a natural fit for many programming models. Beyond this, setting up a replica set with MongoDB was far easier than with any other database I've had to do the same with... Though I'd say getting setup with RethinkDB is nicer, but there's no automated failover option yet.
They also were quite generous by comparing load using non-durable write for CouchDB, HBase and MongoDB against Cassandra's durable write.
From my personal experience many scaling problems that you have with MongoDB once you switch even to a relational database that can't scale out are laughable.
Regarding Mongodb, all I'll say is that I've switched from mysql to mongodb 2 years ago, and I've never looked back. YMMV.
I'm also a user of ElasticSearch and Redis, and looking to add Couchbase to the lot. One size doesn't fit all. mysql and postgres certainly don't fit all either.
Also saying that Postgres cant scale horizontally is not entirely true, it in fact can[2], it is currently more complicated but I learned something when I was investigating how our applications would behave with Postgres backend. Turns out that every instance we had mongo we could run postgres on a much smaller instance. In one instance the data was so laughable that you could just run postgres on the same node that was running the app.
The point of it is that even if you think that you need to scale out, unless you're Google, Facebook or similar company you chances are you don't.
[1] http://www.enterprisedb.com/postgres-plus-edb-blog/marc-lins...
I suspect he did not go over replication, because Postgres technically still fail over support is DIY, although he should. There are two replication methods though which I would like to see:
- asynchronous - this one is fast, but it most likely would have similar issues the other database have - synchronous - the master makes sure data is replicated before returning to the user this should in theory always consistent
You would typically have two nodes in same location replicating synchronously and use asynchronous replication to different data centers. On a failure, you simply fail over to another synchronously replicating server.
Regarding consul/etcd actually those technologies did not do well in his tests, but authors appear to be motivated to fix issues.
That's why I said it's unfair to say "postgresql did well".
Call me maybe is supposed to test all the difficult problems of CAP, which have not been tested at all with Postgresql.
And it's pretty cool that you can make a choice between fast writes or safe writes. You can't have both but at least you can have the choice.
Having said all that, I generally prefer Postgres in almost every possible case.
However, this RethinkDB project looks sexy and with a great potential.
the engineering skill here is the ability to trade off risk vs benefit... I will tell you from my own experience the best designed software systems I have personally dealt with tend to use components somewhat behind the curve.
For me the best software systems are those that are well architected and use the best available technology. This doesn't mean we all should be writing Tomcat, Oracle, Apache stacks just because they are less shiny.
As a small startup, you might not have the time for extensive testing and patching of new hot technologies, which facebook, twitter, netflix do.
And not sure if you've worked for a large company but they largely comprise lots of little startups sized teams. The same principles apply regardless e.g. spiking technologies out, managing risk etc.
What I have an issue with is these stupid generalisations. Less shiny = good, Shiny = bad. The merits of the architecture and technology seemed to be completely ignored.
Which then translates to buggier software as you add more and more new features.
There is a reason we rewrite codebases, no ?
With less shiny I mean that it's old. Not that it is crummy qualitywise. Old code is not like wine. If it was crap then it will be crap now. Old code is like an old house. If it's well made and tended to, and built on solid principles it can last generations.
Extensible domain logic is something that generally does not age well. Old utility libraries with well defined interfaces, on the other hand, are invaluable in technical computing.
If you don't run at the same scale as those guys you won't hit those same limits. But if you do reach the scale of those guy's you will find that your needs are suddenly very much a unique snowflake that will require you either creating something new or heavily tweaking something that already exists.
Plenty of time for that when you reach the scale that justifies it though.
The point is, if you are going to bet your business on a technology, it helps if it has been tested with production workloads in many different conditions and scales. You want to know about as many shortcomings as you can. For many of the use cases that people use things like Cassandra for, they will be tolerant of 30ms++ reads and potential read inversions. Redis is used pretty heavily, but it is relatively simple code and you can trace through the entire writepath pretty easily and get a sense of its limitations (being single threaded is a blessing and a curse, you really need to be careful about bad tenants because a single slow query will cause an availability event for everyone). HBase is used in a few places, but usually only for cold data after they expect it to be read only occasionally and they don't want to use up space on their MySQL pci flash devices for it anymore. There are a bunch more, but they all have some latency, consistency, or availability downsides compared to a traditional sharded B+ tree backed transactional store.
And I can't imagine any of this is relevant today given that MongoDB allows for pluggable storage engines.
I always find it amusing when people bring these issues up because it's like a giant sticker on their forehead that says "I've never actually seriously spent time with MongoDB before". I always go through the configuration of the database I use to make sure it meets my needs. Only seems sensible.
Thats dubious. 1.) When MongoDB was released, none of the drivers used the "safer" settings. 2.) 10gen, at the time, released benchmarks with the "unsafe" settings comparing it to MySQL and boasted that MongoDB was much faster (ignoring the fact that it wasn't acknowledging your writes).
AFAIK, until recently (i.e. the last month), there weren't any such benchmarks released by MongoDB - and then, only for 3.x.
I'd be very surprised if any such benchmarks exist, as you claim.
Disclaimer - I work for MongoDB Inc.
Although to be fair, it's not just MongoDB that is performing poorly.
Almost every single IT company exaggerates claims, says their product is "the best" and "amazing" and suitable for every use case under the sun. Oracle did it. Microsoft did it. Mongo did it. And thousands of companies in the future will do it in the future. It's called Marketing.
And I think you should speak to all these customers (http://www.mongodb.com/who-uses-mongodb) and tell them they don't serve any real world need. I would imagine a few would ask the same question of you.
The day we start jailing people for their marketing hype is... well, it wouldn't be a good day.
Can anyone explain (SQL pun not intended) to me the advantages / disadvantages between rethinkDB and say PostgreSQL?
A lot of the things would also apply to PostgreSQL
... what? ReQL and SQL are both declarative query languages: I don't really see the author is getting at. Is there an implication that SQL isn't declarative?
The only real difference is that the API is based around chaining function calls rather than expressing what is needed as a string - there are many SQL query builder APIs that will let you build SQL queries by chaining together function calls.
The biggest problem with composing SQL strings is that you have to be very very careful about SQL injections, and if you deal with that in a slightly sophisticated, reusable manner you are half-way to an ORM already. As far as I can determine, the ReQL drivers make injection attacks very difficult.
Using a query builder (or ORM) of some sort still allows the escape hatch of raw SQL to do those really crazy things that are sometimes needed for performance, or just because what you are trying to do is rather weird. SQL is a very mature language, it's unlikely you are going to run into something someone else hasn't before.
It seams like a good idea to allow diffrent layers of data access.
This should also be possible to build on top of ReQL, though I don't know any examples.
This is what stored procedures and parameterized queries are for. Even if I am going to do dynamic SQL, I do it in a stored procedure if I can.
I still don't see how you can pass user input from, say, a python string into a stored procedure call without worrying about injections. Or converting between your app's data structures and whatever string is necessary for your stored procedure.
query('SELECT * FROM users WHERE id = ANY ($1::int[])', [1, 2, 3]);
query('SELECT * FROM users WHERE lower(uname) = lower($1)', 'foo');
Where's the injection vulnerability?The challenge is to make sure the column/field `age` is more than 20.
My code is:
query.filter(r.row['age'] > 20)
What's yours? (Hint: start by writing a compliant SQL parser) query.where('age > ?', 20)
If I'm understanding you correctly.I'm also not sure how your comment replies to army's point. The point, as I understood it, is that it is not accurate to characterize SQL queries as steps that tell the engine what to do. SQL is declarative, and leaves the execution plan up to the database itself. army's comments about the API and strings were trying to point out the only perceived difference, which is not relevant to the question of declarative versus imperative.
"SELECT * FROM table WHERE age > 20"?
SELECT * FROM ($S) AS FOO WHERE FOO.age > 20"You already have an existing ReQL query in a string variable. You need to add the age > 20 condition to that query."
Same problem. Comparing apples and oranges, strings and some "live" code. If you put ReQL and SQL into the same category (either as a string or as a thing that represents some "live" running code that you can manipulate at runtime) then it is difficult for me, at least, to really grasp what the differences are between them. SQL is certainly not considered an imperative language, eh?
----
EDIT to respond to the comments below from TylerE and pests:
Oh but you do have ReQL as a string: when you type it into the editor, when it lives on disk as a file of source code. At some point that code becomes live and you can interact with it. The exact same basic transformation happens whether the syntax is ReQL or SQL, just in different ways and at different times depending on how you choose to run it not what syntax it's in. The issues are orthogonal and it certainly fair to demand that we compare the right things.
If you want to say that ReQL is a better syntax than SQL, well, I don't see it (yet.)
If you want to say that the product in question provides a nice way to run ReQL syntax queries in some fashion that is fundamentally better than the way that some other product allows you to run SQL queries, that is a whole different issue (and NOT the one I am addressing in my comment above.)
I hope that makes sense. ;-) Cheers!
Edit to your edit: It seems you are fundamentally not getting it. The ReQL is live code in your native programming environment. That means you can inspect it and manipulate it. SQL doesn't get interpreted (or whatever, it's black box) until it hits the server.
Imagine you're in a world where there are no XML parsing libraries. SQL is a string containing XML. ReQL is a DOM object.
One is much more useful than the other.
You are a comparing a language (SQL is independent to the language you're programming with) to an API.
RethinkDB has API available for three languages: JavaScript, Python & Ruby. If you take look carefully while it tries to be consistent across them, there are still parts that are specific to given language. If you would want to use RethinkDB with a language that is completely different (for example a functional language), assuming RethinkDB would support it, you're guaranteed that the interaction with the DB would be completely different, while you could still use the same SQL language[1].
If you want to compare RethinkDB's API to something similar you should compare it with something like JOOQ[2].
Just to preemptively respond to argument about translating DSL to SQL. Currently modern driver communicates with database using binary protocol, the SQL is compiled on client side. You could actually skip SQL altogether, but then you would lose flexibility of being able to support many other databases.
[1] http://pgocaml.forge.ocamlcore.org/
[2] https://en.wikipedia.org/wiki/Java_Object_Oriented_Querying
(Expanding upon that: We are both correct but not in the same context. There's a context in which what you are saying is true and sensible, and there is another context wherein what I am saying is true and sensible. I can switch between the contexts, so I am not trying to disagree with you, I am trying to give you data to help you to understand this other context and switch between then too. Additionally, this other context is of a higher "logical level" in the mathematical sense than the one we already have in common, and so when you do grok it I can confidently predict both that it with blow you mind and improve your ability to write software.)
Imagine me saying "here's a SQL string, let's see which database can execute it more easily, Rethink or Postgres. Hint: Start with writing a parser to convert it to ReQL".
(Hint: start by learning metalworking)
The query in RethinkDB is very much an expression. In the JavaScript driver you build this expression with function calls. There are other drivers which let you build the expression in a much more declarative way (like my Haskell driver).
I think the current status quo for databases is canned software. And this isn't necessarily bad because neither of the three databases mentioned hide their specs or default settings, the three have very good docs and community willing to help, in addition to companies giving commercial support. Whats your excuse to misuse these products?
RethinkDB writes your data to disk before acknowledging the write but on the other hand can't elect a new primary in case of failure, two completely different features/limitations that might work for someone and not for other ones. Is that hard to understand? Did mongo documentation lie you at some point?
Accept that you are "buying" a general purpose product, the designers thought that their users will need those features, deal with it.
Otherwise build your own database, I know this might sound very hard but I guess in the future we will see smaller building blocks that let you build something that handle your needs like this:
Right now, the situation with database is that we have to convert our internal data structure into a representation that fit the data model of the database we're using (ie rows for relational, document/key for the NoSQL group). I can see the reason the data model has to be that way for scaling, distributed etc... But if I'm happy to scale my database up, and would prefer to have the database storing the data as close to the memory data structure as possible (similar to object databases -- albeit with a boarder definition of "object"), is there any database that could do that?
Otherwise, is there any suggestion on how I could get started to build one?
This is the default, but also note that durability is configurable on an operation-by-operation basis.
http://rethinkdb.com/docs/troubleshooting/#my-insert-queries...
Good defaults are expected in quality software, and are just as important as any other part of software interface, CLI or GUI.
That said, I love the look of Rethink and I can't wait to give it a try.
The hard part has always been verifying that the data is actually persisted to the hardware. And the number of layers between you and the physical storage has increased not decreased. And the number of those layers with a tendency to lie to you has increased not decreased.
For some systems it's not considered to be persisted until it's been written to n+1 physical media for exactly these reasons. The os could be lying to you by buffering the write, the driver software for the disk could be lying to you as well by buffering the data. Even the physical hardware could be lying to you by buffering the write.
In many ways writing may have gotten more reliable but verifying the write has gotten way harder.
Possibly Oracle had fixed 100% of that by the time MySQL came out, but now we're just talking about the timing of adding in safety, again -- and both IBM and Stonebraker's Ingres project (Postgres predecessor) had RDBMS with ACID in the late 1970s, and advertised the fact, so it wasn't a secret.
Except in the early DOS/Windows world, where customers hadn't learned of the importance of reliability in hardware and software, and were more concerned simply with price.
Oracle originally catered to that. MySQL did too, in some sense.
In very recent years, it appears to me that people are re-learning the same lessons from scratch all over again, ignoring history, with certain kinds of recently popular databases.
why we still discussing it at a tech forum in 21st century in Silicon Value? Shouldn't it (together with isolation, ACID, CAP, etc...) be a base knowledge taught in elementary school? Like you can't expect Daddy to buy you a firetruck that Mommy promised if Mommy hasn't been able to talk to Daddy yet... though until Mommy talks to Daddy you probably can convince Daddy to buy you a railroad...
# if execution continues, everything agrees on the state of the transaction
# if execution halts, because of a crash or whatever, you'll come back online at a consistent state from the past
Mongo lets the user decide whether or not to wait for fsync when writing to an individual node. This is not the default configuration. If you want it, you can enable it. You may complain that Mongo has bad defaults for your particular use case. It continues to have bad defaults to this day. Saying Mongodb is unable to acknowledge writes to disk is pure FUD.
Let the downvotes ensue.
That's one opinion fitting one set of use cases. There are plenty of use cases where speed is more important than durability.
Hell, Redis default configs don't enable the append-only log, but you don't see the HN hate train jumping all over Redis. This is because Redis use cases typically don't require that level of durability.
edit for source: cmd+f for "appendonly" https://raw.githubusercontent.com/antirez/redis/2.8/redis.co...
> Let the downvotes ensue.
There's a lot of FUD going around when it comes to Ford Model X car not having brakes enabled. Please read the manual. Ford Model X lets the user decide whether or not to enable brakes or not. [...] It continues to have bad defaults to this day. Saying Ford Model X is unable to brake is pure FUD.
Let the downvotes ensue.
Firstly MongoDB's write durability was set to use the safest option on all of the drivers at the time. So your point makes no sense. And secondly we aren't ignorant users of the system. We are highly technical and as such your analogy again makes no sense.
Can we not have this reddit-ism take hold?
Emin Gün Sirer summarized[2] it best:
> WriteConcern is at least well-named: it corresponds to "how concerned would you be if we lost your data?" and the potential answers are "not at all!", "be my guest", and "well, look like you made an effort, but it's ok if you drop it."
Methinks the world has forgotten that high throughput systems existed long before the web of recent years. Most of what the web world thinks is high throughput is hilariously slow. The ability to run up another instance to scale sideways has ruined people. It doesn't scale in a linear fashion.
a.for("album").in("catalog").filter(
a.eq("album.details.media_type_id", 2)
).return("album")
Or in plain AQL (ArangoDB's query language): FOR album IN catalog
FILTER album.details.media_type_id == 2
RETURN album
The "map" in the second example is simpler, too: ….return({artist: 'album.vendor.name'})
Or in plain AQL: … RETURN {artist: album.vendor.name}
Also, it doesn't really need drivers because the DB uses a REST API that works with any HTTP client.That said, the change feeds are pretty neat and RethinkDB is still a pretty exciting project to follow.
(Full disclosure: I wrote the ArangoDB Query Builder without any prior exposure to ReQL, so I may be biased)
r.table('users').filter(function(row) {
return row('age').gt(30);
})
Could be expressed as: r.table('users').filter(r.row('age').gt(30))
That being said aqb looks pretty cool and quite similar to ReQL.https://github.com/arangodb/aqbjs/blob/v1.10.0/README.md#aql...
// one trip to database using subqueries and Postgres' JSON functions
con.SelectDoc("id", "user_name", "avatar").
HasMany("recent_comments", `SELECT id, title FROM comments WHERE id = users.id LIMIT 10`).
HasMany("recent_posts", `SELECT id, title FROM posts WHERE author_id = users.id LIMIT 10`).
HasOne("account", `SELECT balance FROM accounts WHERE user_id = users.id`).
From("users").
Where("id = $1", 4).
QueryStruct(&obj) // obj must be agreeable with json.Unmarshal()
results in {
"id": 4,
"user_name": "mario",
"avatar": "https://imgur.com/a23x.jpg",
"recent_comments": [{"id": 1, "title": "..."}],
"recent_posts": [{"id": 1, "title": "..."}],
"account": {
"balance": 42.00
}
}I've found the documentation a joy to use
The composable queries alone are enough to make any developer happy. You can pretty much treat your data as if it's in-memory because the drivers integrate so well with the language. The relational model works really well.
Things I have not tried yet are clustering and the real-time support (still need to build this into the lisp driver) but I'm trying to slot some time to do this. One of the projects I'm working on (https://turtl.it) is going through a nice upgrade to mobile right now, and this will include some server changes...I'm looking forward to implementing changefeeds to solidify the collaboration aspect.
Overall I've been really impressed with Rethink over the years, and can't express how excited I am they hit production ready. On top of the DB being great, the team is really nice to work with. They are incredibly responsive on github and were really helpful when I was first starting to build out my driver.
Great post, and congrats to the Rethink team!
My largest project has been running a couple of years now and has accumulated a significant amount of data, and RethinkDB hasn't had any trouble at all scaling with my data growth. I'm running it on servers below the recommended requirements too (512MB DO instances) and have been really impressed with how it handles constrained resources.
"RethinkDB is amazing" - TBD
I don't even think Slava would call RethinkDB "amazing" yet. I have no idea how to make a database, but I know there's a lot of work - and even more trial and error - that goes into making one "amazing."
This is certainly a big step for RethinkDB. But I'd be careful to put Petabytes of data across 200 nodes sharded 500 ways each.
Someone should make an index of such blog posts.
I do see that RethinkDB has some "Overview" and "FAQ" links on its website. However, when I encounter a new technology, I like to read its Wikipedia entry first. Wikipedia is usually more impartial, informative, and actually makes it easier to get a high-level sense of a technology than the tech's own website in most cases. This has grown more and more true over the past five or so years, as even developer-facing websites have devolved into "startup-y" marketing nonsense.
I wonder if there WAS a Wikipedia entry, but it's been deleted by some moderator with an axe to grind? I personally haven't contributed in years due to how unpleasant it is to add new content through all of the Wiki-lawyering. I've also noticed that 5 years ago, when you did a Google search you could rely on the Wikipedia entry being one of the top 2 or 3 results. Lately I see more and more instances where I have to scroll to the second or third pages of results to see a Wikipedia link.
Anyhoo... apologies for the tangential aside. I'm just wondering whether the lack of a Wikipedia entry says more about RethinkDB or about Wikipedia?
I can't find evidence of a deleted Rethink article in Wikipedia, but didn't look hard.
https://en.wikipedia.org/wiki/RethinkDB
Reason: https://en.wikipedia.org/wiki/Wikipedia:Criteria_for_speedy_... (G11. Unambiguous advertising or promotion)
Is there an archived copy of the original page? It's probably best to start with a stub page that contains no advocacy for Rethink and minimal information, and then grow it over time.
> 17:41, 16 August 2013 Alex Shih (talk | contribs) deleted page RethinkDB (G11: Unambiguous advertising or promotion)
I guess there is a higher chance of it not getting deleted if there are a bunch of edits/additions made by different people/accounts. So please, add to it :)
Thing is, some SQL databases have columnar storage, and in there selecting everything, then filtering with an attached function would eliminate the performance benefit of not selecting all of the fields.
This is why SELECT looks like it does. Not to mention it's much shorter than attaching a function for the purpose.
The author also himself acknowledges that:
> The downside is that your queries end up quite long and, for some, rather intimidating.
Ok so they're "quite long" and have less potential for optimizing the performance of. Amazing?
His example of creating specific indexes and views is also not new to SQL.
> There are 3 official drivers: Python, Ruby and Node.
Amazing?
> The query itself didn’t change at all – I could copy and paste it right in. I had to wrap it with connection info and a run() function, but that’s it.
So just like an SQL query, except I can connect to an SQL RDBMS from virtually any language I can think of, and not just a narrow selection of 3 script languages.
I sympathize with author's excitement, but from all his examples SQL feels like it has quite an edge both in availability and in terms of design and fit for the domain than a bunch of JS functions composed together (as much as I like composing functions together in JS).
I realize how much hard work the folks at RethinkDB have put into creating their product. But technology adoption is not driven by pity, it's driven by benefits. For a new type of DB to not be a flash in the pan it needs a lot more than being "stable and fast". It needs to offer significant additional benefits when compared to existing DBs. And I ain't seeing it.
RethinkDB has no access to the structure of the source in order to analyze it statically and work out an optimal I/O read plan. It interacts with the language runtime by providing an API and receiving callbacks to the API from the runtime.
SQL is parsed & analyzed statically at the server, a plan is created based on that analysis and executed. So with SQL it is possible to do so.
With RethinkDB you compose your query in the script, basically, and all of the optimization opportunities end with the exposed API (no function source analysis).
It's not impossible to redesign the API to provide or even mandate static details like requested fields to RethinkDB, and it has a bit of that, but it allows freely mixing in client-side logic and even OP is confused about what it means to have a client-side mapping function.
If they would allow complex expressions to run on the server, it'd become quite verbose to compose that via an API in an introspective way, to the point it'd warrant a DSL in a string... and we're back to SQL again.
Actually this isn't true. One of the really cool things about RethinkDB is that despite the fact that queries are specified in third party scripting languages they actually get compiled to an intermediate language that RethinkDB can understand.
That being said AFAIK RethinkDB doesn't optimize selects the way columnar databases do. I believe it can only read from disk at a per document granularity. But it does have the ability to optimize this in the future.
I would think that would allow Rethink to analyze the structure of the query and perform appropriate optimizations.
Here's the code in question:
.map(function(album){
return {artist : album("vendor")("name")}
})
If this is simply adding a node to an AST, it could be expressed without a function: .map({artist : ['album','vendor','name']})
Using a function for this would be quite superfluous.The restrictions on what language features you can use in lambdas inside queries exist because the query isn't executed on the client, the query in the client language is parsed into a client-language-independent query description which is shipped back to the server and executed on the server. So all the information about the query is available to the server (how much it actually uses for optimization, I don't know, but the query is not opaque to the server; what is composed in the scripting language has the same relation to what the server sees as when you use an SQL abstraction layer that builds SQL and sends it back to the server with an SQL DB.)
So just like an SQL query, except I can connect to an SQL RDBMS from virtually any language I can think of, and not just a narrow selection of 3 script languages.
http://rethinkdb.com/docs/install-drivers/This blog post explains how this works: http://rethinkdb.com/blog/lambda-functions/
Well, sure, because (1) major SQL databases have DB-specific drivers for many languages (often third-party), and (2) SQL uses a well-established, common model for which generic connectivity tools exist (ODBC, JDBC, etc.) so even minor SQL-based databases can go pretty far if they've got just ODBC and JDBC drivers.
But while RethinkDB may only have the three languages with official drivers, there are lots of third-party drivers, and there is documentation on the protocol and process for writing third-party drivers. Obviously, it kinds of loses out where ODBC/JDBC and similar technologies are concerned (though you probably could build drivers for Rethink using them, but you'd probably have to lose lots of Rethink's unique features -- particularly the push feed one -- when using them.)
> I realize how much hard work the folks at RethinkDB have put into creating their product. But technology adoption is not driven by pity, it's driven by benefits. For a new type of DB to not be a flash in the pan it needs a lot more than being "stable and fast". It needs to offer significant additional benefits when compared to existing DBs. And I ain't seeing it.
The key additional benefit compared to most better-established storage technologies seems to be ability to simply set up push feeds from queries. I'd say the demand (or lack thereof) from that is likely to be the determining factor in whether the resources get devoted (first- and third-party) to the RethinkDB ecosystem to bring the kind of conveniences that are seen with more established DBs.
I am very impressed by RethinkDB's cluster management, etc, so I would like to explore it as an option, but is there an easy way to sync my offline (browser based) localstorage-like database to rethink and back again? PouchDB makes this dead easy.
I have a similar use cases as you in mind for my PouchDB setup; being offline for a while and then syncing many things at once. So far I have been testing with high performance networks mostly (including good 3g however). I have also tested going into airplane mode on my phone, then making changes and then going online again after a while. All changes came through nicely. So you can do what you ask: sync as soon as there is a connection.
Indeed when there are no issues regarding conflicts that you have to resolve it seems that Pouch and Couch are a very very good way to have offline <-> online sync.
> It differs from the other driver (rethinkdb) in that it uses advanced Haskell magic to properly type the terms, queries and responses.
For example the driver knows that this query returns a number, and tries to parse it as such:
call2 (lift (+)) (lift 1) (lift 2)
Here are a few more examples from my application: -- | The primary key in all our documents is the default "id".
primaryKeyField :: Text
primaryKeyField = "id"
-- | Expression which represents the primary key field.
primaryKeyFieldE :: Exp Text
primaryKeyFieldE = lift primaryKeyField
-- | Expression which represents the value of a field inside of an Object.
objectFieldE :: (IsDatum a) => Text -> Exp Object -> Exp a
objectFieldE field obj = GetField (lift field) obj
-- | True if the object field matches the given value.
objectFieldEqE :: (ToDatum a) => Text -> a -> Exp Object -> Exp Bool
objectFieldEqE field value obj = Eq
(objectFieldE field obj :: Exp Datum)
(lift $ toDatum value)
-- | True if the object's primary key matches the given string.
primaryKeyEqE :: Text -> Exp Object -> Exp Bool
primaryKeyEqE = objectFieldEqE primaryKeyField
My driver doesn't include all commands of the query language, just those which I need in my product. And I haven't tested it with RethinkDB 2.0 yet.[1] http://hackage.haskell.org/package/structured-mongoDB-0.3 [2] https://github.com/pkamenarsky/typesafe-query
This is a broader discussion to be sure, and it's been had. Given that PG now supports jsonb, it does mean that yes, we get to have these discussions more.
https://groups.google.com/forum/#!searchin/rethinkdb/meteor/...
You can follow the progress on https://github.com/rethinkdb/rethinkdb/issues/3930
So..I suggest you figure out what your requirements are and then use the best tool for the job.
Aerospike is battle-tested in large deployments --- ad-tech, marketing-tech, a few new ones in telecom and fin-serv. Pushing huge load with very, very little downtime. That's what we're the most proud of --- and I'm proud that we're able to offer this killer codebase as open source, after being closed source for the first few years of the company.
Most applications have a huge core of key-value --- twitter, for example --- and need a fast and scalable key-value component. You can start with a single server (on your laptop with Vagrant) and scale up later.
We're adding more types, more cool operations, more indexes this year.
The fact that Aerospike has a basic query system, type safety, flash optimization (Amazon has switched over to being very SSD/Flash centric) support for every language under the sun (three contributed Scala layers --- and we see a lot of Go use as well as the usual Java / Node / Python / PHP / HHVM), Hadoop integration....
Hard durability means that every individual write will wait for the data to be written to disk before the next one is run (in this benchmark, since it only does one at a time).
I don't think any of the other databases in this test is using a similarly strict requirement, are they?
You'd have to run with the currently commented line "rethinkdb.db('talks').table_create('talks', durability='soft').run(conn)" to get more comparable results.
(Edit for clarification: `durability='soft'` is comparable to the `safe` flag in many of the MongoDB drivers. It means that the server will acknowledge each write when it has been applied, but not wait for disk writes to complete.)
db.catalog.find({ 'details.media_type_id': 2 }) r.table('catalog').filter({ details: { media_type: 2}})
Or like this: r.table('catalog').filter(r.row('details')('media_type').eq(2))
For most queries MongoDB syntax and RethinkDB syntax are effectively interchangeable.http://docs.mongodb.org/manual/reference/limits/#Restriction...
Haskell (or F# or any FPL) are, of course and obviously, a perfect fit as front end for a relational system.
http://rethinkdb.com/docs/comparison-tables/
I can't believe new databases are still using this model. Some of my data storage would be 90% keys and 10% data as JSON.
I'm not entirely sure what you're referring to, but the closest I can think of is the move from the memcached interface in 1.1 to the ReQL interface in 1.2. If that's what you mean, I don't think it's fair to say we abandoned our users at all.
The memcached interface had very few people using it (literally single digits). We tried really hard to make it work, but there just wasn't any demand, so we decided to add a full query language, clustering, and rebuild with the realtime architecture. We supported the binary for a while, and helped most of the users migrate from the memcached interface to ReQL (which was fairly easy).
We also helped people migrate to other memcached alternatives if they chose to not to use ReQL. In almost all cases people could quite literally pick another compatible product without changing any of their code.
We took our time, helped people migrate (either to the newer version of RethinkDB, or to other products), and integrated the original architecture and as much of the code as we could into the new and improved RethinkDB. All of this was completely free of charge.
So respectfully, I really don't think you're being fair to us. I'm sorry if this inconvenienced your company, but given the dire circumstances at the time we really did the best we could (and arguably, much, much more than most companies do in those circumstances).
Call it declarative, format it however you want - a purpose-built DSL like SQL is always going to be easier to grok than a Javascript-inspired functional language (for me at least).
However, I am imagining a contest where SQL people write SQL and functional people write functional queries... I would bet money that SQL people could identify basic facts about SQL queries faster than functional people could identify the same facts about functional queries.
Could be a fun programming game.
Because SQL is very familiar and ReQL is not.
r.table('users').filter(r.row('age').gt(30))We did the best we could with JavaScript. IMO the Ruby implementation of ReQL is dramatically more beautiful. Ruby's blocks fit so well that we don't even provide the `r.row` shortcut in ruby:
r.table('users').filter{|row| row['age'] > 30} r.table('users').filter(row => row('age').gt(30))If you python you use python's native constucts, which if you use python you're naturally familiar with.
r.table('users').filter(r.row['age'] > 30)I don't actually agree with this, but one solution to this would be -- as Google has done with several of its not-traditional-db storage products -- to build a library that takes strings in an SQL-like language (SQL with additions/deletions to align with RethinkDB's features and capabilities) and builds ReQL objects from them. But, while it might have utility, its probably not as high a priority for the RethinkDB team as getting the core features right.
My be an interesting third-party add-on, though.
See, for instance, http://rethinkdb.com/api/python/
(You can easily translate most doc pages to the language of your choice using the links at the top of each page.