RethinkDB 2.1.5 performance and scalability
rethinkdb.com
rethinkdb.com
If I could go back in time, I would use RethinkDB for our actual product.
Do you have a link to explain what you mean?
To put it another way, what do the databases you listed have in common with each other that they don't have in common with other NoSQL databases, besides marketing?
In this situation, NoSQL is the marketing term, while a graph database is a clear concept with corresponding performance and capability expectations.
I've found that it's easier and faster to import and traverse edges in SQLite or PostgreSQL than in some of the systems you listed. I haven't tried RethinkDB but it sounds promising. You use words like "hopefully" and "performance and capability expectations", but distinguishing databases based on their hopes and expectations is not at all a clear concept.
The main difference seems to be whether you call things "nodes and edges" or "documents and references" or "rows and joins" in the documentation.
(edit: removed some negativity, and I'm aware some remains)
Similarly, the query language is optimised for graph traversal types of queries, in a way that would not be possible in an RDBMS due to the relational model constraints, and also because some of the query operations would be extremely inefficient in an RDBMS storage engine.
> That sounds like graphs but Rethink is not a graph database (compared to Neo4j, GUN, Orient, Arango, etc).
With regards to the listed databases, some are, but the rest are not (and do not claim to be) graph databases.
Note: It's different people replying to you.
(But I'm a bit afraid of using it. Any armchair lawyers want to tell me if code that uses a GPL database has to be GPLed itself?)
Btw, Neo4J is definitely a graph database, although it might not be what you need.
As long as you access blazegraph via the sparql or tinkerpop interfaces you will clearly be fine.
Yes, "documents and references" and "nodes and edges" is where things become fairly synonymous which is why I asked about RethinkDB supporting relational documents. I honestly think they are doing themselves a disservice by calling it a "JOIN" since that sounds slow and scary. If they wrapped it in a nice API like ".navigate()" or ".traverse()" I think it would be reasonable for them to advertise as a graph database! Tip to anybody at RethinkDB in case they are listening.
* how you use rethinkdb
* what you don't use it for, and why
* good examples resources like ORMs, repos, and data models.
Some things I like/use/reference:
- https://github.com/drhurdle/node-rethinkdb-auth-starter
- Yo Express, provides a nice structured MVC application with thinky and rethinkdb with gulp
- https://www.airpair.com/javascript/posts/using-rethinkdb-wit...
Eventually, we discovered ReQL by itself is just great to work with (especially with rethinkdbdash). Writing your own queries also has significant atomicity benefits over Thinky.
Thinky and raw ReQL definitely work together, don't get me wrong! But as our project scaled we found 95% of our code had become raw ReQL for atomicity, performance, and relational reasons, so it just didn't make sense to use Thinky in future projects. Our original RethinkDB project still contains Thinky, but it'd be cleaner if we just made a complete switch at this point too.
To each their own though, of course!
I'm not using any real-time features yet which makes me a bit sad, but have been using it before real-time was their focus.
I'm the maintainer of the yeoman express generator, glad you like it!
Overall, it's pretty nifty, hoping to find a chance to actually use it. Early on the lack of geo-indexing and automagic failover were features I'd find important for a lot of setups (now has these), but lack of atomic commits across multiple tables/collections not so much.
From https://rethinkdb.com/docs/data-modeling/:
Disadvantages of using embedded arrays: Deleting, adding or updating a post requires loading the entire posts array, modifying it, and writing the entire document back to disk. Because of the previous limitation, it’s best to keep the size of the posts array to no more than a few hundred documents.
RethinkDB looks really nice, was really easy to play around with it when I last did. I'm considering using it for a small project, but wondering what everyone thinks of the downsides, compared to the upsides. (there are pros and cons to everything :)).
Minor failover is automatic, but major failover requires manual intervention to repair tables even when there's no data loss. That's a little confusing and not obvious.
The anonymous function query style can get confusing because it is partially executed client side and then executed server side. This can be a subtle source of bugs.
Source: I am the creator and maintainer of the Elixir driver for RethinkDB.
Wasn't this fixed in RethinkDB 2.1? Or are you referring to something different?
Edit: it's better documented now, but still not obvious. https://rethinkdb.com/api/javascript/reconfigure/
I've never encountered it organically, but did during some benchmarks when I was dropping and adding a bunch of servers.
2. When I needed "more exotic" data types (e.g. but not limited to Decimals) where you want query/analytics logic to happen on the database rather than the app.
3. OLAP types of use cases. I really wanted collation, sort order, query optimizers, ability to "explore" the data in relational ways, and so on.
I suppose using computed property indexes in RethinkDB might work as well, or generating indexes, sort order, and doing query optimisation outside of the DB, but I couldn't figure it out in time, and it also seemed way more complicated than replicating the data to PostgreSQL.
EDIT: Above is for the OLAP use cases. OLAP of some types just don't go well with document databases. For the Decimal type, I stored it as a string, and had a computed property that transformed the string to a float. The float can then be indexed and queried like a number. It was good, until all the other OLAP use-cases creeped in.
Having said that, I found that custom replicas with RethinkDB is easy because it emits data change events. Never expected that feature to be an escape hatch.