RethinkDB raises an $8M Series A
rethinkdb.com
rethinkdb.com
RethinkDB team is the nicest possible group of people one can hope to work with in Bay Area. They have a great combination of often mutually exclusive things: hacker-friendly business model (no ads, tech for cash), aggressiveness and tech-savviness of founders, yet they're humble, honest and nice.
And the product works, well-liked, serving an exploding market, so the probability of failure is quite low by a typical start-up standards.
This is very rare. Make a move.
* Reliability engineer
Skills:
- Good working knowledge of Bash and Python
- Good working knoweldge of C and/or C++
- Patience
Responsibilities:
- Automate testing and benchmarking infrastructure
- Make sure long tests/benchmarks reliably run
- Get to the bottom of stability/performance problems, and
work with the engineering team to fix them
More info: http://rethinkdb.com/jobs/reliability/
* Technical writer/developer advocate
Skills:
- Basic programming ability in at least two of ruby/python/javascript
- Good command of english language, style, tone, etc.
- Ability to explain complex ideas in simple ways in written form
Responsibilities:
- Improve API documentation, guides, and tutorials
- Help support users and bring their concerns back into the product
* Ruby/Rails, Python/Django, Javascript/Node experts
Skills:
- Deep understanding of one of the above stacks (you should
be able to hack on core Rails/Django/Node, know the
conventions, and respect the community)
Responsibilities:
- Write software to make RethinkDB the absolute best possible
experience for your respective stack
- Then tell the community about it
* Visual designer
Skills:
- Produce beautiful web designs and illustrations
- Mastery of Adobe Creative Suite
- Mastery with a tablet and stylus
- Exhibit creative taste
Responsibilities (we take visual design and user experience extremely seriously):
- Design user interfaces
- Help rebrand our website for commercial / community versions
- Incorporate illustration and a unique visual style into our brand
- Creative one-off projects (logo, design t-shirts, marketing
materials, postcards, packaging, generally make things
beautiful)
Email jobs@rethinkdb.com -- we'd love to hear from you.As far as getting started on the codebase, we unfortunately don't have a guide/mentorship program yet for hacking on the internals. The best way is probably to pick a really simple bug and try to fix it. I'll see if we can work on making getting started with core contributions easier.
We do, however, help with relocation (and even take care of the logistics for people from oversees).
I am no expert on databases, so I don't have much to say on that front, but the admin UI for RethinkDB is definitely beyond anything I've seen for databases before. I know that's also something they value tremendously and I'm very happy to see them doing well.
Keep up the great work!
RethinkDB supports join out of the box.
I mean, you realize if data is distributed at large scale, it may take a while till it gets from all nodes the data and joins it...
Now in the cases where if you can optimize the joins, you still have the option of doing it in your code in RethinkDB/CouchDB. I've done that too, and it's usually when I know for sure that I can prune a big collection to a very small subset more efficiently than using an index.
I would still argue that client app is not the right level of abstraction for data join though, unless it is a big performance gain for very little extra complication.
It seems like the addition of joins lends some relational capabilities which in my eyes is very impressive. I'll be watching the advance of this one with interest.
Mongo and Rethink are both single-master/multi-slave solutions.
One is not better than the other, just depends what you need.
And as someone else mentioned, if you want complexity out of your CouchDB queries, you must write map-reduce functions to provide those "views" for you to query. Rethink you can treat similar to a SQL data store and just execute queries against it.
RethinkDB supports MVCC out of the box -- you don't have to do anything as a user to get it. (It basically means you can do writes and long reads concurrently without any locking issues).
Sharding and replication in Rethink is similar to Mongo's architecturally, but is much easier in practice. You can shard and replica in Rethink in one click -- check out the 1m highlights video (http://rethinkdb.com/videos/what-is-rethinkdb/) to see how it works. If you have a little more time, take a look at the 10m screencast for more details -- http://rethinkdb.com/screencast/.
and i think those are very important in high-performance sharding (think-per user_id sharding etc)
A document store flips the default. It makes dealing with data that has lots of nullable columns much, much easier. (It also makes dealing with hierarchical data a breeze)
There are lots of details, but this is the gist of it.
Certainly a lot of applications would benefit from having a full RDBMS they can opt-in to document-style data when they feel like it?
Built-in horizontal scaling is one selling point for non-RDBMS stores, but large systems seem to just shard on top of RDBMSes anyways, right?
It changes the default, which results in a drastically different programming experience. The difference is difficult to describe in the same way a dynamically typed programming language is difficult to describe to someone who's never tried one.
I'd encourage you to try a document store (Mongo, Rethink, whatever) for a throw-away project. A ten minute tutorial walkthrough is worth a thousand HN comments when it comes to stuff like this :)
Here's the Rethink tutorial: http://rethinkdb.com/docs/guide/python/. Just play with it and see if you like it!
Related: The comparison to a dynamically typed language makes me suspicious. I spent a bit of time trying to find any examples of dynamic code that actually provided any benefit. Even read "Metaprogramming Ruby" and was dismayed to see examples of reading a CSV - big deal if I save a few quotation marks. The others were just places where the static type system wasn't good enough (duck typing), or dynamic code was a pain to get going (poor reflection/codegen APIs).
Both document databases and dynamic typing are at their best when you don't understand your problem domain. They let you express what you do know about your problem domain concisely, and then fill in the blanks later on. So in a document database, when you find that you want to record a new bit of data - just add it as a field to newly-created XML/JSON documents, and only display it in the UI if it's present. Or pick a default value if you need to perform computations with it. Don't bother with data migrations, don't bother with schemas, don't bother trying to backfill previous data. Try out your idea and see if it works first, because chances are, it doesn't.
If you always work on projects where the requirements are handed to you, specs are complete, and the problem domain is understood, this will seem terribly irresponsible to you. And it is - if you understand your problem domain, you should capture as much of that knowledge in the software system you build to understand it.
But if you are working in startups, or in consumer web, where you absolutely have to be on the leading edge or die and the only opportunities that haven't been picked over yet are the ones that nobody understands - being able to try things out without having to flesh out all your assumptions is crucial. You will run circles around the people who spend time defining their data model and speccing out their objects. And then when consumer tastes change - which happens quite regularly - you can adapt to them immediately instead of throwing out all the work you did under the old assumptions.
The other bit of context I'll toss in is to get in the mindset of solving a problem that you don't know how to solve and assume that your first 10 solutions won't work. For example, if you're reading a CSV - everybody knows how to do that, dynamic typing doesn't really help there. If you're cloning Stack Overflow, you can probably figure out what your database schema should be. But what if you're trying to figure out a new way for people to socialize over mobile phones? Where do you start there? That's the use case for dynamic languages and document DBs. The problems where technology is a tool for understanding & manipulating vaguely-defined social behaviors.
In F#, adding a field requires "field : type". If I don't care about type checking, I can just add "Props : dict<string, object>" and go to town. Or I can opt-in to the dynamic features and just do "foo?bar <- baz". When I change types around, things either just work due to type inference, or the compiler helpfully points out every place that'd be a runtime error. I've never felt this slows me down. I feel the type checking and autocomplete is worth the tiny amount that specifying a record takes. (I've spent days finding minor issues in JavaScript, stuff that'd instantly be caught by a type checker.)
Databases make it a more cumbersome, and it takes more than one line to start using a new field. I totally sympathize with the flexibility issue there. Even with a document type, most syntax I've seen doesn't have truly first-class querying support (not as easy a column, anyways). And it feels ugly to have some fields defined in schema, and some in a document. But that seems like a minor tooling issue -- there's no fundamental reason SQL can't let me do "WHERE x.SomeDoc.SomeField.OtherField > 5" (perhaps some minor scope resolution issues to ensure I'm not referring to some other multi-part name).
When the schema is not rigid and likely to change on a daily basis I prefer a document database over an SQL one. There are probably other use cases as well, this is just my favourite.
If you're going to concatenate relational data into a document, I'm not sure why a simple KV table doesn't fit the job.
Also, does Rethinkdb have any kind of transactional capabilities?
You have full ACID on a single document, but not across multiple documents. In this way Rethink is similar to other NoSQL systems (except you can do almost any operation imaginable atomically on a single document in RethinkDB).
Enjoy!
A slight aside, but I spotted this (currently broken) integration of RethinkDB and Meteor the other day and wanted to share. It does away with the long poll Meteor is doing on Mongo. (I have no involvement in this project at all.)
As for the polling on Mongo, you'll love what Meteor is shipping this week. Meteor now by default connects to Mongo as a replication slave and slurps up the replication log to drive your realtime queries.
(Not that I'm not a fan of RethinkDB. I've been playing with it since one of the earliest released builds and find it a really lightweight nice database to use.)
I do not know the specifics of this particular deal, nor have I used the product, but I hope the points are of going to be some use.
Let us look at it this way. Assume there are 200 funds out there who can do a series A of this size. That would naturally mean that not every fund is going to be either a leader or someone who spots new trends (there are not enough trends out there). Naturally, a lot of them have to invest in deals in other companies in a hot sector.
A lot of investment is momentum-driven and momentum is often driven by the narrative. You have to remember that as long as a successful exit happens, the fund winds up with a good deal irrespective of whether the public (IPO) or the acquiring company (M&A) eventually profits from it. NoSQL has that momentum at the moment.
A healthy start-up ecosystem can easily support more than a handful of companies in a single domain. Once the narrative for the domain really picks up, even the not-so-great ones (again, I have no clue about RethinkDB) stand a good chance of being acquired as long as there is decent enough traction and the sector is so hot that there is pressure on the GPs to make a play in it.
The later they get into the game, the pricier the ticket becomes, but you get lesser risk too.
And all of this is perfectly OK and fair.
This is a really good breakdown. I can't read our investors's minds, but I'm pretty sure this would be a worst case scenario for them. It's certainly not why we're doing Rethink -- if we thought it would be a #5 company in the space, we'd pack up and do something else (life's too short).
The NoSQL market is reminiscent of "horseless carriages" -- as long as you define a technology by an absence of something, you know you're early in the game. Databases are a fundamental part of the technology stack, and they tend to easily stick around for 20-30 years. We think we can build a long-term open source company that will stick around for that long (incidentally, that's why we take conventions in ReQL so seriously -- we imagine millions of programmers fifteen years from now cursing at us for a stupid naming convention).
It's not hard to imagine groundbreaking features in NoSQL products that nobody is shipping. That's why RethinkDB exists, and we think we won't be a niche product for long.
First up, congrats. Whatever the back story is, getting funded (save runaway revenue/profits/margins) is always a crucial inflection point in a company/product's lifecycle as an enabler for bigger and better things. Whether those will eventually happen or not, nobody knows. But there are a lot of things that only money can accomplish and investment is a key enabler for a start-up that needs capital to scale/grow. If you find an investor who is aligned extremely well, it is a massive bonus.
Historically, a lot of good has also happened from a combination of events that may not exactly be awesome. Outcomes always trump everything else. So don't sweat the mind-reading angle much!
Can't comment much on the technical aspects of the product as I am not even remotely qualified to do something like that.
I'm excited for your team and I have immense respect for anyone who builds an OSS company. There is much that the world owes to numerous companies and individuals releasing code like this and don't get enough credit for it. So, thank you and hopefully it will come together very well for everyone involved :)
-Dave (FoundationDB)
Not really, because "NoSQL" databases aren't really a coherent group of things. "Document databases" is a good name for an important subset, though. As are "Column-oriented datastores".
An investment happens when the investor thinks the company will be worth more to another investor in the future.
That's basically it (and it's basically what you said). The second investor in the sentence above could be another VC, a buyer (M&A), or the stock market (IPO).
There is actually real shortage of innovative databases. For example Redis was released just a few years ago, but all ingredients were here for decades.
It takes years to develop, quality demands are very high and takes long time to build reputation to get enterprise users. Also experienced people are scarce and get 'hired away' by big guys. Very hard field for start-up.
http://www.marketresearchmedia.com/?p=568
http://wikibon.org/wiki/v/Big_Data_Database_Revenue_and_Mark...
From the little I have seen of NoSQL, I like it as a niche use-case DB, not entirely divorced from a RDMBS. Thus, not surprised it has the potential to grow like crazy.
Also, I just started a blog series on Rethink. http://www.realpython.com/blog/python/rethink-flask-a-simple...
Next up will be performance testing.
Congrats, RethinkDB.
I've barely scratched the surface of Rethink in terms of functionality and features; but I definitely see a market for it as a competent and approachable database for people who want something that 'just works'.
The only pain points are that the API documentation is somewhat confusing. But they still haven't really launched so I expect that the documentation will be updated with better examples soon.
Also nice tutorial! I
Show inputs and outputs!
Love the theme you have going. Cheers!
Btw, how do you know that "Thousands of developers are already building applications backed by RethinkDB;"?
PS: I hope we'll see soon more official drivers...
I just trying to compare with my current stack spring mvc 3, spring data mongodb.
We know in two ways. First, people have told us via GitHub/IRC/e-mail/etc. and Shirts For Stories page (http://www.rethinkdb.com/community/shirts-for-stories/).
Second, the administration UI checks for version updates, which gives us some information about how many developers are using RethinkDB and how long they stuck around.
There is no perfect way to know because there are many sources of imprecise information, but we're quite confident about the overall conclusions of the data.
> I hope we'll see soon more official drivers...
We've been trying to keep the surface area of the project low for now. Which specific drivers are you interested in? (we probably won't be able to do anything about it for the next ~6 months or so, but having the info helps enormously)
Java! My current stack is spring mvc 3, spring-data-mongodb. It's not perfect, but at least I don't have to hack my way into basic db ops...
I have been using RethinkDB over the last month in a new project. If you know that a document store is the right solution for you, take a look at RethinkDB. I evaluated it against some of its competitors, and I must say that I was really amazed at the deep engineering thinking that is going into RethinkDB. The ease and power of its programming model (use of AST/lamda functions and like abstractions are awesome), and attention to ease of deployment and manageability (great UI!) is unparalled in like products. RethinkDb is a young product for sure, but one with a very bright potential. In addition, being well funded should help alleviate fears and hopefully help it further gain traction.
Best of luck!
Super helpful on IRC (I've been in there multiple times for help with small problems), seems like an overall awesome team
I've seen issue 1648 closed for 1.12 and issue 97 to be completed. Are these two enough to fix the limits?
Will it be possibly to have, say 100K tables in a single database? In 1000 databases? Is it possible to have 10K databases?
I've read @cofeemug's explanation that a table is a heavyweight object requiring a few megs of disk space. But that's just a few TBs for 100K tables which is perfectly fine for a 32 node cluster.
Also, it would be great of you could make a page like Mongos' "Limits and Thresholds". I understand that you have lots of other things to do but that one is key in making a decision to use Rethink vs other options.
A very crude measurement - I just threw it on a box that I'm 70ms away from, I'm getting insert responses back in 90ms, on Mongo which I HATED (the "old" query language I was trying to understand.. a while back it was thought for just a moment we could get by without the relational algebra) - I was about 200ms - so far so good, but how much can I hammer a particular node?
Looks to be using ~20MB ram on the svr process, works for me..
If you run into any performance issues, please let us know, we'd love to fix them!
You can't quite divide like that. In Rethink
r.table('users').insert({ 'name': 'Bob' })
r.table('users').insert({ 'name': 'Jim' })
is much slower than the equivalent batched insert r.table('users').insert([{ 'name': 'Bob' },
{ 'name': 'Jim' }])
The latter isn't subject to network round trips and can batch disk writes, drastically increasing performance relative to multiple individual writes. You can see a similar effect in most other database systems.There other nuances to this -- the complexity of measuring things properly is why we haven't published benchmarks to date. Check out this doc on insert performance for more details: http://rethinkdb.com/docs/troubleshooting/#my-insert-queries...
I am very aware of the increase in performance of batch inserts and batched disk writes and so on.. as an individual who has worked on the development of a major DBMS.
In effect if i wanted to do 10000 single writes per second typical of most interactive systems, i will need 200 nodes to pull that Off.
In Rethink you can turn on soft durability which will allow buffering writes in memory. That would drastically speed up write performance. Another option is to use concurrent clients. Rethink is optimized for throughput and concurrency, so if you use multiple concurrent clients the latency won't scale linearly.
RethinkDB is much easier to set up and operate, has a much more powerful query language, and is much friendlier for application developers.
Cassandra supports high write availability in case of network partitions. RethinkDB does not. The flipside of that is that you (as an app developer) have to deal with conflicts in Cassandra which makes writing applications much more difficult.
If you need high write availability in case of network partitioning, go with Cassandra or Riak. If you don't, go with RethinkDB or MongoDB.
IMHO the main difference is data model. RethinkDB is a document store, Cassandra is a wide-table store. As of the query language I agree, but this is caused by Cassandra never including anything that would not scale out.
Note that Rethink is still in beta. Lots of companies already use it in production, but we advise people to test carefully until we ship a long term support (LTS) release. We'll also offer commercial support options then.
joe@alchemist~$ rethinkdb
joe@clockwerk~$ rethinkdb -j alchemist:29015
Just guessing, but joe is probably a Dota fan :)http://venturebeat.com/2013/12/16/rethinkdb-grabs-8m-to-show...