In most node.js apps, the best answer is probably a SQL database. Sorry.
If you're working with timelines or other cases where Redis's data models can help you, consider it, though beware that if your data is large, things will get more expensive fast since you're keeping everything in RAM.
HBase, Cassandra, and Riak are all reasonable in similar cases and have their own tradeoffs.
And yes, Couch fits a similar niche as Mongo. You might even be able to use something simpler like BerkeleyDB (quite mature) if you think you want a document store.
RethinkDB may be a nice choice too. It's fairly young but looks like it's going good places.
But your choice should be mostly dependent on what kind of data you're storing and what kind of guarantees and access models you need.
It should not be based on someone on HN telling you "Riak is the best NoSQL database for Node.js" because their idea of what most Node apps need may not be what yours needs.
Another highly underrated solution is using MySQL/PostGres as a key-value store. Just create one table for each entity type, with the primary key as the key and a JSON or protobuf blob as the value. You're using completely battle-tested solutions, you've got bindings in basically every language, you're doing basically the same work (at the same speed) as your NoSQL solutions, but you have a lot more flexibility to add additional indices and can rely more on pre-existing functionality than a MongoDB or CouchDB solution.
That works for some things. However, it's no more a foolproof magical solution than MySQL or MongoDB or Cassandra or Oracle or... It just has different tradeoffs (non-primary key queries will tend to be a problem, you'll have to make your own replication, sharding will be a problem, etc etc).
The nice thing about doing the dead simple solutions first is that they give you time to focus on the things all startups have to do (getting users, building product) and then fall down at the the things that very few startups have the luxury of needing to deal with (scaling, fault tolerance, reporting, alternative views of data).
Throughout the lifetime of my first startup, I was obsessed with the question of "What are we going to do when we need to scale?" It failed because it had a daily userbase measured in the dozens. Then I went to Google to learn how to scale things. And it turned out the biggest lesson I learned at Google was not how to scale things (though I did learn that too), but that you shouldn't scale things, not until you need to. Because the process of designing for scale slows you down significantly, and makes it much harder to develop a system that's usable and performs well under small workloads. Google products take forever to launch, because they have to scale to millions of users from day 1. As a result, their product decisions are very often questionable in early versions. Most startups don't have the luxury of Google's brand name and billions in cash to tide them over that learning process, and need to hit the ground running.
Focus on the problems you have, not the problems you hope to have in the future.
https://blogs.oracle.com/MySQL/entry/nosql_memcached_api_for...
Wait, what? Even if vertical scaling was a good idea, scaling is far from the only reason you should have more than one server for anything serious.
If you do get to the point where you need some redundancy (and don't yet need to scale horizontally), you can proxy all writes to a second server running the same codebase, have it update its in-memory data structures in the background, and hot-swap it over if the master dies.
That's probably the biggest surprise I learned from working in a fast-growing, well-functioning engineering organization. The half-life of code in a market that's actively growing and changing is roughly 1 year, i.e. 50% of the code you write now will have been removed within a year from now. And attempts to optimize for problems you're going to have in a year, rather than the ones you have now, actively make things worse because you inevitably have a different product direction in a year, and baking in last year's speculative assumptions just means there's more code you have to work around.
You also understand that most of the advice easily accessible on the Internet comes from people trying to sell you something, and so they have a vested interest in you adding many layers into your software stack that you don't need?
If you work in an actual engineering organization that has a clue what they're doing, mmap() is your best friend, and the more layers you can cut out of the stack, the better off you are.
https://news.ycombinator.com/x?fnid=cjVXpi8HxVR5TTze3bqSCa
Unknown or expired link.
Oh I remember now...My guess: because by relying on in-memory data-structures you can't do what any half assed php forum do, ad hoc queries.
Anything you can do with SQL you can do with in-memory data structures. If you're interested, I'll be happy to take any SQL query and convert it to some Python list comprehensions on arrays of dicts.
BTW, do you miss Java's more advanced structures (say MultiSet) when programming in Python/Go?
I think I'd miss these a bit more in Go because the built-in datatypes are privileges in some of the language statements, but I haven't written enough Go code to really feel their absence.
Sure, but unless you also do some indexing manually, you can't really query your whole dataset when it start to become too big.
CREATE TABLE mongodb (
key VARCHAR(256) PRIMARY KEY,
value JSON
);Also, see https://github.com/umitanuki/mongres
Postgresql speaking the mongodb protocol.
This is currently a prototype. The following operations are supported.
db.collection.find()
db.collection.insert()
Two methods only, well, that's too little.https://postgres.heroku.com/blog/past/2013/6/5/javascript_in...
I don't know how stable/performant it is (I've never needed to use it), however...
Well, this is the thing; 'NoSQL' is really a pretty unhelpful term. It tends to just mean "not relational", and covers a vast number of things.
So, for instance, you might be okay with having to have your data set fit in RAM (with MongoDB you'll suffer if it doesn't, anyway), and not care too much about availability. In that case, Redis might be good. Or maybe you care deeply about availability; in that case, one of the Dynamo paper databases might be good, if you're willing to put in the work dealing with the consistency issues. Or...
I could go on for a bit. 'NoSQL' is verging on a meaningless term.
On the other hand, Redis, Cassandra, Riak, and many more are also excellent NoSQL databases. But none of them, including CouchDB, are excellent at everything. What are you planning on making? You can write a lot of different things in node.js. If you're writing, say, a blogging engine you probably should look into flat files, or maybe Postgres, and forget the NoSQL kool-aid. :)
I really don't see how MongoDB beats Postgres for running a basic blog. And while it doesn't prove anything, I note that Ghost (which has been getting a lot of press as a new, shiny, node.js based blogging platform) is backed by SQLite of all things. Why is it obvious that they should have used a document database instead? What advantages do you think that would have given them? Because of the top of my head I can't think of one.
1. If you are doing multi lingual site, you can store your multiple language content in a single document instead of futzing around with {lang, content} tables 2. If you want to do custom form/content, it is trivial to do it in a document database instead of relying on key,attribute tables. 3. Just store your theme in a single document, which can include various html templates, css, etc. To export or import a theme is also easy - just stuff the whole document into the db. 4. If you want to add plugins to enhance the capability of your blog/cms, they can have their own nested document inside their target document. Everything is contained.