Choosing NoSQL For The Right Reason
blog.fatalmind.com
blog.fatalmind.com
The first, aimed at every developer/application out there, is that some NoSQL solutions (document stores being the most obvious to me) have less friction with OO languages. In other words, they are simply more productive to work with. Embedded types, arrays as first class objects, and schemaless design all help you write simpler infrastructure code that takes less time to write, test and maintain.
The second reason is that some NoSQL solutions provide specialized features/capabilities. For example, what you can do with an [Ordered] Set in Redis simple isn't possible with relational databases of moderate (say 100K rows) size. You either have to lower your requirements (not real time), or increase your hardware. There are various examples...Solr does better indexing, Hadoop does better data processing....MongoDB does good logging and geospatial stuff.
The right way to look at it, imo, is to see RDBMS as specialized systems. Most people should almost always opt for a more productive general purpose solution upfront (say MongoDB), and then only turn on a specialized supplementary system (say Hadoop, or PostgreSQL) when they have those specialized needs.
Choose SQL for the right reasons.
RDBMS have a lot more features and support many more usage scenarios out of the box than any NoSQL systems and SQL offers more features as a query language than any other NoSQL query languages.
Of course, NoSQL query languages are simpler than SQL by design because the programmer is expected to write a program or a script for complex queries instead of using the query language. But then, you have to write your own execution plan, and think about atomicity and consistency, etc.
NoSQL are clearly the specialized systems, not the general ones.
That said... some NoSQL may work better with OO, but so do some RDBMS schemas. I've had experiences where it's simple and almost 1:1 between object and table, and other times when I'm bending so far that I wonder why we're taking an OO approach at all.
I don't think that I'd disagree with your claim that an RDBMS is a specialized system, but I'm not sure that various noSQL solutions are any less specialized.
I agree, choose SQL for the right reasons, but you could just replace that with "choose your persistence approach for the right reasons".
It's outdated though (MongoDB stable is 1.8.1 already).
One gripe I have is that NoSQL is becoming the default by most PaaS providers. I have a Duostack account and they offer 3 database choices: Redis, Mongo, and MySQL. In their documentation they recommend using persistencejs which doesn't work with the new versions of node.
Honestly, if I would have to develop a revision control
system, I wouldn’t take an SQL database as back-end.
But then, suddenly, Fossil, which seems to be doing pretty okay. (Admittedly, it does a lot more than manage blobs.) http://fossil-scm.org/index.html/doc/trunk/www/index.wikiNoSQL forces discipline which is a good thing, preventing "late" performance optimization when in production.
If you have a table with four rows in it, would you like the database to give you errors when you do a full scan on it?
What you want is an option to throw errors when a query takes too long to execute and most products already have this in the form of a statement timeout parameter.
On a separate note, the balance between "late" optimization and "premature" optimization is thin indeed... These test systems need to run with the same data and on the same hardware as your main database to give meaningful results, so it's generally unavoidable to do some kind of performance tweaking on the already deployed code. And of course as the data distribution and its amount changes, you need to keep on optimizing...
In MongoDB (and maybe others too) there exists an option to fail when doing a table scan so that you know immediately while developing that there exists a potential for things to go awry at some later point. You can fix that now by either re-thinking your data structure, re-thinking your query, or adding appropriate indexes.
MySQL's slow query log is good for reactive development instead of proactive development.
I think its an indispensable feature.
I still don't see what a setting that gives errors when a particular access method has been chosen would be good for.
You don't care about the plan, you care about how fast the query runs.
You'll always have to do reactive development when trying to get good performance. There is only so much simulating and testing you can do and real data volume and distribution changes all the time.
Many developers have been burned by trying to optimize their queries and indexes too early. The query planner, although sometimes unpredictable, is generally better than developers at predicting the performance of a particular execution plan. Table scans are also preferable in certain situations.
Relying on the fact that RDBMS have sophisticated query planner (as opposed to MongoDB) is not a sign of "reactive development". It just means that you should test your queries (and your indexes) with close-to-production data, and, as the data grows and evolves, continue to test, because the best execution plan might change.
At least for interactive apps, NoSQL wins.