NoSQL is a Premature Optimization
smoothspan.wordpress.com
smoothspan.wordpress.com
However, the truth is that most business datasets are relational, and because of that relational databases do an excellent job of storing and accessing that data in an efficient way. I completely agree that people who try to force what would be a simple relational database into a complicated NoSQL system to gain scaling in the future are making a huge mistake.
Is it really for free? While I tend to agree with your point, one of the wrinkles here is that the older SQL systems have been beaten to death in production. We've had all sorts of 'little problems' with some of these newer systems. And this is not just about runtime issues, the 'hardened product' benefits reveal themselves in more secondary things like education, troubleshooting, monitoring, add-on products, and backup.
But they haven't. SQL Multi-master replication systems with no single points of failure really aren't that common, and that's what everyone really wants. Everything else imposes huge production costs.
It's true that anytime you write a function that maps some set to the set of logical values (true, false, null, maybe, ...) you're dealing with relational data. So while there is no clear cut border, at some point it becomes obvious that you would benefit from a Relational Data Management System. The danger is that your particular data set might out grow a particular management system.
When you switch to a "NoSql" solution all you're doing is deciding that you need a more purpose built managment system for your relational data.
I don't know about that either. Graph DBs have been around a while and are ideal for relational data (you don't have to mess with tables and joins because everything is explicitly joined), but Web developers are just starting to realize their power.
If I am developing a website that uses the Facebook Graph API (http://developers.facebook.com/docs/reference/api/) or the Twitter API to store my users' friends and followers (and many modern auth systems are doing this), a graph database is ideal because you can use a graph query language like Gremlin (https://github.com/tinkerpop/gremlin/wiki) to do all sorts of cool stuff, like find a user's friends of friends:
g.v(user_id).outE('friend').inV.outE('friend').inV
This says for the graph database "g":1. Start at the vertex for the user_id
2. Follow the outgoing edges labeled "friend" to the incoming vertices (these incoming vertices are the user's friends)
3. From the user's friends, follow the outgoing edges to the incoming vertices one more time to get their friends.
This is so much simpler than messing with joins in a relational database, and it's much more performant on large data sets.
I disagree. I have no idea what you just said and I'm sure I could figure out the join faster than trying to figure it out.
This script will return a list of all the user's friends:
g.v(user_id).outE('friend').inV
All you're doing is starting with the user node v(user_id) and following the outgoing links that are labeled "friend" to the adjacent nodes (the "incoming vertices").In a graph, each "edge" (each link) has an outgoing vertex and an incoming vertex. The outgoing vertex is where the edge starts from, and the incoming vertex is what the edge points to:
outgoing_vertex -- edge("friend") --> incoming_vertex
Here's a more concrete example... james -- "friend" --> julie
In this example, "james" is the outgoing vertex, "julie" is the incoming vertex, and they are linked together by an edge labeled "friend".Again "vertex" means the same thing as node so every user in the graph is a node (a vertex).
it's really quite easy once you become familiar with it
This was really my point. Things are always simpler when you're familiar with it and I happen to be very familiar with SQL. I just don't think ease should be a highlighted feature of NoSQL.
Don't get me wrong though. I use a few NoSQL systems in my company. But I didn't pick them because they were easy to understand.
Why do you say that? -- see "MySQL vs. Neo4j on a Large-Scale Graph Traversal" (http://markorodriguez.com/2011/02/18/mysql-vs-neo4j-on-a-lar...).
You can represent a graph in almost any data structure, including a relational database. But the difference between a graph database and everything else is that in a real graph database (like Neo4j), each node has an internal/local index for its adjacent nodes so it doesn't have do an external look up for each traversal step.
Watch this video on "The Graph Traversal Programming Pattern" to see what I'm talking about (http://vimeo.com/13213184).
"However, no attempts have been made to optimize the Java VM, the SQL queries, etc"
Emphasis being on optimizing the sql. We have run tests comparing neo4j and postgres, and postgres comes out with greater throughput for our data set, where our database implementation was done people who know postgres extremely well. Where you will see especially great differences is aggregate queries, such as if you want to count the number of a certain type of connections coming into a set of nodes, and then sort these nodes by that number. A sql database is much better at stuff like that.
Gremlin has significantly improved what you can do with graph aggregating and sorting:
// count incoming friends for each node and sort by most friends
m = [:]
g.V.inE('friend').outV.groupCount(m)
m.sort{}Actually, I've never worked on any project where our data set only needed simple key value mapping and no relationships at all.
I think the author of the blog completely missed the point of what Michael Stonebreaker (VoltDB), Jim Starkey (NimbusDB), and Paul Mikesell (Clustrix) said. All three companies make scalable, relational, ACID compliant, SQL databases. None of them are advocating using NoSQL.
Not to put words in their mouths, but they are saying you need what MySQL gives you. You should not have to move the relational and consistency logic to your application, that's the job of the database. The problem with MySQL is not that it is SQL, it is that MySQL doesn't scale. In the case of Clustrix their database is MySQL compatible, it uses the same library interfaces (e.g., MySQLdb in Python), the same command-line tools, and it can participate in a MySQL replication setup as master or slave. The difference is it actually scales, so you don't have to shard.
With MySQL, once you reach the performance limit of using massive hardware, lots of read slaves, and memcache, the only thing left is sharding. With sharding you are breaking up your database into lots of databases running on different machines. With separate databases you lose a lot of the relational ability, or you move it into the application. What the "NewSQL" databases are offering is the ability to keep a single database, and when you want to scale performance or capacity you just add another node to the DB cluster.
To those that say NoSQL is easier and more flexible than SQL, that entirely depends on your data. You should pick the database that fits your data, which might be multiple databases for different datasets within your organization. Facebook uses a lot of NoSQL, but they also use a lot of MySQL. They have obviously decided that some of their data works best in NoSQL, and some of it works best in a RDBMS no matter how hard it is to scale.
So build your app using whatever works best for you. If that's CouchDB and Redis, cool. If it's MySQL, go for it. Once you've proven that there is demand for your app, and it takes off, then deal with scaling.
Sharded MySQL means you're tying your architecture together with boogers and spit. But the problem isn't the relational model itself, and there's no good reason that model shouldn't be able to scale.
NoSQL is not just about scale (although it can be). In my opinion the main benefit of giving up relational constraints is flexibility. Find me a relational system that can credibly do offline replication to mobile devices. Each of the options I’ve seen are full of gotchas. But if you replace the relational model with an MVCC document model, you can suddenly solve an entirely new class of problems.
Read about how Apache CouchDB is being used in rural Africa to bring collaborative data and web technologies to health clinics that don’t have reliable internet access: http://radar.oreilly.com/2011/03/couchdb-zambia-healthcare.h...
These same patterns are applicable to mobile connections in the 1st world.
Scalability is one piece of the puzzle; picking the best tool for the job is a much bigger piece. Work with what you are comfortable with, but realize that an RDBMS may not be the best way to deal with a graph. Logging is probably not best done through an RDBMS. We have tools that we can use to solve problems in a more efficient way, so we should use them.
RDBMS are not a panacea and neither are NoSQL database solutions. Use the right tool for the job.
Each application is different and you need to pick the appropriate tools.
Or for MongoDB: http://www.mongodb.org/display/DOCS/Advanced+Queries
But you can't set constraints that enforce such relations, which is the whole point of the relational model.
user_to_contact_sets -> keyed on user's key
contacts -> keyed on a contact id
For searching I'd use a real search engine and index the contact data as it is stored and deleted.The beauty of using a schema-free KVS you can have a much more flexible contact document that can have embedded one-to-many objects that would require a number of extra tables in SQL for data that will only ever be associated to a single contact record.
Here's an example of a hypothetical KVS storing contact data:
// API:
db.store = function(bucket, key, data) { ... }
db.fetch = function(bucket, key) { ... }
search_engine.index = function(key, document) { ... }
// Store a Contact record:
var contact = {"id": "eric-moritz", // lexicographically ordered key
"name": "Eric Moritz",
"emails": [
{"type": "personal", "value": "eric@example.com"},
{"type": "work", "value": "workin@example.com"}],
"addresses": [
{"type": "home":
"address1": "111 A ST",
"city": "somewheresville",
"state": "somestate",
"postal-code": "90210"},
{"type": "work":
"address1": "111 B ST",
"city": "somewhereelseville",
"state": "somestate",
"postal-code": "90210"}],
"phone-numbers": [
{"type": "mobile", "value": "555-1212"},
{"type": "work", "value": "555-2323"}]
}
// Fetch the user's contact set
var contact_set = db.fetch("contact_sets", user_id);
// If this contact is not already indexed, add it and store the set
if(contact_set.indexOf(contact['id']) < 0) {
contact_set.push(contact['id']);
db.store("contact_sets", user_id, contact_set.sort());
}
// Store the contact into the contacts bucket
var contact_key = user_id + ":" + contact['id'];
db.store("contacts", contact_key, contact);
// Index the document into the search engine
search_engine.index(contact_key, contact)Even a basic chat software, you'll want to relate users to messages. Even session management.
Basically, Nosql is not just about scale but flexibility. Give me a schemaless engine anyday
I think a lot of our perceptions about technologies are influenced by the things that sit between us and those technologies. A lot of the talk about this particular issue has nothing to do with Relational databases or key/value stores.
In other frameworks, you use your tool of choice to change the database schema, then build the project and you're done. All the "model" stuff regenerates automatically. Thus, schema/model changes have no pain associated with them whatsoever, so there's no particular advantage to using nosql vs sql.
The added bonus of a build step is that you don't need to keep regenerating the same classes and SQL code every page load, but it's not really any different from the programmer's standpoint.
For prototyping purposes I'd like to see a comparison of building a twitter clone in MySQL, MongoDB and there's already one in Redis http://redis.io/topics/twitter-clone MySQL will require the most boilerplate here, setting up schemas etc.
If you primarily want to persist objects and indices that don't have good or easy analogue in the SQL world, a NoSQL solution is good solution because you don't have to bother with the unneeded SQL overhead.
If you've doing something that maps reasonably well into the relational model, well, that's what you should use. Rational technology has been around a long time and has been scaled to huge levels. And it seems fairly clear a social network maps pretty well to a relational model.
The truth is that our traffic could wildly exceed our most hopeful projections and we still wouldn't nearly be in the position of orgs like Twitter, Facebook, or Netflix, stretching the limits of what an RDBMS can do without becoming a nightmare to maintain.
However in reality not all DB's are created equally. It comes down to this: do you want to spend a known amount of time hammering in a screw, or do you want to spend an unknown amount of time optimizing your state of the art, but unproven screwdriver.
This is important for people starting a business. You don't start a business to test out new technology, that's what weekend projects are for. They need to be focusing on the product and not spending time 'just getting it to work'.
Mysql/MC aren't sexy, but it takes a trivial amount of time to set up and tune them.
There are lots of caveats to this; the entire argument can be nit-picked to death. But I think it holds for the generic webapp template that just needs data to persist and then be served back to users.
That being said, I think it is - wait for it - premature to call NoSQL an unwarranted early "optimization." For our real-time analytics app, we use Redis because we are just storing key (page name) value (hits). With millions of writes a day for relatively simple data, it is nice - and in no way premature - to run Redis on a medium-sized instance. Data retrieval is fast and after a year and a half we haven't had any problems.
Other apps, use MySQL exclusively. For our many-many forests of tag, entry, source, author, etc etc relational data, we rely on the easy oversight of MySQL and the rich ORMs that have gracefully evolved to cut through the jungle.
What's more, there are many applications that use both Redis and MySQL. My Silver gem https://github.com/tpm/silver wraps MySQL requests in a Redis caching layer to speed up queries. This is not terribly novel or that different than, say, memcached but it does what we need it to.
The point is that just like apps are as multiform as the imagination is competent, so should be the tools and strategies we use to create them and solve their problems. Yes, people will use NoSQL incorrectly. People use MySQL incorrectly and inefficiently more than not. Hell, you can use YAML in a bad way. Those aren't the cases we should focus on. We should focus on all the systems that do work.
I'd like to see, for every mind-thinky post about whether or not MySQL is a death trap or NoSQL is a nu-wave panacea, a post about how some company is making the technologies that they have chosen work for them. At the end of the day, there is no uniform march toward optimal global technology efficiency. Just the small victories of a startup that needs to record cat GPS movements, a newspaper that needs to analyze a million pages of leaks, a broker who needs to pour through SEC dumps. NoSQL, NewSQL, SQL, HithertoUnknownUnPostSQL: I say let em all in. Someone will find them useful.
What then?
My other question is that I've read several places that NoSQL makes most sense for full-text documents. And yet I see company after company with little apparent full-text data boasting about their NoSQL use. I am especially aware of this as someone in the job market seeing terms like Hadoop and MongoDB just thrown around in the job description. How? Why?
Hierarchical and network stores were popular before RDB was invented; although RDB was slower, the big advantage was that you could transform the data into whatever form you need (instead of having to use the specific hierarchy used to store it). At least, that was the key benefit according to my reading of Codd's 1970 paper.
But if you just need to persist a data model, don't care too much about the schema, and come up with the solution of, "I'll just serialize my hash table to a CSV flat file", that's just taking the ultra-simple solution, not really an optimization. If anything, pulling out the big iron of SQL when all you need to do is serialize a hashtable would be the premature optimization (or at least, premature architecting).
NOSQL isn't one catchall group. There are different tools there for different jobs. His arguments rest on the assumption that NOSQL is hard(er). If it's not hard, and it solves a problem, then it's not premature optimization, it's just optimization.
There are certainly situations where using a NOSQL database to start with gives you a competitive advantage and makes things easier down the road. This same argument is made every time a new technology comes onto the scene. People who made a living using something else will argue that learning the new thing is too hard, and isn't worth the time. If he doesn't want to bother with learning about new tools, that's fine, but it's silly for him to argue that learning them is pointless or a waste of time.
the flexibility and schemalessness of NoSQL is actually a consequence of the scaling possibilities - it's hard to manage horizontal partitioning and maintain consistency between multiple tables with foreign keys and rollback transactions in a distributed environment; and easy when you have only one table with two fields: key and value.
when, like in mongodb, the value is actually a tree-like document, and you need to have an index on some of the element, you gradually return to schema, because this element should be in the same place/surrounding in all indexed values.
i am saying that flexibility is not why you should choose NoSQL, but actually the scalability that you might need later.
(and non-relational data is kept ok in blob fields or in files and indexed by lucene - if you want (i don't))
His first point is 'all NoSQL is less mature'. Second point is 'there is no advantage to NoSQL except scale' Third point is 'you'll have time to scale later, which includes moving to NoSQL'
None of these arguments is novel. Developers pick NoSQL for other reasons as well.
I didn't find a pioneer tax to setting it up. It worked extremely well for non-uniform data. I found its design easier to handle for failover.
This entire debate is old. NoSQL is a tool. His arguments add nothing to the debate.