Mongo DB is web scale
xtranormal.com
xtranormal.com
To the armies of bloggers parroting slogan after slogan, ricing benchmarks so far removed from real-world applications as to place themselves somewhere on the spectrum between meaningless and malicious in the name of pageviews, well, that's not helping anyone.
It's great to see people get excited about new technologies. But if tech/startup culture is one that embraces and celebrates the fail, then we damn well better also talk about the different areas where many different datastores are not appropriate choices, and where some of them downright break, risking extensive downtime or corruption, require unreasonable amounts memory for indexes, or multiple machines to provide for durability, or return inconsistent "postmodern" results, or what have you -- without it being understood as a personal attack on any one individual, company, or sector.
The productivity of this video is debatable, sure. But perhaps we can appreciate that it parodies the "magic scaling sauce" image that the "tech press" (face it, we're it) has given surprisingly young but maturing technologies. There's no magic sauce, no drop-in answer to "web scale," and certainly nothing fraught with difficulty. It's data. If you have a ton of it and it's valuable, it's worth mountains of expensive programmer time and effort to ensure its integrity, accessibility, and utility based on the storage and query requirements of an application.
If you read Chris Date's books (I recommend starting with _An Introduction to Database Systems_), he really hammers this point.
There are people who hate SQL because it's too relational, and there are people who hate SQL because it's not relational enough.
Anyone watching this might believe that MongoDB has a major flaw where data isn't written immediately and you have no idea if it has successfully been stored to disk. Whilst it's true that inserts are not immediately written by default, you can a) change the startup config to set the delay time b) force a write of all pending changes from the command line c) force the write from your call to the insert/update/remove/etc method in all the libraries and most importantly, d) request the library method wait and return the response from MongoDB so you can determine if the write was successful or not.
This means you have complete control over when you need fast inserts at the expense of potential data loss, or when you need to be certain the data has been written.
It's certainly not the same as writing to /dev/null.
Maybe the developer can have other ways of guaranteeing consistency.
MySQL had severe data consistency issues before InnoDB came around and even today, on InnoDB, there is a variety of situations that will cause silent data truncation or silent data loss.
The funny thing about relational databases they model relationships fairly loosely, mostly understood and expressed by the user, not the system. Object systems have a stronger idea of a one way relationship with a reference anyway.
How many systems should care about relationships. All of them can.
Here in Australia they want to put everyones health records online, and connect all the hospitals, doctors, drug stores, testing labs, etc. to those records.
So far they have spent heaps, $500 million'ish, on a system that does this:
http://www.abc.net.au/rn/healthreport/stories/2010/2975642.h...
I think this is a good place for something like a non-relational database. You just create a sequential list of events for a person and you are done.
You can use something like CouchDB. If you loose your network, you can have a local replica at your local hospital and doctors office. Everything re-syncs later.
I think the whole thing is in analysis paralysis because they are trying to systematize the health universe so the meaning of all data is unambiguous. While that sounds grand it will take forever.
I blame the whole failure on a drive to systematize data, which SQL databases foster. If you just view health records as a bunch of bits of paper, and just say we want to store them in one place, in a way that is vaguely sequential, the problem is much simpler.
CouchDB is after all web scale, shard friendly, and there is no impotence mismatch. After seeing this video I am always going to want to say it that way.
It's simpler, and tremendously less usable.
Absolutely. To go back to the original point, I don't think it's document-centric or not that causes the massive cost overruns. Instead it's that big consulting companies know that they can absolutely ROB the government blind and get away with it. Here in Canada we've seen the same farce with the long gun registry, with the health records here in Ontario....just about any government-related project is a boondoggle.
It is great perceived importance that makes a boondoggle.
A few employees, government or otherwise, doing a project of no great significance, will not waste much.
Being insignificant creates a need of efficiency.
It's pretty clear - i.e., his use of the term 'schema' - that he meant in-DB relations, rather than in-code relations.
"Why not write to /dev/null? It's fast as hell"
"Does /dev/null support sharding?"
NoSQL has become an increasingly broad term used to indicate some other storage strategy than traditional SQL databases provide. NoSQL databases tend to operate on a more basic level and omit things like transactions in favor of performance. Schemas are also disregarded to allow arbitrary storage requirements to be addressed with ease; although, you lose the traditional guarantees that a schema can provide.
This is just my opinion, but I think it's a lot easier to setup a clusterfuck storage nightmare using a NoSQL database. If you're unfamiliar with SQL and NoSQL, you're most likely going to end up in a safer and more recoverable situation going with something tried and true like PostgreSQL.
If you're not sure whether you need a NoSQL system, then you don't need one. They serve a use case and performance requirement that very few websites demand. That being said, when the need does arise, they can be an indispensable for scaling out a large website. Like anything else, they're not a magic bullet that you just drop into your storage strategy and go on your merry way. Rolling them out requires lots of planning just like any other large scale deployment does.
SQL represents 40 years of experience designing general data stores that work for the widest set of applications. SQL databases solve many very hard problems which you are probably not aware of.
If you jump into NoSQL first you will be reimplementing SQL features in your application code, and doing a shitty job of it because you have experience with data stores that actually solve these problems well.
The reason for the existence of so many NoSQL databases is the rise of web applications and the need to scale massively. However the majority of apps will never need to scale beyond a single well-tuned database server anyway. By the time they do you will have hard problems to solve regardless of what data store you used. The advantage of SQL is that it's a fantastic hedge on the evolution of your data usage patterns because it is designed to support ad-hoc queries well, and the schema prevents bad application code from thrusting your data into chaos at the first occurrence of a small bug.
Realistically if you knew you had to build an app for 5 million daily users, and you knew exactly what it was going to do, then an SQL database very well might be the wrong choice. But in the real world you have a long road ahead before you hit that scale, and you'll have real data to determine what kind of alternate data stores can best handle your load. Personally I'm a huge fan of redis, and its ability to scrape bottlenecks off a MySQL database in a piecemeal fashion.
This is a false dichotomy.
Just because you want to use some unstructured data (which was your original example) doesn't mean you need a new data store that's optimally suited to that. You can store documents just great in an SQL database or in the filesystem.
'Best tool for the job' is an oversimplification.
The nice thing about Mongo is you can query that hash without having to have an index table.
To the parent: thanks for the edifying comment. I do have some experience with SQL and relational DBs, but was thinking of using noSQL for some projects. Your point to thoroughly learn the relational model is well taken.
Okay, so we have a domain object, Foo. Foos represent individual instances of a Foo that a user has, but we want to keep general information about the different standard types of Foo, so we also have a relation between Foos and FooTypes. Oh, and each FooType can have a few different sizes, and some FooTypes are the same sizes as each other, so we also need a FooSize. Not only do we need to relate the number of sizes that each FooType could have, but when a User has a Foo, we gotta know which sized one they have. All this stuff... it's complicated.
It probably would have been much easier for me to have just done each FooType up as a document, with an embedded array of sizes, and then each Foo gets a document, with its own copy of the data. Yeah, there's nothing saying that you can't store de-normalized data in a relational database, but if you're not going to use its features, why not just use the tool that's designed for that use-case?
The only solutions i have found are to check the length of the distinct query, which takes too long for a large result set, or to write a map reduce function which takes longer than I'd like and is a large amount of code for functionality that should already exist in the db.
(I withstood the urge, you probably should too)
If you want a lot of speed and can tolerate the possible (even if unlikely) data loss, use NoSQL.
If your business requires that your data is guaranteed and always up-to-date at any moment, then use RDBMS.
For instances, what CouchDB treated as a major bug, is the accepted behavior of many relational databases. (Eg, data isn't lost, but must be recovered via a long-running process should there be an uncontrolled shtudown.)
Riak and Cassandra also have modes that treat durability as paramount, and give you better assurances than MySQL or even commercial RDBMS products.
"If you can type, you can make movies."