Schema-less is usually a lie
blog.mongohq.com
blog.mongohq.com
If the DB is determining the schema and your code just mirrors it, you will inevitably do things like triggers, stored procedures, etc. that essentially put application code inside the DB. This makes testing and maintenance of such things all but impossible, even if it does make the DB queries fast. At the same time, your application still needs to mirror the schema of the DB or things will break.
"Schemaless" (query schemas) isn't better per se, but having the mindset of putting the schema in the code means you can write tests to ensure the schema over time. Depending on your perspective, one might be better than the other, but as they say "knowing is half the battle"
I don't believe this is necessarily true. I think it depends on the structure of the enterprise.
The traditional relational database seems to center on what could be called the "traditional large enterprise". This large enterprise tends to have many projects sharing the same static data and here a single store for that data makes sense and that store can and should have a standard interface, which can be embodied as stored procedures if necessary and would have to be well-defined enough that individual applications can deal with it.
That's the "traditional large enterprise". Today, however, we have large companies where single application has become the essence of the company (Google, Facebook, etc)and the one-datastore, multiple-applications model no longer makes sense and it does make sense to move more data manipulation the application level.
Still, when making pronouncements about what works best, I think context is important.
Which "one application" is the essence of Google?
While there are reasons that running all those applications for, say, Google on a shared RDBMS backend isn't the right answer, the reason isn't that Google has a single application that uses all its data and so doesn't have to worry about coordination between different applications using the same data.
Maps is maps, Search is search, YouTube is YouTube, GMail is Gmail etc. Those are not apps providing different views of the same data, in the way the parent describer enterprise apps.
If testing and maintenance of database objects (tables and other relations, as well as procedural code in triggers, SPs, etc.) is "all but impossible", the problem is with your DB maintenance policies and practices, not with where you are putting code.
It may be the case that lots of places have bad database maintenance practices (just as lots of places have bad maintenance practices for non-DB software), and it may be that in those environments, when you have a limited scope of influence, routing around those bad practices is the best solution. But it is a mistake to present that as a general solution.
Developers have been maintaining such apps for decades without any problem. And everything is testable , even a stocked procedure.
If you dont care about data integrity, then make your application responsible for your schema...Some do care. Non-rel databases guarantee 0 data integrity, with very little performance gain over relational-databases.
Finally most frameworks allow developpers to generate the db schema during development without writing a single query. with all the cache layers like redis and other goodies , there is little to no reason to use non-rel databases.
guarantee ZERO data integrity?
IMHO the database these days SHOULD be fairly agnostic. Also normalization, which is the core of your so called integrity actually REDUCES performance. At my last job, our main VIEW into the data took over 28 join operations for some highly normalized data (much of which had to overcome some bad data in the db).
We setup a MongoDB database for searching against, as well as being able to pull up a single record without dozens of join operations, and it was a LOT faster, against real-time data... the search that was replaced was a batch process that recreated a single table every half hour. In this case MongoDB was a much better fit.
Sorry, but SQL databases don't guarantee any data integrity either. It's up to the developers that implement those schemas... and the fact is, for most of them, they are better off doing that in their primary application code.
Also, if you are using an ORM tool to "generate" your schema from code, then what advantage does said schema's "integrity" give you?
I really wish MongoDB had support for a strong schema in the DB besides indexes (in addition to existing support for schema-less). I only actually want schema-less at most 20% of the time, but I am stuck with it for everything outside of indexes.
I guess what I want is a schema-ish database, where I can say "these 5 fields are required and must conform to these rules. Anything else is fair game."
http://thebuild.com/blog/2013/07/02/postgresql-as-a-nosql-da...
For that matter, if you use NodeJS, then said API and the backend structure are fairly trivial. I used NodeJS to create an API for queries against MongoDB, and it allowed for me to normalize and check data in said queries against the data. It also allowed me to do programmatic elimination of sensitive portions of data from the front end without much effort at all.
It seems what you really should be creating is an API that your application uses. I'm a proponent of the Data Storage Layer being as dumb as possible. If it weren't for the built in indexing, and common access structures, I'd be more inclined to roll my own. I do think that MongoDB does a nice job of striking a balance between say Couch and MySQL (I intentionally use MySQL here instead of PostgreSQL or MS-SQL). I also think that RethinkDB within a year or so will likely be a better option for many of those thinking about MongoDB.
'...but every time I see "Schemaless" in MongoDB, I think "oh, so you're implementing schema in your application?"'
But if you only have one application, you don't need centralized documentation.
Now that's quite interesting statement. I'd say it is true for humans but false for computers and it's paradoxical to think about why this is.
If computers were an intelligent as humans, you wouldn't have to worry about giving your program any structure because the computer could change that structure later. But sadly, spaghetti code isn't what it's cracked up to be.
Similarly, an application where don't bother thinking about your schema beforehand isn't going to be application which you can change easily later.
So, whereas a schema based application that models a user with an address is unable to cope with being given two addresses. A schema-less approach where it just stores blobs simply stores what it was given. If code that reads this is unable to make sense of multiple addresses, it will raise an error. Not necessarily unlike a user being told to send a package to an address, but given a list of addresses.
[1] http://www.infoq.com/presentations/We-Really-Dont-Know-How-T...
I think that's all you need to read in this post. It's... rather thin IMHO.
MongoDB specifically allows for additional indexes to be used as part of general queries, or the aggregation methods available. MongoDB even has some nice built in features for querying against geolocation data. They do have some limitations, but in general it works quite well.
Your analogy comparing non-sql to a trashcan is a bad comparison... a better comparison would be to a set of file cabinets.
I find that MongoDB tends to perform at least as well as an SQL based solution in general use. Some use cases are better suited to SQL (anything with multiple records in a single transaction as a hard requirement for example). It really depends on your needs.
In my last job, we were presenting search, and display of classified listings for cars. The normalized data required, iirc, 28 join operations to get most of the data for a single record (for display), about 12 iirc for the search support (not including geo/location based searches), and a second lookup for related data.
This could be replaced by a single query in a non-sql database. There is a real cost to these kinds of structures in SQL... There are a LOT of use cases where a single record structured as a complete object is much better than having to break up said structure into dozens of fields.