MongoDB Days
gaiustech.wordpress.com
gaiustech.wordpress.com
Talking about the relational model and "sound mathematical underpinnings" is fine, but I've seen dozens.. hundreds?.. of production relational databases and they've all been monsters. Most have been partially or significantly denormalized for good/bad reasons. All have their own warts, usually significant.
If people end up normalizing their MongoDB as the author suggests - Well I'd expect that. It's pretty rare to get your DB model right first time. Any thought that you can is probably tinged with madness. If you can have a good stab at it and iterate you're ahead.
Plus if you've ever worked with a TPS (or TM/2PC) you'll know they're a nightmare. Any volume systems I've worked on dispensed with them and use a reconcile/compensate mechanism.
MongoDB has gone down a replica/sharded model that focuses on Map/Reduce. How successful this is, well that's a different argument - but comparing it to an old-school big-iron mentality is a waste of time.
So you keep your normalised schema for production and have a methodologically denormalised schema for reporting.
Mind you, computer science courses generally don't teach OLAP -- stuff like dimensional modelling. My university didn't when I was there, though they've rolled it into a larger "Data Mining" course alongside MapReduce and friends.
I was arguing from practical terms and my own experience. Even in established, DB-heavy, "we have hundreds of DBAs" organisations, you can't have a porcelain schema. It evolves over time, breaks, get optimised and de-optimised.
The Oracle shops I've worked in don't have one schema, they'll have dozens. Usually because they can't shoehorn emerging requirements into the databases they have. Or because of project timelines. Or they need to be performance tuned in some special way, which in the case of Oracle could be considered a dark art.
These orgs may have just one big data warehouse, but that becomes a dumping ground. It's possible to do your dimensional modelling (I worked on star schemas in the past, but that's about it) -- but it's a big ball of mud to try and weave together.
Even then your OLAP is under-stress because it's got a lot of ground to cover. Usually they start out as marketing engines, but get co-opted into all sorts of things. The most common and troublesome is regulatory reporting. Once your data warehouse gets used for something like regulatory reporting you're a bit stuck as nobody wants to touch it.
Then, because point to point integrations between the databases is brittle and cost-prohibitive -- they break the golden rule and take data out of the OLAP and put it back into the OLTP. I won't name names, but this is commonplace.
SalesForce is one of the biggest Oracle DB users (I read biggest at one stage, but can't find the reference). Even then the SalesForce data model is so denormalised you could call it one big table[1].
I use MongoDB day-to-day and it's fit for purpose for what I do. That said, I'm not a rabid fan and an alternative/hybrid approach is always on the horizon. What I don't buy is the "the DB world has already solved all these problems", particularly when things like "relational integrity" and "two phase commit" get thrown in the mix - truth is there is plenty to be solved.
[1] http://www.dbms2.com/2011/09/15/database-architecture-salesf...
I don't know any MongoDB users that I think would be better off on Oracle - especially for the reasons cited in the article.
My own view is the critical pieces will converge. If I could have PostgreSQL with a greatly simplified replication/parallelisation solution I'd be a happy man.
Equally if MongoDB could do a COUNT query in even double the time of psql. Or compress their keys. Or have fine-grained locking... (All of which I'm sure they will get to, just when).
Maybe that's MongoDB 5, PostgreSQL 14, Rethink 3 - I've got no idea and I'm sure I'll look at them all. However, I'm not the guy rocking up to a MongoDB conference and coming away with the conclusion "Oracle had that in 1988".
My argument is not that MongoDB users would be better off on Oracle. It's that Oracle users would not necessarily be better off on MongoDB, since Oracle actually does many things that the makers of MongoDB claim it can't. If MongoDB had been made by MS, everyone would call that "FUD"...
If your intent is to rebut FUD, then statements like "But just remember that these kids think they’re solving problems that IBM (et al) solved quite literally before they were born in some cases" don't contribute. You're just meeting rhetoric with rhetoric.
If there is something specific that MongoDB is blatantly or arrogantly ignoring from Oracle (i.e. the original article) -- technical, business model, whatever -- Then call it out.
I was using MySQL for real work back in the 1990s. I remember then, them saying, you don't need foreign keys - just check it in your application. You don't need transactions either, just handle failures in your application, yadda yadda. And of course, MySQL these days supports all these things (with InnoDB, an Oracle product) because these things weren't there "for lulz" in commercial databases, they were there because people needed them and saw value in them. Now I am getting a complete sense of deja vu with MongoDB. It will have to add transactions. This is inevitable. It will have to add enforced schemas, row level locking, ACLs, and other features besides. If we remember, then in a year, let's touch base again, and you'll see I'm right. This is what I mean when I say the lesson of MySQL.
And in 10 years, there will be another product, let's call it ZongoDB, and the cycle will repeat again. You don't need this, you can't do that, they'll say, and the MongoDB crew, by now themselves old geezers (this means, in IT, in their 30s) will roll their eyes.
PS I may be an Oracle partisan, but I recognize too when Oracle implements something IBM did back in the 70s...
There's more of this sort of criticism in the following old thread, "SQL Databases Don't Scale":
https://news.ycombinator.com/item?id=690656
where a few commenters say (somewhat unpleasant) things like:
- I find that this type of FUD comes about from people that aren't good at designing and implementing large databases, or can't afford the technology that can pull it off, so they slam the technology rather than accept that they, themselves, are the ones lacking. Most of them tend to come from the typical LAMP/SlashDot crowd that only have experience with the minor technologies.
- For me, thousands of transactions per second and 10s of terabytes of data on a single database is normal. It's unremarkable, it's everyday, it's what we do, we have done it for years. And I know of installations handling 10x that. It's only people who's only "experience" is websites that whinge about how RDBMS can't handle their tiny datasets.
- Mr. Wiggins article would be better titled something like "ACID databases have scalability problems, especially cheap ones startups use"
How true are these criticisms nowadays? Is open-source still far behind, or is it (as I think) more than good-enough for 98% of use-cases?
edit: Thanks for the responses, sounds like I'll be trying out Postgres for my upcoming personal project.
The foundation of CouchDB queries are materialized view / continuous query. PgPool provides result caching.
Getting there... http://www.postgresql.org/docs/9.3/static/rules-materialized...
(The rest of his examples are complicated Oracleisms and you'd probably get pretty far with MVs.)
Regardless, I had no idea DB2 was so smart. I guess you get what you pay for.
I cant answer your question, but looking for feature X in other products is the wrong approach IMHO. You're often better served by looking at what problem feature X is solving, and how that problem is other places. Sometimes the problem doesn't even exist.
Behind Facebook and Twitter there's a big usage of MySQL. Not even PostgreSQL, MySQL.
Of course it's a pain for them, but it works and they're profiting from it.
The question is, how much they are saving by not using Oracle or other proprietary DB? They priced themselves out of the web, now they can't "cry to mummy" about it.
"For me, thousands of transactions per second and 10s of terabytes of data on a single database is normal"
It helps when you have dedicated top notch hardware (especially true some years ago) and the startups have to work with EC2 and EBS (ok, there are some better choices, still)
http://www.oracle.com/us/corporate/pricing/technology-price-...
Processor License cost (excludes support):
* Standard Edition: $17,500
* Enterprise Edition: $47,500 (that's the one with Materialized View Query Rewrite)
* Partitioning option: $11,500
* Advanced compression: $11,500 (basic compression is apparently slow?)
Per processor.
This is why all these other databases exist. Very few businesses, certainly not many startups, have the kinds of value of the data stored to warrant that kind of cost.
I believe in the beginning of the web, Oracle wanted to push a per-user pricing model.
Yes, if your website has 10 users, it pays 10x$PRICE, 100, 100x$PRICE
Nobody could work with that (with Oracle)
Perhaps if somebody could offer {DB2|Oracle|Informix|Sybase} As A Service, like what many providers do with PostgreSQL and MySQL, it would be a different story for startups.
My profression involves moving MSSQL and Sybase databases to MySQL and MariaDB, then I go and read some of the features from the article in Oracle and DB2 documentation and think to myself "I ain't helping fucking nobody", the higher-ups in enterprise and startups usually only see upfront costs, not long term benefits.
As per your first question; they are getting there. PostgreSQL now has a built-in function for managing materialised views, in MySQL/MariaDB that is still hand-rolled.
MariaDB however does have Virtual Columns and Dynamic columns (to tackle nested tables, as per the article).
Look to MariaDB for the MySQL developments, as Oracle are bringing very little to the table in terms of matching MySQL features to Oracle (not surprising). But they are resolving important security, and infrastructure issues and improving InnoDB heavily.
Also PostgreSQL does have HSTORE, for storing JSON data types, very tasty.
The article also reminds me of how a father and son went to a Microsoft presentation in 2000, where Microsoft showed their solution to the tricky problem of integrating multile backend servers. Their solution was to have front end tiers close to the client, and the client getting thinner. The son was very impressed. The father said 'that's what IBM did before the 70s!'
Seriously though, we are using MongoDB with great success at StartHQ (https://starthq.com) having done a lot of work with relational databases before. It's a great fit for startups where the schema is constantly evolving & the amount of data stored can be quite small.
Also, by talking to the DB directly, without an ORM, we can keep things really simple. I dread to think of what the same code would look like if we were to use a relational database, either with or without an ORM.
Now, if I have this complex document normalized into several tables, how am I going to easily shard several tables, all such that I can execute successful joins that only need to execute on one leaf node? What if I start reusing a small piece of data in one of these normalized tables? I might be forced to go between network nodes to get this data.
Normalization is like premature optimization. I can take any program and modify it such that every piece executes optimally fast, but I am likely to compromise on clarity or to add complexity while doing such a refactoring. In the end, it probably got me no real-world performance boost that mattered. 80/20% rule and all.
Same thing with normalization: automatically making all my data fully normalized from the start is a like a bad premature optimization habit that we are forced into with relational databases out of A) sheer habit and school teachings, B) lack of easy support for nested structured data.
The MySQL guys are very into sharding too but of course without hash joins, you can't join big tables anyway, so this "optimization" is easy because it costs you nothing.
Yes, I can totally see how Postgres and hstore can run circles around key-value storage.
And by "all that", I would love to know how to use the open-source versions of Postgres or MySQL to have transparently-sharded tables with joins AND have that be in a replicated environment where I can do real-time failover.
It seems all anti-NoSQL rants are the same whining and refusal to understand. "Sound mathematical base"? Really? So, if I can't describe something mathematically (I can, by the way) it's not worth it? A computer program is a mathematical description, there you have it.
The relational model breaks for very common use cases nowadays. Yes, maybe you think it's fun to do a query across how many tables to get the information you want, but if your website has a non-trivial traffic then the solution is usually to add more cache.
That's (one of the reasons) why PostgreSQL has hstore. Beyond the fanboy insistence that you can do everything with relational DBs, Postgres have accepted the reality that you need a more flexible data structure.
Edit: yes, please continue showing your contempt while I have to code around the limitation of relational databases.
"MongoDB solves problems MySQL didn't solve at the time."
How do you migrate documents created before you changed your schema to documents created after? Or do you deal with an entire collection of documents all wildly varying in shape?
Progress!
If you have a system setup where different parts run on different services you don't want to have to co-ordinate and resync all the applications when their underlying data structures change.
Web services solve this by having a agreed contract of communicating between systems, service A knows of a better way of accessing its own data that service B knows of accessing system A's data.
- Easy to set up (including replication)
- Fast
- Data in "JSON" format (it's BSON internally)
- Javascript (used for map/reduce mainly)
- Good availability of libraries
Theoretically, CouchDB would have been better, but I've heard some bad stories (remember UbuntuOne)?
I don't have much experience with other NOSQL DBs (except for Redis, it's great, but I like to call it a "DB Toolkit" - I'm not dissing it, I love it)