Announcing MongoDB 3.0
mongodb.com
mongodb.com
It's good to see they haven't sacrificed the hype while working on other features :-/
(Not affiliated with them in any way, I'm just testing rethinkdb since I have to use Mongo at work and hate it. I especially like their "stability" page: http://rethinkdb.com/stability/ )
It seems ridiculous, really -- developers commit to something because it saves them hours at the start, versus considering the months of work following -- but many technologies followed this curve. MySQL was a lot easier to start with than pgsql, so it won the early battle. PHP was a no-brainer to start with, etc.
The lesson is that your tool should default to essentially no configuration, no security, and simple beginnings, and it tends to pay off.
Though I really like RethinkDB, the lack of automated failover handling is about the only concern left keeping me from them. Because of the tooling around it, I'm using ElasticSearch om my next project over both of them. I've also used Cassandra in the past, which was great in some cases, harder in others.
For the past couple years, I've tended to pair an SQL database with a key-value/document store depending on the needs. SQL is the canonical source, and saves are mirrored to the distributed database for wider reads. Most no-sql databases sacrifice parts of CAP/ACID in different ways, and can fit certain growth needs.
When you are faced with paying a over $30-50k/month (hardware, licenses, bandwidth, redundancy) just to host your sql database on hardware big enough to handle the load, or distribute... distribution becomes an option. For a company with under a million a year in gross revenue, spending $600/year for the database alone is not an option. Even for larger companies it's probably not the best idea.
There's something to be said for ease of configuration and administration in that mix.
That's a lot. Could you elaborate what kind of hardware and software one gets at that price level?
Here's a couple disproving your claim:
https://www.scalebase.com/the-state-of-the-open-source-datab...
http://www.datastax.com/wp-content/uploads/2013/02/WP-Benchm...
Those with different data access patterns may not be as fortunate.
As a side note, the hype from MongoDb the company equals the anti-hype you get on places like HN, so I guess they even out.
After using MongoDB for years (very happily too, I'm not one of the anti-Mongo crowd), I ended up switching and it's been great.
So...Mongo "can be the new DBMS standard for any team, in any industry", and they don't even support transactions yet?
For some, sure, okay, maybe, but nearly everything I've done with databases falls quite nicely into a relational model, which doesn't seem to be MongoDB's strongest point.
I'm pretty sure that Facebook is using a wide database (similar to Cassandra) with reverse indexes generated at the application layer. Yahoo, it depends on the developer group, but something like mongo in use wouldn't completely surprise me. Google has published a lot of white papers on how they use data (see Cassandra and Hadoop from some of those lessons).
MongoDB is itself a bit of a bridge... replication setup is far easier than with a typical RDBMS, and you can break data into documents for simpler queries. Mongo as a persistent caching layer is also an option depending on your load and needs with SQL as authority, mongo for public search/display.
I also said, with a caching system in place. (data and/or output). It really depends on how interactive and static your needs are, which will vary... but most sites are mostly static and very cachable content...
For an example, look at StackOverflow... they were really huge before they needed to start breaking their data up more... Of course their SQL servers are enormous by anyone's standards.. the point is you can scale vertically a lot, and if you can handle a certain amount of downtime, or reduced functionality (if db is down), you can go a long way on a single database server... add caching, add replication... all of this before needing to break data up. (In mostly read scenarios)
If you think that automatic sharding is the answer to handle high loads then your are in for a very exciting surprise specially when things go awry and all that missing ACID compliance comes knocking...
As far as ACID, many applications don't need it, for example in a game application users will forgive you if you lose their latest scores, but in a banking application users will not forgive you if you lose their last deposit. So it depends on the application.
You've said 'need automatic sharding' multiple times now, but you've still yet to explain why you need it.
Yes, sharding is important for write scaling. But it doesn't "need" to be automatic.
Plenty of setups have completely capable manually sharded workloads running on MySQL. Some would consider the manual approach a plus, vs. relying on the database to automagically do it.
It can be nice to have automatic sharding, assuming it works properly. But nice and needed are two different things, an I think it's an important distinction. You're basically saying if you shard you have to go with mongo and you can't go with anything that requires you manually handle it.
This point is repeatedly made about mongo, php and so on: to me it sounds like "This gun is so easy to use, it allows me to start spraying bullets around with no training and no knowledge of how it works!" or "This car is so simple I can go 150mph on the freeway the first time I got behind the wheel!"
Databases are massively powerful systems and they are used almost exclusively for business applications and matters of import. Why is "I can use it to make long-term consequential actions without any idea wtf I'm doing!" a plus point?
It's "basic" as in your data being your foundation, not in rudimentary. And to be fair, most of the things that aim to replace SQL tend to be about as complicated. TANSTAAFL.
It's also where people get into trouble. After they've done this and built their MVP they don't go back and rethink their data decisions. Then scale issues occur, among other things.
But it's also very easy to change your schema while iterating to serve your use cases better. But it's up to the users to change things and to learn which types of patterns scale better and to actually learn how to use the database to their advantage after gaining some experience.
Learning how to model documents effectively is rather important.
InnoDB by the way is an Oracle product.
When Google, Facebook, Twitter, and almost every other big player picks you to underpin their platforms I think you might have made some good choices.
( Yes I know all those companies also have their own DB's, but remember that they almost all started on MySQL and only switched when they out grew it )
However, MongoDB can not guarantee any operation that affects multiple documents will affect them all - if there was a failure after some documents were updated there is no rollback.
However, between atomic single-document updates and the ease of implementing optimistic locking patterns, you still have options.
Multi-document locking and transactions over a cluster of numerous shards is difficult. This is one reason why distributing documents is easier than rows of data. I'm not even sure they will ever have distributed transactions (using a 2-phase commit) because of how slow they are. However, single shard transactions should be easy to implement. But it will be use-case specific I imagine - you'd have to lock by shard-key to guarantee related documents are on the same shard I gather.
A few months ago they closed up some of JIRA (comments and a bit epics/sprints), so some issues would be created and have no comments exposed. I thought they were working on transactions in this release (while they were still around 2.7.3 or so), but I guess that's not the case.
Lastly, I wonder how this will impact TokuMX, both positively and negatively? Really whether they'll be making their crown improvements into a pluggable storage type/engine, and if users would be able to take that and plug it into Mongo without patent issues on their fractal tree indexing.
Nonetheless, exciting news for some of us who are building on web-scale, MVCCABCD databases which operate at the speed of hype and web-scale :)
EDIT: saw that Tokutek has created TokuMXse, only after my post
You can't just throw random data into it and expect it to work. Proper document modeling and understanding write concern solve almost all the problems people have with MongoDB.
It still has a ways to go but so many of the comments here are issues from way back in v2.0.
Regardless, it's still very popular and only gaining popularity. And improving.
This is why I'm not a fan, its "sweet spot" , in my experience, is not very big and not very interesting. It requires a small data set where at least the indexes (but ideally the entire "working set") fits in memory on a single server, in which case it can get pretty good read throughput. But there are a lot of options that perform great in this use case, including just mysql or postgres. Maybe the flexibility of being schema-less is a win for some people, in which case great, I'm glad it's available as an option.
But when the company claims things like it's the best database for every use case, people are bound to use it wrong, get burned, and then hate it. Where I used it it had been the exact wrong choice to make, a data set bigger than ~2x a single server's memory (and growing) and with relatively high write throughput, and it was an absolute nightmare.
And 40% more hip. Gotta love those numbers. Also, don't forget about the new WiredTiger storage engine - sounds way cooler than, say, InnoDB! Hype is strong with this one.
[0] - https://courses.cs.washington.edu/courses/cse444/08au/544M/R...
I'm not a fan of bashing any product, but after using MongoDB in several projects, in legitimate document database use cases, I just can't find any bright side.
So much claim without any real information to back it up.
I don't understand all the angst against this technology. If you need rollback-able transaction-guaranteed, exactly-once consistency, normalised schema with triggers left right and centre, you knew long ago that this wasn't the tech for you. Why is everyone so negative? I for one cannot wait to be able to store 4x more data on the same SSD, and using less RAM. In my view, Mongo is a massive enabling technology for startups on limited budget.
Overall, I enjoy working with MongoDB, because it (generally) maps directly to your object - there is no need for an additional layer (such as an Object Relational Mapper (ORM)).
However, you have to be more careful with your data structure. For example, having sub-arrays in sub-arrays is probably not a good idea.
I will be happy to share more, so feel free to look up my profile.
I get that some people have had bad experiences with Mongo. Sounds like you're one of them. Why not share your experience rather than just spread FUD?
It's worth noting that a lot of the complaints people had about Mongo were with things that were clearly documented (e.g. no confirmations on writes by default back in the day).
We started having issues around 100GB that forced an upgrade to SSDs. That lasted us til 2TB, at which point we switched.
Did you use RAID on your HD's? How much RAM did you have? What types of CPU's and how many? Were you updating documents often or writing new ones? Did you have compound indexes built properly? What was your latency requirements? Was MongoDB able to saturate your resources? Did you read off of secondaries at all? Did you shard your DB and if so what was your shard key? Was it random? Could you use it for reads too? How much of your data set did you need to possibly use at any given time?
If you had to "upgrade to SSD's" at 100GB you either had amazingly poor provisions, were doing most everything wrong or both. Messages like yours indicate to me a poor user - not a poor database.
10gen couldn't find fault with our setup other than our need to do updates. Which is weird, updating a database? Inconceivable. Anwyays, spinning disks apparently couldn't deal with the seeking. Our lock percentage was incredibly high, which makes sense given our "high" write load. It was during this evaluation that we found many of their "critical statistics" were non-deterministic (like the padding factor, rerun and you get all values 0.0-2.0. Probably fixed by now).
We had 1 index: primary. We did no other lookups. MongoDB managed to saturate only the IO subsystem. We did read from secondaries, but that didn't help overly much given that the write percent was high on them too.
On their recommendation we upgraded to SSDs. That dropped our lock percentages and it stayed low. We never made it to sharding because we simply stopped caring about MongoDB. I inherited it and was responsible for it, but ultimately decided to get rid of it in favor of Cassandra.
It's not a "poor user" when you follow their best practices of the time and have a poor experience. It (is|was) a poor database. The storage engine was braindead, their acquisition of WiredTiger admits that. It's at least improving, ever so slowly.
The problem with updates in MonogDB is that they are inline and flushed to the filesystem that way. So if you have documents which are growing then the old ones have to be 0'd out and a new one written, which results in a ton of wasted space and an unbelievable amount of thrashing. As you saw, ridiculous amounts of IO time and even with SSD's you get pretty bad SSD wear on your disks.
There's ways around this by pre-padding documents or using buckets if you are pushing into arrays but I can see why you didn't bother. Also, the write lock back then was system level - yikes.
Around 2011 MonogDB was incredibly awful. I will say it has improved manifold since then.
My use cases may have helped (they were very appropriate for MongoDB, but I suppose I would have chosen something else if they weren't).
Further, what's your read / write pattern? Do you do replication? Do you have requirements around writes being visible to reads on slaves near immediately? What is your data integrity requirements? All of these play into whether MongoDB is the right fit for someone's needs.
The point remains, MongoDB is good for some use cases, but if you buy into the marketing hype of "the new DBMS standard for any team, in any industry", then you deserve the inevitable pain you will endure (that said, I do feel bad for your likely eventual replacements who have to clean up the mess from your poor choices).
Of course it's not good for everyone's use cases though. Nothing is. Did anyone honestly believe that? I mean, even with the hype, did anyone truly decide to not evaluate their _database_, and believe the ad copy blindly?
Ultimately, I wouldn't have chosen MongoDB if it was bad for my use case. I tried it and other things -- read the documentation, made benchmarks and test cases -- and chose it after a period of evaluation. To do anything less for your database seems irresponsible (I'm looking at you, people who were surprised about the original default write semantics because you didn't read the docs).
Tangentially, the personal attacks ("which is fitting if you are a MongoDB advocate I guess", "that said, I do feel bad for your likely eventual replacements who have to clean up the mess from your poor choices") add nothing to your argument. They make you look foolish. Also, my replacements are still using my system (quite happily), over a year after my departure.
Granted, their product is not excellent and it has major flaws and lacks some common DB features, but they have the advantage of developer-ease-of-use and low barrier to use.
I think this product is simply not yet "DONE".
Plus, I think you'll be able to just buy Toku's storage engine if you want it when they make it pluggable with the new MonogDB storage engine API.
Pros: It's great for building your MVP. Super easy to get up and running, super easy to work with at a superficial level (I'm talking about storing data and querying here more than ops and administration), dataset can grow to a fairly large size without you having to think about your database at all, freeing you to think about your business. Mongo was my first brush with "schemaless" and that's a big advantage particularly for early-stage projects where things are in flux. I also enjoy using JS as the query language, because I like JS, but YMMV.
Also for certain use cases, particularly the single document case, it's probably one of the best solutions available. However, in my experience, this ends up being a limitation if your data doesn't fit nicely into a single document, and the types of data that do fit this use case are rare.
The cons you've probably heard before: no transactions, no joins, not ACID compliant, until very recently writes locked the entire database, scaling/sharding just works until it doesn't.
You'll get easy replication/failover (rolling updates of the DB server are easy), ease of use (with node.js for example) and very decent performance...
Good for an application where you need frequent access to a "row" except that row is complex, with subobjects, sets and arrays (in this use case, a relational model performs much worse because of joins)
Not so good for an application where you need frequent access to lists of tabular data and list/details access (where relational models shine)
Your mileage may vary. Some use it for logging, for analytics, for which I have no need or experience. GridFS also works very well (but I use nginx as a cache in front of it).
I was confused about why there was no links to documentation on the .com, until I stumbled onto the .org, which lists actual features right on the home page.
Also striking: no mention of 3.0 on the .org home page at all, and on the downloads page, you have to scroll down to "Development Releases (unstable)" to find any reference to the 3.0 builds.
https://www.mongodb.org/downloads
Quite the split personality between the two sites.
Many more people are collecting and working on larger data sets than they would have ever dreamed of years ago. I credit MongoDB, in part, to allowing developers to do this and encourage them to try. Sure it has it's limitations and may not be as great in areas as other databases, but, wow have a lot of people tried and succeeded.
tl;dr, If reverse survivor-ship bias is a thing, I think mongodb has it.
That said, there are situations I'd be more inclined to reach for ElasticSearch, Cassandra or others. As soon as RethinkDB has it's automatic failover story in place, I'd put it above MongoDB though, slightly nicer developer interface, and admin is much nicer than others.
Then again, if PostgreSQL got replication with failover in the box, I'd probably use that far more.
The marketing copy still pretends that TokuMX doesn't exist, though - it's had these features and more (including transactions) for quite some time now.
Just give us facts, plain and simple. What it improves and how it improves it. When there are that many adjectives about the project, it just causes me to tune out a bit.
Was this supposed to be part of a sales deck or something? (An as aside, I use mongo and know its pluses and minuses so reading all that hoopla is just unseemly.)
[1] https://blog.serverdensity.com/does-everyone-hate-mongodb/
In reality, MongoDB works well enough to be where they are, armed with the money they have.
MongoDB gets a bad rep, some of it their own fault for sure, but it's not as lousy as some people here make it out to be. Careful object modeling goes a long ways. Wired Tiger will be a massive improvement on their current storage engine, which is pretty awful for write-heavy loads to be sure.
The default since is to get an ack from the server. You can control this on each write and make it more durable (make sure it is replicated to 1, majority or all replicas) or faster (just get to the server). The default is to be written to the primary server and put into the transaction log so if it doesn't get committed to disk (possibly 60 seconds by default) it can recover. The transaction log is flushed (fsync) every 100ms by default and is configurable. You can also specify that the write is only acknowledged after the transaction log is synced. Anyways the default is it is put into the transaction log and then acknowledged.
I guess I'll give it another look then.
But I'm still a little wary of a database built by people who ever thought that such behavior was acceptable for a database...
I also agree, having a setting of "fire and forget", although useful in some (limited!) use cases, is not a sane default.
I tried getting the most out of you. I treated you like a princess, I indulged you, your wishes became my wishes and your thoughts were my thoughts. I stopped listening to all that criticism around you and thought of you as of a misunderstood child, even as you would refuse to do even the easiest tasks one could imagine.
I gave you everything there is to give, but you broke my heart. You left me in the most critical moments. I trusted you, but you would go your own way. Tears were shed and countless sleepless nights were to follow.
Remember that night you disappeared without leaving a sign? I sent you messages which got a response only after many hours. You didn't give a damn about my needs. One you told me: "it's not me, it's you". And I believed you, I truly did.
It's over now. It's been some time without you, and I'm getting better. I have discovered, that not everyone is like you. Some DBs care, they really do. You can trust them, they give you their everything.
I'm still struggling falling in love again, but it's getting better. Don't write me back. Goodbye.
Other things like continually pointing to each other, not figuring out a master for hours has made some people tear their hair out, despite the documentation saying it should behave in a sane way.
The result is that many developers who bought into the hype train are now salty because they encountered the limitations at unexpected and unfortunate moments that might have cost them uptime or data.
For most applications I usually use Postgres.
Then they chucked it over the fence to Ops who were expected to make it work in Production, and everything about it was found to be woefully inadequate, and the Ops guys were forced to carry the can. Those guys feel - quite rightly - that they were shafted by MongoDB, both the company and its fans.
Without telling you.
There's "one lonely dude complaining on the internet", which you can safely ignore. And there's "half the audience has horror stories to tell", which should probably be a red flag.
- had a global write lock.
- mmap was the extent of their io and memory management
- called rand() in in client to decide if it should report errors
- said a write was committed before even sending it out
- released benchmarks that measured how fast they could put data into local RAM, but compared it to real dB's
- lost data/corrupted on 32 bit builds and dismissed that as being by design
- is generally better off replaced by postgres but hyped by people that don't know any better
That's off the top of my head. I'm sure there are competent users and it makes sense in some cases but the general feeling is that it's over hyped for the little it is. Their early mistakes make people very cautious about trusting them with important data.
And by over hyped, I'm also taking about people that go on about how amazing it is to have a document store, but ignore the obvious idea of using pg and making a table consisting of (key, document) and using XML or JSON operators on top.
I might be totally wrong.
Edit: I guess for the same reasons I am being down voted without comment.
It's sympathy karma because the poor guy got dumped. We've all been there.
looks away in the the distance CVS. I had such high hopes. You're dead to me now. sigh
I wouldn't have minded it as much if it wasn't so long and including some detailed issues OP faced with Mongo.
It speaks to me. Substitute MongoDB for PHP or vb.net or w3schools.com in that post, and it is my frustration. With these tools, I feel bad. I feel sad. I feel they are inferior. But I can never find a real, convincing, proper argument. I have to go through great lengths to stay objective, and it's never quite satisfactory.
Why? Because I'm not good at exploring my thoughts purely objectively. I'm human. When I use a tool, I feel passion. Sometimes positive, sometimes negative. When I use Lisp, it burns like fire. But when I try to explain it, I have to admit: there are only downsides the eye can see.
When I use PHP, vb.net, or MongoDB, I'm on Antartica. Without a jacket. It's cold. It's sad. There are no penguins. I feel bad. But I cannot tell you why. I have to admit: they get the job done. Millions use them. The greatest of the greatest sites and applications are built with them. Software that you use every day. Around the world.
So perhaps it's just me. As painful as it is to admit, I simply cannot find the words to match my feelings about this. And I feel alone.
Then I read this. It warms my heart. He knows what I mean. He speaks my language. Had he had a good, objective argument, he would have used it. But he didn't, because he only has bad feelings. It is precisely the fact that he doesn't have anything objectively negative to say, yet feels so badly about it, that speaks to me so much.
MongoDB might be objectively good, but I find it noteworthy that so many people consider it subjectively bad. This is useful information to me.
That's why I voted it up.