MongoDB as a better default data store
blog.gregweber.info
blog.gregweber.info
> * relational tables
> * transactions
Shudder... I think I'll hold on to my relational tables and transactions as the default core of my systems, and optimize when I see that they're not fast enough, thank you very much.
1) Your tests have to be document aware. Models that are embedded within a document are very difficult to test in isolation from the rest of the document. You need to build the whole document before you can test any persistence related functionality.
2) Refactoring relationships is difficult. When you start building an app you don't know how the model is going to turn out. With a relational model that's fine, you can add/remove relationships easily. With documents your models are organised into hierarchies that they have to be aware of. The code that handles this hierarchy is incredibly brittle. If you move a model from one document to another, or just it's own document, you're going to have lots of code to fix. Maybe it's a case of needing some new refactoring patterns, but rest assured, they'll be far more complicated than refactoring patterns for relational stores.
3) Mongoid and the uncanny valley. It looks like activerecord, but I spent an inordinate amount of time working around bugs and methods that you expected to be there that just weren't. Admittedly this was before the stable version was released(the recommended version at the time) so things may look significantly better now.
Persistence (which is available in MongoDB) has nothing to do with testing of a model (a single document inside a collection). Some actually see that as an advantage of using a schemaloose (yes I just made that up) storage like MongoDb. I don't need to have a whole document defined in order to test the functionality of an embeded object in that document. Yes, you're relying on the application alone to enforce the model, but that's a trade off made up front regardless.
As for modeling and the cost of refactoring, I think it depends on more than just the storage facility you're using. The code that handles your modeling in MongoDB is only as brittle or robust as you choose, same as a SQL solution. Yes, it helps if you think more up front about how you are going to query the data that you store using MongoDB as this will help to minimizing refactoring later on. However, I think the pain associated with moving part of a table to another / new table in SQL is no less painful than moving an embedded object or other data structure around in MongoDB. In both cases you're moving data and affecting changes in your application.
I think 3) is better now that version 2.0 of Mongoid has been released (and is stable). The developer I am working with complained about earlier versions also. I ported hundreds of lines of ActiveRecord model code, and it mostly just worked.
I am not really experiencing pain around 1) either. I am using factory girl and find it easy to build any required data. It could actually be beneficial that you must focus on creating a group of properly linked objects. I have run into trouble before with messed up/difficult to create object graphs in ActiveRecord tests.
2) Is a good point I am now mentioning in the post. I will let others argue with you about its severity :)
I should say I think mongodb/mongoid can be very useful. My opinion so far though is that it's most effective as a performance optimisation strategy for a stable application.
And similarly - most of the data I work with is naturally relational. Having the database make checks to ensure that invalid things aren't happening is a lifesaver against coders making a stupid mistake (and everyone makes one from time to time).
------------------------------
mongo now has a 'safe mode' option that protects against this, just FYI
> An application that frequently times out (page loads over
> 30 seconds)
An application like this is a good candidate for review what the hell is going on, before making any decisions about changing the stack.Go back to school, or hire capable developers.
And an article followup touching on the the basis of his tweet: http://metaduck.com/post/3564002672/asynchronous-iteration-p...
Per collection locking is coming soon as well https://jira.mongodb.org/browse/SERVER-1240
But if per document lock could be available, then mongodb performance would be much more impressive. Though as far as i understood it's not possible with memory mapped files for some reason.
Yes, with the small caveat that Mongo doesn't respond to any queries[1] while you do that. So better don't try such a bulk operation on a production Mongo.
[1] http://2.bp.blogspot.com/_VHQJkYQ5-dY/TUO3RAn8SNI/AAAAAAAABq...
Edit: Not sure why it's getting downvoted. You can reproduce it fairly easily by performing bulk writes - esp. deletes, but also mongoimport or just a write-loop. The global lock is a known problem.
Of course the original author seems to think that most relational databases lack performance and / or features. There is no one size fits all situation. This is why companies like facebook, netflix, etc. use combinations of relational databases and nosql databases (although i think netflix is almost completely over to simpleDb now...)
In my experience there are always ways to tweak something in your RDBMS setup to improve performance, like tuning queries or your indexing strategy, or skipping using an ORM for super-critical stuff, or looking at I/O latency or blocking issues between your app and DB machines.
I personally don't have deep operations experience, and I am sure your individual experience is valid.
I think that in part, MongoDB is maturing- it now has single server durability and replica sets. Single server durability means it now can be a default, whereas before you had to commit to 2 servers as you allude to.
People seem to report it is more critical to have enough memory, and that it wants more in comparison that a SQL database, so perhaps that limits its ability to be a default database.
The issue of course isn't whether you can adjust an RDMS based infrastructure to suit your needs, but how much effort that will be in comparison to an alternative.
In a database with a schema, if you saved a string to an integer field, the database could either complain or coerce the string to a integer. Either way, whenever you read a row, you know what type the value will be whenever you read data a row from the database. Another issue with schemaless is simply mis-typing the key (column) name. If you have a field called referrer, but you perform a query using referer, MongoDB won’t know you have done anything wrong. In a SQL database you will get an error on insert and on querying. An ORM can instead be the one to tell you about the error.
The MongoDB docs do say that I am supposed to use the term "ODM". I will update my post. thanks.