You Only Wish MongoDB Wasn't Relational
seanhess.github.com
seanhess.github.com
Denormalization is still important. Taking that comment example a bit further, if you need user records tied to the comments, you're going to want to include all the user information that you need for displaying the comment in the comment record itself. Mongodb doesn't have joins. So if you have 200 comments on a post by different people, you're going to have to grab those users from mongo in a separate call, and if you're sloppy about it, 200 separate calls.
If your users really need to change their user names, you might run a job that updates the comments async.
I use MongoDB for my startup (yabblr.com) and have quite a relationship data stored in my entities.
For example, my user entites have an array of commentId's stored inside of them.
Whenever I fetch 1 or 1000 users from the database, a framework that i designed will look at the entire list of users and build a single list of commentIds based on that list.
Then I do one more lookup on the database to find all comments in my new list of comments.
Then a final pass over the users looks up the commentId associations and inserts the actual comment object into the user object.
Finally the list is passed back to the consumer. Its 2 database calls. It does require that there are 2 iterations over the first call (one to gather up comment Ids and one to associate comments with users) but since its being done in the application layer, its much easier to scale.
That way you could easily do db.posts.$comments.find({"name": username})
to retrieve all posts by a user, or
db.posts.$comments.find().sort({date:-1}).limit(10)
to retrieve the latest 10 comments on the blog.
Link to the feature in MongoDB's issue tracker: https://jira.mongodb.org/browse/SERVER-142
Any possible work-around proved to be complicated and tedious--likely requiring heavy application-side logic--if not inadequate. Going hybrid was an option (RDBMS for critical parts, and NoSQL for other non-critical parts), but that may not have been a prudent way to start building this thing, since it would be difficult to predict the portions that might be correlated or codependent in the future (where a JOIN between the RDBMS and the NoSQL DB would be inefficient, if not impossible).
As a payments platform, we cannot afford to lose any transactions, and every penny must be accounted for, so a transactional and ACID-compliant data storage is a must. Thus, we decided to go with a traditional RDBMS. With that said, and looking towards the future, I am considering the use of an ultra-high throughput RDBMS such as VoltDB for our system... A VoltDB vs. MongoDB (or any other NoSQL DB) would be interesting.
So at the end of the day (yes end of the day processing), ACID is needed.
Well, I was referring to payments/transactions, not banking. As I specifically said, we are building a third party payments aggregator (TPPA). So, we handle transactions; we're not a bank.
In our case, there are many scenarios where ACID is needed. A couple of the most basic are the need to rollback a transaction (e.g. if it was cancelled, or if it did not succeed) and the need to recover from a failure (specifically, to get back to a consistent state to resume a failed transaction).
Design your database (or organically grow its needs), and you'll wind up breaking things apart appropriately.
I'm pretending there's two types of data - the "nexis" of the star (comments), which needs a fast DB, and the leaves (users), which are small enough to be cached.
http://blog.fiesta.cc/post/11319522700/walkthrough-mongodb-d...
sql databases scale pretty well too.
http://www.mongodb.org/display/DOCS/Retrieving+a+Subset+of+F...
> t.find({}, {'x.y':1}) { "_id" : ObjectId("4c23f0486dad1c3 a68457d20"), "x" : { "y" : 1 } }
And the mongo docs even have an example pretty detailing what this page wants (paging of comments)
db.posts.find({}, {comments:{$slice: -5}}) // last 5 comments
So, what am I missing?
{ "_id" : "one", "comments" : [ { "name" : "a", "asdf" : "1" }, { "name" : "b", "asdf" : "2" } ] }
The following query returns the comment "a", not "b"
> db.posts.find({"comments.name":"b"}, {comments: {$slice: 1}})
So, you're right, for paging, $slice works fine. For more advanced queries (ranges, the comments by user example), you're stuck. It sounds like 10gen is working to correct this though.
At least they have tickets open for it.
You sometimes want to just get the blog post object too, not all the comments. It's worth it to have to do two queries and get the extra selection power vs one round trip to the DB, but you have to potentially sift through a lot of data to get what you want (or get back a lot of data you don't need).
This is, in my opinion, the perfect mix of the relational/document schemes. You can nest things when they don't grow continuously, but you still get most of the querying power of a relational database.
I feel this should be clarified a bit (since you most likely have a more complicated scenario in mind) but it is entirely possible to only select certain fields to be returned from a query. So if a 'comments' document was inlined within a 'post' document, we could easily opt to only have the latter returned by mongod. Such a query might look like this:
db.posts.findOne( {'_id' : <...>}, { comments:0} );
This would definitely cut down on the amount of data transferred over the wire, but I am not sure if there is any performance hit within mongod.http://www.mongodb.org/display/DOCS/Advanced+Queries#Advance...
The syntax is a bit wonky, but I believe it does exactly what the author says is impossible.
It becomes an issue when you have extremely large arrays in the document (e.g. thousands of comments).
Is there some sort of repair process in mongo that prevents destructive updates in the case of 2 concurrent writes?
However, they do support partial updating. So you only need to write the specific part of the document that you are actually changing.
I wonder if the NoSQL crowd is affiliated with the REST crowd. :)
You usually just make an object in whatever format your app wants, save it, and you can write query against it.
Which isn't to say there aren't things that Pg is better for. Full text search comes to mind.