Fighting The NoSQL Mindset
yafla.com
yafla.com
Maybe his numbers are off and his technique is wrong, but this article isn't making a one-sized-fits-all mistake.
It's pretty clear that this guy things Digg is a bunch of idiot engineers - I mean, why else allude to "rudimentary comp. sci. knowledge"?
These articles might have a central point, but when it's surrounded by a bunch of opinion, it like watching TV news (e.g., Fox News). The real information gets drowned out.
Digg's description of their entire setup seems a bit unusual -- from how they've defined their tables to their query methods. It seems at least somewhat likely that they were not making optimal use of their technology.
I'd be curious what the traffic/content difference between Digg and Stackoverflow is. Stackoverflow uses an architecture very similar to what Forbes proposes here and they have plenty of capacity on rather unremarkable hardware.
http://www.alexa.com/siteinfo/stackoverflow.com#trafficstats
Edit: To expand I'd imagine SO serves a lot of static "how do I" google hits for logged-out users which is about as easy to serve as it gets. Digg puts a higher emphasis on the logged in experience and commenting which reduces the ability to cache.
First these addresses didn't have a zipcode field, but now they do. What about relationships between records, isn't there going to be redundant data? What if the user changes their email in the account record, do I have to manually update their contact record? This is what foreign keys are for. OMG, Where am I ... Where's the exit? We're all gonna die!
That's basically what I go through every time I think about it. It's a big scary fully loaded gnarly nested hash table aimed directly at my foot.
I'd love to see the technical details of how Digg used MySQL, why it was slow and why NoSQL is better, how they use it. The more of this I see and the less real information the more I think they didn't know what they were doing in the first place.
If you can switch from one to the other and back again, how are they not solving the same problems?
I can store JSON in a MySQL column and you'd say I'm using the wrong tool. But I decided to define a schema instead, then I'm using the right tool? It's pretty arbitrary. Some requirements do lean one way or the other but most of the time there's not much of a difference.
The areas where clearly an RDBMS is the correct solution or where NoSQL is clearly the correct solution aren't an issue. I think there are huge areas that do overlap. NoSQL is being advocated in far more situations than just where it's clearly the correct solution.
Yes they both store "data", but there are types of data that one is well suited to deal with, and alternative types of data that the other one is better for. On my team we've already split out data into two sets, one that lives in ACID RDBMS and one that lives in NoSQL.
If you are dead set in using one solution for everything, my pity for you.
Let's bring up the age-old vehicle analogy. Any vehicle fundamentally solves the the same problem but that's naïve outlook that doesn't acknowledge any specific use-cases. I can drive a truck down the street to get a Coke or ride my bicycle across Canada but that doesn't mean bicycles and trucks are meant to solve the same problem. The differences are obvious when you need to move a pallet of Coke (or 50).
If you have pallet of Coke to move, a bike and no driving licence that you don't care what the bike was meant for.
If you have to hammer a nail but you only have pilers in your close vicinity you hammer it with pilers and don't care what inventors of pilers meant them for.
I'm not trying to draw any parallels to RDBMS vs NoSQL just pointing out that what often matters in real life is what problems are solved with which tools and why. What the tools were meant for is pie in the sky.
It's not the physical world and we basically always have whichever tool we need available to us. I mainly use open source software so that may not be true for everyone I suppose. That's their choice though.
When I test the performance of a query, I query a live mysql node (or at least a live replicated standby) of some data that's not actively being used (to give a realistic "cold-cache" scenario, even though the caches aren't necessarily cold).
If digg used this method, it would completely account for performance discrepancy. Digg did not release a benchmark, and trying to treat their findings as a repeatable benchmark is wrong.
Yet they released record counts, schema, and then performance numbers, and then used their results to demonstrate the failure of the RDBMS (which they led into by saying that it, as some given philosophy, optimizes writes at the cost of reads, hence their poor read performance).
Many of the comments in here are baffling. Digg specifically used the hammer of NoSQL to pound the nail of their database needs, replacing MySQL. They've made a big deal about this. So why the noise about "they're different, man?" And now the petty whines about benchmark methodology when Digg made concrete claims about RDBMS systems?
MySQL does not do certain things well. However many mature RDBMS's have solved these problems and its worth pointing that out before you drop the entire class of tools for the next shiny tech.
Eventually, you learn not to criticize projects that you aren't actively in the trenches with. There's almost always some subtlety you're missing, and the existing team is too busy fixing it to correct your misconception. I bet the Digg team is looking at these comments (well, if they have time) and thinking to themselves, "We tried that a year ago, and it didn't work. If only they knew..."
It's actually pretty depressing to see content like this from an alleged Database expert. I've never been able to get 1:1 results between "lab" tests and actual deployments, and I've gone to much greater lengths than the author to simulate workloads.
That's all in the article.
Now the Joe Plump or whatever guy is telling people that you should sort in PHP. That is the level of expertise of Digg.
I actually flagged the parent article. It's a troll, and not worth anyone's time. RDBMS is good sometimes. NoSQL is good sometimes. Do we really need _another_ holy war?
The specific case -- IIRC -- Digg mentioned was a query which required MySQL to generate a temporary table too large to sit in memory, so it ended up being done on disk. Moving some of the processing out of the DB query and into PHP avoided that, understandably resulting in a huge performance difference (in-memory versus on-disk has a way of doing that...).
Uh, using language like "bottom-feeder RDBMS" and saying that an engineering team is using it "horribly incorrectly on clearly comically deficient hardware" makes this article highly biased, and trollish. It's pretty clear that this guy things Digg is a bunch of idiot engineers - I mean, why else allude to "rudimentary comp. sci. knowledge"? These articles might have a central point, but when it's surrounded by a bunch of opinion, it like watching TV news (e.g., Fox News). The real information gets drowned out.
That's a pretty good reason to flag a story. The entire argument is based on the premise that others are idiots: hubris of which we are often guilty. After you've been in the trenches for a while, you should realize that there is usually a pretty reasonable explanation for seemingly stupid problems.
And a few choice statements from the article:
> I would say Digg's case is an example of a bottom-feeder RDBMS product (apologies for being incendiary, but why does the problem always come down to MySQL? These examples always end up being "we moved from MySQL to NoSQL" rather than "We moved from Sybase ASE to NoSQL"), used arguably suboptimally on unpowered hardware,
> went contrary to the demonstration that even a mediocre machine can beat their results.
> Nonetheless, it is a warning sign of a foundational product issue.
> Decent database products like SQL Server even allow you to include
> So either MySQL is an atrociously bad product at the larger limits, which ample evidence seems to point as a truism,
> Please get away from the compiler and save the world from your monstrosities until you have some knowledge of these basic concepts.
> Alternately you can just clutch onto NoSQL and bleat about how it changes all of the rules anyways, which is the route quite a few have decided to pursue
Okay, I'm done. Point is, dude is straight up trolling about how MySQL sucks. This article does nothing but fuel the fire of yet another flamewar, and so it gets flagged. I'd like discussions to remain sane around here.
Maybe MySQL really isn't a decent database product, or not for high performance needs. How many of his statements are troll-ish in that light?
The OP repeated this reasoning more than once.
He showed that the facts cited by Digg do not make sense unless we take into account poor database technology, poor database configuration, or poor database skills, or all three.
Perhaps you disagree with that. However, that's not what you said above. The OP also discussed the nature of this micro-benchmark, and it's relevance despite his own poor knowledge of the actual data characteristics.
So in other words, he has already directly addressed your concern in advance, more than once on that too as a matter of fact. Considering that fact, you haven't actually responded to his article, you just wrote a "tends to" point about micro-benchmarks, I think it's pretty clear that Dennis Forbes knows a thing or two about benchmarks.
Blah blah blah to empty air, this comment page is pretty much a fact-free and nuance-free flame war anyway, so what's the point, sorta embarrassing for the esteemed HN crowd.
I am not a database guy. I also don't know enough about Digg's set up to say with authority if these comments make sense. So I specifically wrote "I feel," "tends to," and "imply" because I'm not comfortable making an absolute statement about the issue.
However... I don't see how this test is in any way relevant. a 30GB database? Running on totally different hardware?
In any case, re-reading the article again, I see that relevance paragraph now. I guess I missed it the first time around between all of the flaming, trollish comments about both NoSQL and MySQL. But I still don't see how we can extrapolate this test in any way to imply anything about Digg's practices at all. Then again, it's 8:30am.
What a unclever excuse for missing the point.
It still doesn't change my original point, however. Just because he acknowledges that the benchmark is unrelated to what he's talking about doesn't excuse him from the fact that it's unrelated to what he's talking about.
It also doesn't change the fact that the article is still a troll, regardless of the correctness of his benchmark.
I'd rather be able to remove barriers like having to design a schema, and get some early efficiencies to develop my app fast and iterate.
Though, my experience in scaling every site that needed to be scaled has concluded with sharding. So MongoDB sort of fits there as well with its autosharding capabilities.
That said, Cassandra won't always fit my needs. Sometimes I really do need sophisticated queries and arbitrary transactions. I've been enjoying this guy's articles, because he's bringing up a lot of good performance tips for those times when a relational database is the right tool for the job.
That's because there's a zillion MySQL installs out there, with users that talk about them. On the contrary, there are a lot less Sybase installs out there and their (corporate) users don't talk about them. Go figure that you only hear about MySQL. But please, keep on spreading the FUD; that just gives us the edge of using a free, OSS, system.
DISCLAIMER: This is not a high-fidelity reproduction of Digg's situation
And it's probably not even a low-fidelity reproduction. The article gives us no reason to suppose he actually knew or understood the problem Digg had. He just shows that it was not a trivial one, as that would've been easy to solve.
sed s/NoSQL/NoMySQL/
and avoid the confusion.
Need more be said? Seriously?
MySQL is slow on writes because you'll get a random write for every index you maintain (+the table it's self). MySQL's replication will help scale reads but does nothing for writes. Cassandra is actually said to be slower on reads and it's thought people will already be using memcache so it's not a problem.
The article doesn't seem to have any writes going on while he is reading.
Bath water and baby gone without a second thought.
Not to mention that later the author says:
> SSDs change everything.
To which I say: "Or you can clutch onto SSDs and bleat about how they change all of the rules anyway, which is the route this author has decided to pursue."
I am dipping my toes into couchdb, just to see what all the noise is about.
Still getting my head around map,reduce and the fact im writing in javascript. All that aside the biggest exciting factor for me.
It is so much easier to write custom functions for it than SQL. ( Mysql, and yes I only tried via phpmyadmin )
Uh, he said "fan of SQL Server", which is Microsoft's SQL product. Fans of SQL Server tend not to be fans of mySQL at all, since SQL Server makes mySQL look like a "bottom-feeder" in the original poster's words. Not that I disagree.
Now run the same queries you were doing again on your test machine, but simulating 50K users online at once. Oh, and don't forget about thousands of writes per second, which was conveniently not part of this test. What's that you say? The performance is suddenly complete shit? Color me shocked.
I'm not convinced that read performance has to suffer for writes. Reads don't have to block writes or block for writes - use the NOLOCK hint.