The OP repeated this reasoning more than once.
He showed that the facts cited by Digg do not make sense unless we take into account poor database technology, poor database configuration, or poor database skills, or all three.
Perhaps you disagree with that. However, that's not what you said above. The OP also discussed the nature of this micro-benchmark, and it's relevance despite his own poor knowledge of the actual data characteristics.
So in other words, he has already directly addressed your concern in advance, more than once on that too as a matter of fact. Considering that fact, you haven't actually responded to his article, you just wrote a "tends to" point about micro-benchmarks, I think it's pretty clear that Dennis Forbes knows a thing or two about benchmarks.
Blah blah blah to empty air, this comment page is pretty much a fact-free and nuance-free flame war anyway, so what's the point, sorta embarrassing for the esteemed HN crowd.
I am not a database guy. I also don't know enough about Digg's set up to say with authority if these comments make sense. So I specifically wrote "I feel," "tends to," and "imply" because I'm not comfortable making an absolute statement about the issue.
However... I don't see how this test is in any way relevant. a 30GB database? Running on totally different hardware?
In any case, re-reading the article again, I see that relevance paragraph now. I guess I missed it the first time around between all of the flaming, trollish comments about both NoSQL and MySQL. But I still don't see how we can extrapolate this test in any way to imply anything about Digg's practices at all. Then again, it's 8:30am.
What a unclever excuse for missing the point.
It still doesn't change my original point, however. Just because he acknowledges that the benchmark is unrelated to what he's talking about doesn't excuse him from the fact that it's unrelated to what he's talking about.
It also doesn't change the fact that the article is still a troll, regardless of the correctness of his benchmark.
When I test the performance of a query, I query a live mysql node (or at least a live replicated standby) of some data that's not actively being used (to give a realistic "cold-cache" scenario, even though the caches aren't necessarily cold).
If digg used this method, it would completely account for performance discrepancy. Digg did not release a benchmark, and trying to treat their findings as a repeatable benchmark is wrong.
Yet they released record counts, schema, and then performance numbers, and then used their results to demonstrate the failure of the RDBMS (which they led into by saying that it, as some given philosophy, optimizes writes at the cost of reads, hence their poor read performance).
Many of the comments in here are baffling. Digg specifically used the hammer of NoSQL to pound the nail of their database needs, replacing MySQL. They've made a big deal about this. So why the noise about "they're different, man?" And now the petty whines about benchmark methodology when Digg made concrete claims about RDBMS systems?
Eventually, you learn not to criticize projects that you aren't actively in the trenches with. There's almost always some subtlety you're missing, and the existing team is too busy fixing it to correct your misconception. I bet the Digg team is looking at these comments (well, if they have time) and thinking to themselves, "We tried that a year ago, and it didn't work. If only they knew..."
MySQL does not do certain things well. However many mature RDBMS's have solved these problems and its worth pointing that out before you drop the entire class of tools for the next shiny tech.
Now the Joe Plump or whatever guy is telling people that you should sort in PHP. That is the level of expertise of Digg.
I actually flagged the parent article. It's a troll, and not worth anyone's time. RDBMS is good sometimes. NoSQL is good sometimes. Do we really need _another_ holy war?
Uh, using language like "bottom-feeder RDBMS" and saying that an engineering team is using it "horribly incorrectly on clearly comically deficient hardware" makes this article highly biased, and trollish. It's pretty clear that this guy things Digg is a bunch of idiot engineers - I mean, why else allude to "rudimentary comp. sci. knowledge"? These articles might have a central point, but when it's surrounded by a bunch of opinion, it like watching TV news (e.g., Fox News). The real information gets drowned out.
That's a pretty good reason to flag a story. The entire argument is based on the premise that others are idiots: hubris of which we are often guilty. After you've been in the trenches for a while, you should realize that there is usually a pretty reasonable explanation for seemingly stupid problems.
And a few choice statements from the article:
> I would say Digg's case is an example of a bottom-feeder RDBMS product (apologies for being incendiary, but why does the problem always come down to MySQL? These examples always end up being "we moved from MySQL to NoSQL" rather than "We moved from Sybase ASE to NoSQL"), used arguably suboptimally on unpowered hardware,
> went contrary to the demonstration that even a mediocre machine can beat their results.
> Nonetheless, it is a warning sign of a foundational product issue.
> Decent database products like SQL Server even allow you to include
> So either MySQL is an atrociously bad product at the larger limits, which ample evidence seems to point as a truism,
> Please get away from the compiler and save the world from your monstrosities until you have some knowledge of these basic concepts.
> Alternately you can just clutch onto NoSQL and bleat about how it changes all of the rules anyways, which is the route quite a few have decided to pursue
Okay, I'm done. Point is, dude is straight up trolling about how MySQL sucks. This article does nothing but fuel the fire of yet another flamewar, and so it gets flagged. I'd like discussions to remain sane around here.
Maybe MySQL really isn't a decent database product, or not for high performance needs. How many of his statements are troll-ish in that light?
The specific case -- IIRC -- Digg mentioned was a query which required MySQL to generate a temporary table too large to sit in memory, so it ended up being done on disk. Moving some of the processing out of the DB query and into PHP avoided that, understandably resulting in a huge performance difference (in-memory versus on-disk has a way of doing that...).
It's actually pretty depressing to see content like this from an alleged Database expert. I've never been able to get 1:1 results between "lab" tests and actual deployments, and I've gone to much greater lengths than the author to simulate workloads.
That's all in the article.