When I test the performance of a query, I query a live mysql node (or at least a live replicated standby) of some data that's not actively being used (to give a realistic "cold-cache" scenario, even though the caches aren't necessarily cold).
If digg used this method, it would completely account for performance discrepancy. Digg did not release a benchmark, and trying to treat their findings as a repeatable benchmark is wrong.