Increased traffic? I don't think I understand the question
:|
I was trying to determine what component broke first.
An app like reddit has most of its backend tweaked or rewritten continuously as it becomes more important to find and alleviate some new bottleneck. If you do it right, you predict the next bottleneck before it becomes one but there will always be something.
Maybe writes take too long. Why? You've capped the performance of the disc in your DB master (or more importantly of the most expensive single set of discs you can afford). Maybe adding app servers proportional to site-wide traffic stopped helping you keep up at some point. Turns out you're network bandwidth bound. Maybe requests are too slow, but apps aren't using much CPU. Why? Because you're network-latency bound. Hey, it turns out that some tight loop was doing single memache GETs when it could have been doing a single multi get.
That last one is just a general performance bug, but that's just it. Every bottleneck that keeps you from just adding resources is a performance bug. Assuming your app is not infinitely fast, there's always something.
http://blog.reddit.com/2010/01/why-did-we-take-reddit-down-f...
We were already using memcachedb as a persistent keystore, "caching" precomputed listings and comment threads. When it started becoming more trouble than it was worth to maintain, we decided to try cassandra as a beefier replacement.