Step 1. Measure, measure, measure: response time, page load time, query time
Step 2. Identify bottlenecks: Why isn't your page load snappy? Too much JS? Round-trip to DB is killing you? Your mongrels are overloaded?
Tip: Use YSlow for frontend measurements and take most of its advice
Step 3. Fix bottlenecks: Lots of techniques to fix faster frontend performance, start with YSlow above. Architectural and backend issues can be solved with a combination of hardware (load balancers, more web boxes, more db boxes) and software (better caching strategies, evolve your DB architecture, break up your application into smaller micro-apps that scale independently)
Step 4: Go to step 1 and keep at it.
I'd have to say that if you're not doing Step 1, you're going to fail now, or eventually. As one of my first mentors taught me:
"You cannot improve, what you don't measure"
1. caching is your friend
2. you probably need 2 load balancer (instead of 1) for high availability
3. probably need a SAN or shared storage somewhere when you have several web servers
Split static content serving from dynamic. Can start off on same server but requirements for each are so different it pays to have tuned processes serving each.
Build customized, svelt apache or use nginx.
Get separate DB server, get lots mem for DB server, tune DB to use that memory. Few/None are setup up out of box to use 4-8GB+ well.
Get more mem for http box(es), use memcache/similar to take advantage of that memory.
I only point this out, as load balancers and additional servers can be expensive.
About 20-30% of stuff posted there is worth reading, which is pretty good =)
The bottlenecks are very rarely where you think they are. Even when they are, there is typically lots of low hanging fruit hanging around that you don't know about until you profile.
Also, you should have some sort of monitoring software running aggregating stuff like server loads, connection stats, number of db requests, memcached hits/misses, etc. Cacti I find quite good for this.
imho you need to design with horizontal scaling in mind
Given 8 x 3Ghz core / 32GB / 5TB machines for cheap you're packing as much power as 15+ servers of a few years ago. Even quite large web sites can be run off a few beefy (but cheap!) servers. This means money saved on network infrastructure, power, space, management, etc.
Probably 80% of the Alexa top 1000 could each be run off less than 20 modern beefy machines.
-Phil Karlton
Second step is usually cache - local, in-application cache. It's pretty much auto-pilot now, writing a cache function.
Cache helps, but not as much as you'd think. When it gets serious though it's still useful because its shape makes it easy to go to step three: mix the cache with a bit of refactoring. This can mean simply pre-fetching: there is a huge difference from doing 20 "select * from order_items where id=xxx" and just one "select * from order_items where order=yyy". And with the cache in place it usually means only a few extra lines of code. (This example is for building reports, not for displaying orders. I'm not that of an idiot to get each order_item separately).
If things get heavy refactoring can go as far as breaking the abstractions in place. Last month I squeezed an extra second by sending to the browser a big table in batches, while I was computing it. Not something I want to do every day, but occasionally it's fun.
This is also the step when I usually realize I made a stupid mistake along the way which caused the delay.
The only problem I met which was mostly impervious to this was a design decision. In one app I decided (as I usually do) to compute stuff like balances and debts on-the-fly. It usually (well, always) is a good idea, because any intermediary result is a source for bugs. But this particular app grew rather unexpectedly large, with the result that computing an overall balance takes 3-10 seconds. It's not a huge problem for the client, but my pride suffers. I'm still looking into this one - I think some refactoring may make it ok.
Without traffic, there's no point in scaling.
Here's something I recently wrote about speed and web apps: http://bens.me.uk/2009/fast-web-applications
Check out:
I also tried to read about in the first chapter of the upcoming CouchDB book, but it only contains vague generalizations like "Erlang makes you scalable automatically" or "it is built on the same technology (HTTP) as the web which is proved to be scalable".
I am still not convinced.
1. Host all your static files (images,css,js) on Google App Engine
2. Avoid hitting your database. Cache everything possible.
At the very least, make sure each system bus (i.e. PCIe/PCI-X for commodity x86) slot is populated with a host adapter (ideally a hardware RAID controller) which can saturate that slot.
For each host adapter, ensure that each port has enough disks behind it that, even with worst-case contention (not fully-random, unless you've somehow managed to magically induce 512-byte transfers), they can saturate the the upstream system bus, even after RAID processing is done by the adapter.
After that, upgrade to a server with more slots and repeat the saturation.
Repeat once more for the largest single server available, while starting the project to implement (perhaps write from scratch) a distributed database.
I have yet to see anyone make even a concerted effort on the first step, instead skipping to the second part of the last step, probably because that appears, at first glance, to require only a Small Matter Of Programming.
sure, one day you'll have to build a nuclear reactor and tap a dam to provide all the power and cooling for the datacenter space, but it's easier than a redesign.