Fighting Back After Hacker News Took Down My Site
plainlystated.com
plainlystated.com
I don't mean to level criticism solely against the post author for this, because he is not alone, but I find this mentality inexcusable. This is especially true if you're using WordPress, for which there are a multitude of caching plug-ins. All of which will offer orders of magnitude better performance than allowing the blogging engine to build the page from scratch every time.
In other words, caching is always a priority. It's not an add-on or an after-thought, it should be part of your design.
Consider a scenario where someone asks you to add up three arbitrary numbers:
123 + 456 + 789
Now, imagine you type those in to a calculator to add them up. It takes you a few seconds. The output is 1368. You can easily remember this number, so the next time someone asks you what the result of 123 + 456 + 789 is, you can just say 1368. Not caching is like keying the numbers in to a calculator every time, rather than just relying on your memory.
I know this is a rudimentary example, and I know that most people running a WP blog probably know how caching works, but even if you don't "need" it today, why would you leave your blog set up so that it's constantly re-building content that can be cached by simply installing a WP plug-in?
I implore you. Make caching the second thing you do (after security) when setting up or building your web app.
However, lots of people get confused by caching plugins when they're editing templates and things.
Idea: Wordpress caches pages. It checks to see when its template files were last touched. If they were touched after the cache, it regenerates. Otherwise, the cache is served.
You can have that for free, Wordpress.
If I just want to set up a personal blog that I don't plan on promoting or spending a lot of time on I wouldn't want to spend any extra time setting up things that don't help me with my primary goal, writing blog posts.
If you're not using WP, you might use something like Tumblr or Posterous. In that case, caching isn't your problem. If you rolled your own blogging software, well, you already violated the principle that blogging is your primary goal.
It also increases the barrier of entry to posting something, so I'm less likely to post something inadequate.
I find that mentality kinda odd.
Plus, I'd grown lazy from my lack of traffic :)
http://www.viper007bond.com/2010/08/10/why-wordpress-doesnt-...
I will also quote myself: "I first learned to disable Apache KeepAlives in 1998. Yes, 1998. It's disheartening that Apache still ships with it enabled by default. It has always allowed a relatively small number of lingering clients to completely DoS your server."
Key points:
• Make sure that wordpress supercaching is on. You can verify this by looking at the last line of the HTML that comes back which has a timestamp for when it was generated.
• Turn off KeepAlives.
• Set MaxClients to 8.
• Use monit to check to make sure it can connect and restart apache with a kill -9 if it can't. (This is optional, but helps if you have some random thing that ends up taking a very long time to execute and eats up connections.)
With that in place even a much smaller server can easily handle a HN top story without breaking a sweat.
LOL what? what's your mpm? did you enable this to keep the backend from blowing up from too many queries? surely you can handle more than 8 connections at a time. is this the proxy layer, and if so are you using web caching on top of wordpress caching?
if monit is restarting apache every time it can't connect (i hope you have a long timeout) you're denying service to a lot of people. connections are supposed to queue so they don't get dropped.
My point, specifically, was how low you can go with a cheapo VPS. We hold up fine during an HN spike. (We're B2B and not a destination site, so our usual load is trivial.) Even during an HN spike you're getting tops of 2-3 visitors per second, which can be dished out reasonably well with 8 workers.
The monit thing kicks in after a 30 second timeout. With the configuration above, we don't get that because of load, but rather when something else has gone wrong (specifically there's a wordpress plugin that our internal status blog uses that sometimes hangs). But given the original poster's issue of apache getting so out of control that it took him several minutes to get a live SSH connection and a system load of 60, having monit kill things (and restart them) is a preferable stop-gap.
(Note: Our actual customer facing stuff is quite different; there we're using multiple servers behind an nginx proxy and using a combination of Rails, Sinatra and Java services. The basic web stuff is segregated off from those primarily for security reasons.)
If you have some free time try deploying your LAMP stack with Buildroot and uClibc. The application size ends up being around an order of magnitude smaller, but i've never bothered checking if private RSS on clunky apps like PHP or Perl is minimized at all.
edit Also for prefork we used to have some scripts that would monitor processes to see if they 'went crazy' and wouldn't ever return, and reap those processes so Apache would fork a new one so we didn't have to restart the whole server. When you're under peak load and you restart a server and it sends all those clients to all your other servers which are already almost at their peak things get very nasty very quickly. I know you only have the one VPS here but for applications with many servers it can be handy.
HTTP keep-alive connections are faster (saves 3-way handshake on subsequent requests and allows pipelining).
The problem is that Apache implements them in a horribly inefficient manner (thread/process and all its memory kept in use just to hold on to a socket).
If you use nginx, lighttpd or shield Apache with haproxy/varnish, then you can easily have keep-alive enabled and clients will see better performance.
nginx: rails & sinatra are on passenger. php is php-fpm. nodejs is proxied
Just look at the query count even without plugins.
Add a few plugins and it's a total mess.
Seriously, there's no reason to stay on a traditional style server setup and you'll never go back once you've tried it. It can be pricey once you ramp up, but give it a shot as you can literally pay by the hour while you play.
Still, I think there's value in figuring out this stuff by investigating efficiency gains. If I was having to do real tricky stuff (which would probably be beyond my expertise) then the cloud would make more sense for me.
P.S. I upvoted you. I also don't get why people downvoted you.