The Slashdot Effect from the Other Side
cmdrtaco.net
cmdrtaco.net
We used to ask people all the time if the would send us their sanitized logs so we could see the "reddit effect", or at least give us the aggregate data.
I still get a lot of traffic from r/food, but it was relatively miniscule compared to the traffic I got from a post in r/reddit.com (which no longer has any incoming traffic).
I should see if I can get one of my posts to the top of reddit. So far all my best submissions were on the reddit blog.
Global news sites can experience far bigger surges when a major story breaks. Around 10 years ago CNN started switching out its regular home page for a version that stripped out video and most images when major stories broke. Not sure if CDNs and better Internet video technologies have reduced the need for this ...
or in other words after 9/11 melted servers everywhere.
That is one of the coolest feelings in the entire world for a little nerd like myself. Usually
tail -f /var/log/apache2/access log
is running on one of my monitors all the time, so I can pretty much watch visitors as they come into the site (I don't even see the code, I just see blonde, brunette, redhead). When I've gotten that sort of traffic, I can't make sense anymore. It's the difference between taking a drink from a fountain, and sticking your face down a geyser.Not only that, but to then watch your code stand up to the traffic is cool :).
Like a webdev version of your first kiss, I suppose :)
"Hello, Foobar ISP, can I help you?"
"Hi, I'm the reason your servers are melting down right now."
"Ah. Let me connect you to the company president..."
Told him the detailed story, he got a kick out of it, and deleted the attracting file.
([1] - AP took issue with what I thought "fair use" of the Elian Gonzales pictures the day they were published. No argument; I didn't expect concocting them into an animated time-lapse sequence would get _that_ popular.)
If you read the EC2 forums[1] for any amount of time you get used to seeing post after post of "hung sites" running on the Micros when a constant level of demand is made on them (not the intended usage model[2])
Under load the Micros frequently go into a catatonic state with %st ('top') climbing to the moon as the VM environment provisions the brunt of the server's resources to the other VMs on the machine and starves out any hungry Micro instances (as designed).
I didn't think you could host anything on them with regularity. Maybe a swarm of them behind an ELB, not not a single one... anyone else had the same experience that Malda had?
[1] https://forums.aws.amazon.com/thread.jspa?threadID=58323
[2] http://docs.amazonwebservices.com/AWSEC2/latest/UserGuide/in...
Also "tens of thousands" of requests is not really a big number anymore. Nginx will happily serve a couple thousand per second even from a micro (until it gets choked by the EC2 resource beancounting).
Are you hosting dynamic content (PHP, Java, etc.) or is everything really heavily cached in static files/varnish and hosted out by nginx quickly so as to avoid saturating CPU or I/O?
There is also a wordpress site and another database-request-heavy site with 500 requests a day.
The problems I had was that the access.log and error.log quickly ate up the 5GB space of the micro instance :-) Causing mysql to stop operating and me to scratch my head for a few minutes.
5GB is not a lot .. you want ruby1.9 with the default setup? That will be 180MB, and so on :-)
400 * 1 million / 1024 / 1024 = 381.5 GB
381.5 * .14 = $53.41
Disclaimer: I haven't used them, but I've heard good things about the service.[1]
[0] https://www.cloudflare.com/
[1] http://news.ycombinator.com/item?id=2631019 http://news.ycombinator.com/item?id=2561341
For a localized readership, it is fine since all the data will be coming out of a single regional location. If you have global readership and need locality, that is where CloudFront comes in handy with 20 POPs around the globe that data is distributed to.
Serving off of S3 is not that expensive, even if your site goes gangbusters.
Serving off of CloudFront can surprise you because of the sometimes unexpected number of origin pulls that can occur as your data is expired from the edge locations.
You are sharing cache space on each edge node with every other CloudFront user, if your content is red hot it stays in the cache, but if it is low-volume (which is a relative term to the other traffic coming out of the node) it gets expired much faster, sometimes hours so any future hits for it will pull (download) it again from S3 to CloudFront, then back out from CloudFront to the client.
The performance implications aren't horrific (they can be for video) but the cost can double what you are paying with enough origin pulls occuring requiring redownload over and over and over again for the same files from the edge locations.
I have seen this catch a handful of people offguard to the tune of $100s of dollars or thousands on the forums over the last few years because they didn't realize this could happen... they just looked at the bandwidth rates on the site, multiplied by their payload sizes and thought that was the fixed rate.
My cheap webhost managed to handle it pretty well. No issues. Then again it was static HTML.
I just read taco's site. He doesn't mention what CMS he is using, but its Wordpress. I wonder what cache plugin he's using. That's pretty important stuff, just as important as what instance of EC2 he is using.
I will put up some stats later, but HN brought in 10,000 hits, Slashdot brought in about 600 people. I think the English vs. Japanese thing scared away quite a few people (ironic considering the content of the article..)
(As an odd note, I did panic after hearing about Slashdot but not after seeing people from HN - go figure)
/. still relevant..!