Where did Foursquare find their engineers? I hope no one lost their job here but this is pretty elementary stuff.
Where did Foursquare find their engineers? I hope no one lost their job here but this is pretty elementary stuff.
Further, there are lots and lots of different things that we need to be monitoring at any different time to make sure that everything is going ok and we aren't about to run into a wall. Automated tools can help a lot with this, but these tools still need to be properly set up and maintained.
I'm not saying we didn't screw up. We had 17 hours of downtime over two days. We screwed up bad, and we feel horrible about it, and are doing a lot to make sure that we don't screw up the same way again.
But it's not because we're morons that never thought about the fact that we should be monitoring memory usage. We just got overwhelmed with the complexity of all that we're doing at once.
-harryh, foursquare eng lead
In additions to hiring scalability experts (many will claim to be experts but most aren't), talk to your investors to find outside technical advisors who have worked on and are still working on similar problems.
Find advisors from Google, Facebook, and other places who have dealt with these issues. In my experience, you have as much to learn from people who have made mistakes as from people who have succeeded.
Also, in my judgement, you guys are a little too risk-tolerant for your significance (first major Scala 2.8 upgrade, largest MongoDB deploy...).
It's not like your 66GB is hot all the time (or even big by any sort of measure), so that makes no sense to me, and that part of things needs to be further explained by 10gen and/or 4sq.
That said I'm looking at monit too. I hear it's quite nice and has less of a learning curve.
Get that tool in place NOW while you are small, so it's there when you get bigger... so you can hire people to watch it while you move on to other things.
You will always have a lack of time, and always be busy. Spend a weekend putting up Nagios, setting up alerts over SMS or XMPP or Twitter or whatever you want, and then move on. This is fundamental to systems management.
Dont' over think it - just get nagios/cacti (again, try groundwork) up and running and graph EVERYTHING you can... yuo'll be very, very glad you did.
We can help you all with these problems.
Here's how fast/easy it is:
1. Create an account (~30 sec)
2. Add your cloud credentials (~45 sec)
3. Install the monitoring agent (~90 sec/node)
4. Create a CPU/Memory/Disk monitor, query-targeted all your servers (~60 sec) (example: "provider:EC2")
5. You get an email whenever the monitors you created reach the thresholds you set
There are a lot of other cool features - check 'em out on the site: www.cloudkick.com
All our plans are free for 30 days.
Sharding needs to be managed in a way that evenly distributes the data. I would never do it by user ID unless a proper analysis of the data showed it to be a fairly even distribution.
That doesn't excuse anything, but these oversights can happen even when you've got primo talent on board.