So, that was a bummer
blog.foursquare.com
blog.foursquare.com
As a next step, we introduced a new shard, intending to move some of the data from the overloaded shard to this new one.
We wanted to move this data in the background while the site remained up. For reasons that are not entirely clear to us right now, though, the addition of this shard caused the entire site to go down.
To those who use MongoDB - does this sound like something that might have been caused by MongoDB itself, or Foursquare's use of it?
Flickr works the same way (not by coincidence, since we have several former Flickr engineers on staff).
I've never had either problem though... but if I ever need to shard I plan on doing it based on object ID. Then one request can be handled by multiple databases, "for free", increasing both throughput and response time.
But it sounds like the shard was still up. As per the sharding FAQ[1]:
"What if a shard is down or slow and I do a query? If a shard is down, the query will return an error. If a shard is responding slowly, mongos will wait for it. You won't get partial results."
They said they had a performance issue on the overloaded shard so perhaps it's possible that mongo believed the shard to be up, when it was instead overloaded. This meant any queries just waited, taking down the whole site.
As I said, this is speculation and as a massive production user of MongoDB ourselves[2], I'm interested to know more.
[1] http://www.mongodb.org/display/DOCS/Sharding+FAQ [2] http://blog.boxedice.com/2010/02/28/notes-from-a-production-...
"For reasons that are not entirely clear to us right now, though, the addition of this shard caused the entire site to go down."
an approach to CAP that chooses consistency over availability. I wonder why as it isn't a bank and they could have chosen the availability instead.
Sounds similar to some of the problems we're facing with MongoDB too. Indexes fragment a lot, especially if you delete documents frequently. You have to reindex everything eventually.
MongoDB offers background indexing (except for _id fields!) but it's not really an option if you need the indexes at all times to keep executing queries. Also, you'll need to reindex _id eventually and there's currently no way to do that without blocking the server.
MongoDB still has problems with locking too much for certain operations. The 10gen guys keep on introducing "yields" to those operations, but the simple (and fast) database/process-level locking could very well cause this... (not a 100% 'yes', but it seems possible)
Clearly, as their blog post indicates, they were unable to trace the root problem.
To me, the worst feeling in the world as a developer is when there's a major bug in your production site, and you can't figure out exactly why it happened. Then even after you get the site working there's that pit in your stomach of "what if it strikes again?"
> UPDATE Oct. 5 8:01PM: Our server team is still working to resolve the problem. The issue is related to yesterday’s outage.
But I would like to know if Foursquare has a commercial support contract with 10Gen. If they didn't, why not? especially for a service that big? If they did, how was it that 10Gen took that long to fix the problem?
Then again, it's no excuse for bad software. I've only used Mongo on small sites so far, and have been loving it.
In addition, it's very, VERY difficult to scale writes against a single object, such as Justin Bieber's profile data, say, if you've got a view counter on it. You can either serialize writes on read like Cassandra does, which has it's own drawbacks (the more writers an object has, the more expensive reads become), or you can have single-master-for-an-object sharding like MongoDB employs and most other production sites (Facebook, Flickr, etc) use.
1) Use connection pooling at the application layer to prevent overloading the DB of any specific shard. This means that if a shard has 16 CPUs, having 16 connections sounds reasonable. Additional connections will not give you more performance. This means you need to queue and throttle requests at the application layer and with some thought you can probably figure out what to do with the waiting users - show partial results? show a nice whale? A "loading please wait" sign?
2) If you didn't do #1 and the DB is getting overloaded, my normal response is to start shooting down connections. Oracle has separate unix process per connection. MySQL has its own way of shooting connections down. Put up a small script that will kill the correct percentage of sessions to prevent overload on shared resources. This will generates lots of errors and will cause a percentage of the users to hate you, but you won't be down.
2) This sounds like a great way to create data inconsistencies, unless you've got very tight constraints on your database, which is impossible in a sharded scenario.
I agree though, that ultimately they should have had some way to "fail whale" instead of getting overloaded.
http://groups.google.com/group/mongodb-user/browse_thread/th...
"Basically, the issue is that if data migrates to a new shard, there is no re-compaction yet in the old shard of the old collection. So there could be small empty spots throughout it which were migrated out, and if the objects are small, there is no effective improvement in RAM caching immediately after the migration." - Dwight Merriman (at the link in the parent).
"The kernel is able to swap/load 4k pages. For a page to be idle from the point of view of the kernel and its LRU algorithm, what is needed is that there are no memory accesses in the whole page for some time."
-antirez from http://antirez.com/post/what-is-wrong-with-2006-programming....
I'd love to know the root cause behind this specific issue. Was this a behavioral issue within the user base, or a technical problem that routed check-ins to this specific shard more than others?
Since they mentioned they partition their shards by userId that would probably rule out their routing process. I wonder if there was some event that caused a certain sharded subsection of users to start sending so many checkins?
And since this was a subset of userIds on the same shard - could this have been a targeted DOS or SPAM event?
I'm making a very conscious migration to MongoDB so I'm very interested to hear what the root cause of this was.
--- commenter
I can see the lock-in concerns with AppEngine, but an AppEngine level of abstraction seems so much more appropriate than manually deploying/configuring an entire infrastructure of proxies, load balancers, web servers, etc. Especially when an error can take down your whole site, like in this example.
But you know there's tons of tickets, tons of priorities and finally shit happens and some tasks are placed on top of the pile and become priorities.
Sucks that such a popular service had such trouble. I look forward to reading any additional posts they write explaining in more detail exactly what happened.
However, as a business you really want to give ALL your customers HA. Its not just a reputation thing, its a "we love you all equally" kinda attitude.
As for MongoDB, we ve been using in production for small insignificant things. FWIW, they have replication http://www.mongodb.org/display/DOCS/Replication and some cool new features like Replica Sets for failover and redundancy. Maybe they missed a trick?
I think the apology post was totally fair and he did categorically mention ".. This blog post is a bit technical. It has the details of what happened, and what we’re doing to make sure it doesn’t happen again in the future.". They could have dilly-dallied with words and said "we had a technical failure of a data nature" and that would have been just been plain stupid. So thanks for the detailed technical write up and hope there is more to follow.
Seriously though, are they THAT important?
I'm not blaming a flat-out bug in this case (the cause of the severe part is still unknown?), but it could also be architectural or operator error.
Edit: I found this reference to the 22hr outage that occurred, and I remember this outage, but I don't ever remember it being a 3 days outage.
http://www.internetnews.com/ec-news/article.php/137251/Cost-...
Fun fact: the "Steve Abatangle" quoted in the article is yours truly, and the author of the piece is Dan Lyons, now AKA Fake Steve Jobs.
No, OLTP was not at all new in 1999. OLTP probably means something other than what you think it means.
There is no way a 'regular user' is going to understand what a shard is, nor should they care. Ok, they explained sharding in laymans terms but then went on to talk about "reindexing the shard to improve memory fragmentation issues"... Woah, that means nothing to 95% of users.
If you experience down time and you want your users to be sympathetic then you got to explain whats going on in terms they will understand. Sure, include a technical explanation at the bottom for those inclined, but not as part of your main body.
Would love to hear suggestions on this topic.
Does 4sq have an engineering blog?
It broke. You fixed it. But, "Can I expect this to work again?" "Reliably?" All I heard was that it was broken.
It sounded a lot like, "something broke, it took a long time to fix."
It's silly for someone on Hacker News to say "well I thought the level of detail was fine" - of course you would, like the rest of us you're a technical geek. The point that seems to be lost is 95% of FourSquare's userbase ISN'T!
Also FourSquare is one of those startups that, in addition to the YC startups (for obvious reasons I guess), people give a little more favoritism to then perhaps other startups of equal quality/interestingness.
How do you think it could have been better? We struggled a lot trying to decide how much technical detail to include. We decided that including more information (even if a lot of our users didn't understand it) was better than "something broke, it took a long time to fix." Would love to hear suggestions on this topic.
Well, I'm not suggesting you wrote "something broke, it took a long time to fix" - I'm all for transparency. But if you are going to be transparent you need to communicate at a level at which that transparency can be understood by all of your readers. I'm sorry if some people on Hacker News don't get that.
So ok, here's how I would have written your post (for time sake I just did the intro - I'd have repeated the technical description after this block of copy):
Yesterday, we experienced a very long downtime. All told, we were down for about 11 hours, which is unacceptably long. It sucked for everyone (including our team – we all check in everyday, too). We know how frustrating this was for all of you because many of you told us how much you’ve come to rely on foursquare when you’re out and about. For the 32 of us working here, that’s quite humbling. We’re really sorry.
Below is an explanation of what happened and what we’re doing to make sure it doesn’t happen again in the future (a more technical explanation for those inclined appears further below)
What happened As you can imagine we store a huge amount of data from all of your user check-ins. We split that data across many servers as it's obviously far to big to fit onto just one. Starting around 11:00am EST yesterday we noticed that one of these servers was performing poorly because it was receiving an unusually high volume of check-ins. Maybe there was an incredibly popular party that we missed out on! :)
Anyway, after trying various things to improve the performance of that server we decided to try to add another server to take some of the strain off the original overloaded server. We wanted to move this data in the background while the site remained up - however for some reason when we added the new server the entire site did go down. Ouch!
We tried all sorts of things to ease the strain but nothing seemed to work. By around 6:30pm EST (phew, what a day!) we decided to try one final idea, which fortunately worked. Yay!
However it took a further 5 hours to properlly test our fix, and so it was only by around 11:30pm EST that we were able to bring the site back up. Don't worry, all of your data remained safe at all times, and that hard-won mayorship is still yours!
...
Anyway, if people disagree that you should always communicate with your customer at a level they understand, then I'd urge you to read http://steveblank.com/2010/04/22/turning-on-your-reality-dis... or http://www.readwriteweb.com/start/2010/05/is-your-startup-to... (pitching to investors, media or customers - it's all the same issues).
Also keep in mind I tried to edit their original post, as kinda suggested. I'm not sure I'd have written any the post quite in the way that they did - but I tried to work with what I had.
Also considering starting a separate engineering blog where it would probably be appropriate to go into more detail for those that are interested.
-harryh
I think end users have a problem if the one thing they need to know is expressed in a way they cannot understand. What they need to know is when the site is going to be up again and what the likelihood of it happening again is. Once they know that, I don't think they have a problem with added technical detail that's not meant for everyone.
To echo others, I'm interested to read the more in-depth post-mortem.
Good luck!