Zynga Moves 1 Petabyte Of Data Daily; Adds 1,000 Servers A Week
techcrunch.com
techcrunch.com
- 215M users.
- Up to 1000 servers per week.
- Say, they've been doing this for 30 weeks.
- So they need 1 server for every 7167 users? Is this normal for mmorpg type games?
- Are 215M people playing around the clock?
I'm just a little baffled by this number.
For example, we developed a product for Turner on top of our platform. Take a worst case scenario - everyone in the game moves using arrow keys, and they're all moving. A typical room might have 25 people, each broadcasting their location roughly every 0.5 seconds. This location update is broadcast to all 24 other people in the room. So each person generates 2 messages into the server per second, but results in 48 broadcast messages per second. Of course there are 25 people in the room, all of whom are moving. So we have 2 x 25 = 50 messages into the server per second, and 48 x 25 = 1200 out going messages per second.
A typical node of our platform might support 4000 players. If each room supports 25 people then we have 160 rooms. Every second the server needs to broadcast 1200 x 160 = 192,000 or nearly 200,000 messages, and receives close to 8000 messages.
Movement packets are easy. We can route these through the system quickly, but then think about the "slow" requests, such as writes to the game state persistence store (Mongo DB) or slower still, writes to the SQL database.
With 8000 incoming messages, and a typical 64 thread processing pool each thread needs to churn through 125 messages per second. Or one message every 0.008 seconds. Those slow SQL writes might take 0.5 second, during which time 0.5/0.008 = 62 messages have backed up.
Many of our games (and im inferring Zynga's) store huge amounts of persistence / state data. Where has the player been, whats completed, timers, friends, etc, etc. Much more than you'd imagine - and thats not including the analytics data. These slower writes very quickly add up and reduce throughput.
Of course there are lots of things to optimise the number of messages, but this starts to give an idea for how quickly messages back up in a MMO type environment.
If you've bothered to read this far you might be interested in a short blog post I wrote about our casual MMO architecture http://dubitplatform.com/blog/2010/3/4/under-the-hood-dubit-...
Some MMO designers actually really want to bring that number even lower, though lots of other people at game companies are looking at ways to pack more users on each core. But from a designer's perspective, if you typically get less than 1/50 of a CPU slice per user, maybe 1/100, and much of that is taken up by doing basic bookkeeping, it's hard to put in things that even 1990s RPGs took for granted, like NPCs with at least passably interesting AI.
Zynga games are more comparable to social websites than WoW.
Movement is one already mentioned, but in general, AI isn't only turned on in combat. NPCs move outside of combat - most creatures in the world wander about or follow a pre-defined path. I don't recall if they did on WoW, but I know in Everquest creatures used spells of their class both in and out of combat, such as buffs and heals - a creature in Everquest would even heal and buff up their allies. If monsters are buffing, you now are tracking those buffs & their durations - when we added monsters buffing to our multiplayer game, we made them infinite in duration to get acceptable performance.
Another operation that could be taxing is calculating line of sight (LOS) - this factors into any action an NPC or player would do. You can't have monsters attacking players, or buffing allies, through walls.
Buff tracking as a performance hit is surprising to me, though. I'd presume you were tracking monster lifespan - with this and an offset for 'AI interruptions' (like for chasing players as the buff wore off) you could recover the buff times and remaining durations without touching the monsters when they were out of player LOS. I can see LOS as a more taxing system, but monster LOS has no z-axis and the dungeons don't seem to require too many nodes for a pathing graph...
At any rate, most MMOs have rather... shallow... gameplay mechanics, and I don't think it's really computation that's holding them back.
"...it takes roughly 20,000 computer systems, over a petabyte of storage, and over 4600 people. Using multiple data centers around the world, this works out to a total of 13,250 server blades, 75,000 CPU cores, and 112.5 terabytes of blade RAM."
It doesn't say if that's total or just US (since WoW is hosted by different partners in Asia, last I paid any attention).
1. The servers aren't all application servers. There's load balancers, memcache servers, lossy memcache servers, etc.
2. It's the cloud. It includes large instances, HA-Proxies, memcache servers, etc. 3. There's very high write characteristics for social gaming's data store, but the read is similar to normal web application.
4. He took an aggregate server #, and generalized it over a period of 30 weeks. It's directionally accurate, but I wouldn't take it literally.
We created a multi-player rpg for mobile phones at startup weekend a few year ago. It actually did ok and it was growing pretty quickly. I hosted it on AppEngine to start and the resources just started going sky high. Players were initiating an attack at least once per second (not to mention all of the other server calls like messaging, weapon purchasing, scanning to see who is in your area).
I thought we were going to have to reengineer in order to try and make it as efficient as possible. The thing is, these addictive multi-player games that are all about throwing data back and forth end up taking huge amounts of resources.
We eventually ditched the game, but the site is still up at http://killyourneighbor.com
There were some issues with a high number of writes, but it has got better.
The unscheduled downtime also killed us.
But really, it was just the sheer amount of traffic. Real-time multi-player games that people have with them all the time can just create a ton of requests.
Btw how is it possible that 10 or 20 years ago using megaflops or MIPS (or other sane measures of computer power) was normal and now we are here talking about "servers"? :)
Edit: I'm searching around the internet if there is some new general measure of computer power that is able to reflect the real-world performance in non scientific applications. No luck so far.
When Zynga launches a new game, they can get over 1M new DAILY players on that game within the first week. [1]
For their games, it seems that 7% of daily players are online at any given time. [2]
So within the first week of launch, you have, about 100,000 people playing, simultaneously.
Let's say the client is sending 5 requests per second to the game server. And let's say that this week, all 1000 servers went to the new game.
Now, of course that's just an AVERAGE across the day. In reality, there will be a cresting factor, like any distributed system, and sometimes the load will spike, say, 8x.
Is it ridiculous for a single server to process 500 requests per second on average and 4000 requests per second peak? That seems pretty standard by my book.
[1] http://techcrunch.com/2010/06/14/pincus-frontierville/ [2] http://techcrunch.com/2010/08/17/zynga-launches-first-locali...
PS: The key to pulling this off is running the same simulation on the client using the same random seed. User clicks attack blob with fork. > Client says: attacked blob at (timestamp, with fork) > Server says: blob took X damage at (timestamp) random seed at (value) > Client shows: A long string of actions that adds up to that value and or requests a new game state if something does not add up.
You mean per minute, right? Either that or I'm way worse at SCII than I thought.
My original point was less about the actual numbers, and more about the exercise of estimating server load, instead of just blindly declaring that they must have crappy code.
That being said, launching 1000 nodes for non-batch jobs is still pretty remarkable.
Also keep in mind interactive multiplayer games with persistent state are different beasts compared to usual web applications. Games are very write heavy and do not cache well.
Really this probably means that one week they added 1,000 but the actual average number is much lower.
Makes for a good talking point worded that way!
http://gigaom.com/2010/06/08/how-zynga-survived-farmville/
Zynga uses Apache PHP on the front end, memcached for active user play and MySQL on the back end. It uses memcached to store key value pairs to deal with active user play during sessions and then later writes it to disk.
Or would a node fail rare enough for the data loss to not matter much?
The people playing Farmville are enjoying themselves and expressing themselves. Would it be better if they deployed that cognitive surplus and attention towards Wikipedia or a similar project? Of course.
But it doesn't make me sad that there are 215 million people in the world who have a comfortable enough life that they can afford to dabble in some casual games. In fact, I hope that in 100 years we can have all 7+ billion humans striving for self-actualization instead of food and water.
We used to batch commands in the client and submit every 6 seconds. It was still a crazy number of writes, because basically every click is a database update. MySQL is really bad at that sort of work load, but I migrated almost everything to Redis. Really fun problems to solve!