"There's no way to exactly estimate how many servers we'd need. So we literally bought thousands of them, and all the equipment and networks to go with it," Perlman told employees. Those servers, he said, came with lengthy contracts -- contracts that tied OnLive's capital up in maintaining servers that few (if any) users were actually using. "If you've got 8,000 servers and 1,600 users, how could we ever get to cash flow positive, right?"
-- http://www.joystiq.com/2012/08/18/documenting-the-death-of-o...
(How do the GPU clusters even work? And, at $2 it seems way more expensive.)
Cluster GPU Quadruple Extra Large Instance
22 GB of memory
33.5 EC2 Compute Units (2 x Intel Xeon X5570, quad-core “Nehalem” architecture)
2 x NVIDIA Tesla “Fermi” M2050 GPUs
1690 GB of instance storage
64-bit platform
I/O Performance: Very High (10 Gigabit Ethernet)
EBS-Optimized Available: No*
API name: cg1.4xlargeThe other big thing is that GPUs have very limited scheduling (i.e. no pre-emptive multitasking), so it's virtually impossible to virtualize into multiple VMs efficiently. So when you get a GPU machine you end up with isolated hardware which doesn't share the same economics as other Amazon solutions. As GPGPU become more popular, this should, however, improve.
I also question the need for one VM per client. There's a very small set of inputs (input devices from the client) and a lot of static output (rendered frame). Why not share the same VM for a set of games that share the same GPU?
2. Worst case, at the very least you sell your excess capacity as cloud computing resources so your servers aren't sitting there doing absolutely nothing.