Amazon Orders More than 10,000 Nvidia Tesla cards
vr-zone.com
vr-zone.com
I also remember thinking that a 64kb row size for DynamoDB was very odd.
I wonder if these things are at all related.
Early generation of NVIDIA gpus did not an automatic Caching mechanism or could not for CUDA, I forget) that could help solve this issue. But they did have memory available locally on each compute unit where you could manually read / write data into. This helped reduce the overall read/write overhead.
Even when the newer generations have the caches, it is beneficial to use this shared / local memory. Even when the shared / local memory limits are hit, there are alternatives like Textures in CUDA, Images in OpenCL that are slightly slower, but significantly better than reading from DRAM.
If the article is correct, Amazon paid 15 million for those cards which will be out of style in about two years (not that they have to get rid of them, but something faster, easier to maintain (if Nvidia starts opening up to Linux), with more memory and less power usage will come out. They'll have to fork over a large sum of money again to keep their top "on demand computing" title.
Amazon's cluster GPU right now has two Nvidia Tesla Fermi's in it. I'm going to assume Amazon will split their new cards into twos and fours, at about half of each. That's ~1750 new computers that are going to load up. Looking at the current rates of the cluster, it's $2.100 for an hour of the normal, I'll say it will be $4.200 for an hour on the jumbo with 4 GPUs.
They paid $15 million for just the cards. They need to get 2380952 hours of usage out of the machines to break even on the cards. They need to log 1360 hours per machine to break even, or have someone run all the machines at full bore for 56 days. As the cards are the most expensive component (assumption), and the total price of the computer will be about the price of one of the cards, we'll add a little bit of over head for all the other things they need to do to make it work - 120 days of full time use to break even on an investment of about $25 million (they need to buy lots of other things to put all the GPUs in, and worry about all that heat, and have a place to put it all, and have people install the new computers, etc...). I wonder what the actual usage of those clusters are, and if they've had anyone sign a deal saying we'll use the cluster for an entire month. That's a beautiful maneuver though, say CERN didn't want to do all the data analysis from the LHC in house because by the time they got to this part of the experiment, their technology they purchased previously would be way out of date. Just let Amazon do it. They will always have the latest technology, and you'll have an inexpensive way of leveraging that power.
Assuming they can make it all work (and I'm sure a lot of their decisions now are strategic decisions aimed at future investments) this is a great time to be a computer user, log on and get the best for a couple hours for a couple dollars. Instead of shelling out $1500 on a new computer personally, I could log a ton of EC2 hours getting significantly faster, more powerful machines, that never get 'stale', and their lives are much happier (my computer probably doesn't do anything "intensive" 70% of its life, whereas the EC2s are probably pushed a bit harder than that).
If you're working on applications that will need to be using the gpu regularly, you can build a system with 4 gtx580s for about $3,000, and one of those systems will outperform 2, maybe 3, aws gpu instances, which will run you about 1000 per month each. The ownership number does not include data center/power/etc., but I still think buying is better value if you'll be using it a lot.
Now, if you're running gpu jobs sporadically, aws may make sense, but you should really look carefully at this, it's not the same value relationship as hosting web servers on aws (which I'm a general proponent of).
Although, to be fair, that may change if they really do pass on some of their savings from this deal to the user.
The reasoning may have been to focus the GTX series more on gaming. Or it could be more sinister to push more people towards their costlier Tesla Line. Considering that they came out with the K10 which has terrible double precision performance, but incredible single precision performance, I think they are heading towards multiple Tesla lines and want to push the GTX series away from the serious GPGPU computing.
Disclaimer: The following is my company's blog. The post is authored by me. http://blog.accelereyes.com/blog/2012/04/26/benchmarking-kep...
I have a 590 and I'm pretty happy with it, been eying the 690 but I wasn't able to find any real world non gaming benchmarks.
Amazon offers a heavy usage deal which comes out to ~ $11,000 per instance if used 24x7x365.
You could argue that the cluster / machine you set up would be useful for more than a year. This is true to an extent, but at the current rate of development GPUs become obsolete rather quickly and suddenly having a cluster on the cloud sounds more appealing than going through the process of updating your machines every 18-20 months.
They might not run their hardware at a profitable capacity right now, but in 3-4 years when some labs are looking at replacing/updating their current systems, they will look at Amazon as an solid alternative and then things will turn up. And until then they can't half-ass it with too low capacity because when labs do some trial projects, if they come back with a "sorry we're out of instances" error, they'll decide they can't trust Amazon.
I'm facing this very problem myself right now with some GPU-bound calculations. The issue for me is that the software I'm using isn't free, so I'm stuck running it on the one machine that I do have a license for it to run on. It's a very frustrating situation to be in, as the hardware isn't so great.