Combine Multiple AWS Instances into a 16-GPU Monster Machine
bitfusion.io
bitfusion.io
I'd like to see something in the cloud thats bare-metal / full access to GPUs (Maybe a good idea to start one). For scaling higher with a very large number of GPUs, you'd need Infiniband but at some point there is going to be a bandwidth tradeoff.
It would be interesting if someone could run some benchmarks of these instances versus a physical server.
We're adding support for other clouds, particularly ones with higher-end GPUs so feedback like this is good to know.
What if you limited people to writing in a domain specific language: one that ran distributed on this infrastructure? How would that make it different than folding at home, for example?
"“Monthly Uptime Percentage” is calculated by subtracting from 100% the percentage of minutes during the month in which Amazon EC2 or Amazon EBS, as applicable, was in the state of “Region Unavailable.” Monthly Uptime Percentage measurements exclude downtime resulting directly or indirectly from any Amazon EC2 SLA Exclusion (defined below)."
Amazon EC2 SLA Exclusions... (v) that result from failures of individual instances or volumes not attributable to Region Unavailability
Presumably if you're using a cloud hosted in people's basements, if someone's basement server dies, you'd just pick one from someone else's basement, so this model could provide better availability than AWS.
Something like a combination of Freenet, TOR, BOINC, blockchain etc. technologies, using the current "legacy" internet as a backbone, where anyone can voluntarily offer their computing and storage resources to the network at varying levels of participation.
Say you could offer your laptop as a simple discovery/directory node to simply help others connect and find stuff, and your desktop as either a static-content serving node or as a computation node that can host distributed applications, like SETI@home or web apps like Facebook.
Maybe even reward cryptocurrency to those who offer the most resources.
I wouldn't be surprised if certain services became more centralized, offloading computing to the cloud. Consumer grade devices would become thin clients (like back in the day). NVIDIA has hinted in this direction, I can remember something about GaaS (Gaming as a Service).
The bigger problem we ran into is that AWS instances use a lot of small GPUs, which don't scale well using a lot of deep neural network tools (e.g. theano). It was never really a viable option for us.
Just go to the custom link at the bottom of the page, the link is: https://console.aws.amazon.com/cloudformation/home?region=us...
There you can select any number of clients and servers. For example: 5 clients and 1 server (many to one), or 5 clients to 5 servers (many to many).
Is it the fact that GPU code already runs in parallel streams that makes this possible?
We are not limiting our software to AWS so you can built this kind of service on any kind of cluster by installing our software directly from https://boost.bitfusion.io - I say cluster, because we have played with the idea of thin devices accessing remote GPU instances in the cloud, but over public networks the network performance was a limiting factor.
If you haven't, I could probably contribute some strong and weak scaling testcases.
Did you guys think about building it further out to provide a GPU load balancer for multiple frontend machines running Cuda / OpenCL?
Since the GPU client is now abstracted from the GPU devices by placing the GPUs across the network. It seams like time-sharing should be next logical step.
Check it out: https://console.aws.amazon.com/cloudformation/home?region=us...
Hopefully you'll see some good uptake.