A distributed supercomputer using Parallella boards
supercomputer.io
supercomputer.io
Basically, a narrow set of applications that leverage these sorts of chips best and are easily distributed can benefit from this setup. It will probably have more raw performance per watt than most @home setups. Yet, I'd guess lower potential in number of nodes given fewer Parallella users on the market. I hope they benchmark it with many, realistic software that can be compared to other parallel computing methods. This would give us useful information on whether to use something similar for a local site or another @home, non-profit effort.
Clearly we don't have the volume to make this as big as @home, but let's not forget how expensive it is to run a big computer. The folding@home project claims 36,000 TFLOPS of performance. If we generously estimate 1 GFLOPS/W for most home computers, then we are talking about a cost of ~$36M/year for the people donating cycles. This would be a smaller system, and if we get all 10,000 Parallella boards in the field hooked up, the electrical cost would only be ~$50K/year.
I am not going to get into the whole "what is a supercomputer thing":-) I am never going to beat Nvidia at that game. http://nvidianews.nvidia.com/news/nvidia-launches-tegra-x1-m... What I do know is that 180,000 CPU cores working on one problem is one insane distributed computer.
I'm still waiting for Adapteva to combine their work with SGI- or NUMAscale-style tech to have a UV-style system full of CPU's, your accelerators for general use, and optionally FPGA's for special purpose stuff. The CC memory + bandwith + ton of Adapteva IC's per node would = incredible performance per dollar and watt. Maybe. ;) Any plans for integrating with ccNUMA architectures?
I forget how long that's been but I have to admit to be being really disappointed that there hasn't been much demonstrable progress. There's still no 64-core version of the Parallella board. There was some talk about an update to 16-core Parallella board but I don't think it's shipped. As far as I know there's still no way to connect multiple Parallella boards using their "e-Link" ports. And lastly no way to assemble a small Adapteva cluster with a more favorable ratio between the Adapteva and Zynq (as far as power, heat and money go) than a stack of the 16-core Parallella boards.
Obviously I expected too much and I guess I'm being a tad unfair. That I've seen the dev team have always been decent, reasonable, and polite; so I really hate to write anything negative at all. I just thought they'd be much further along by now and that I might have a mini many-core cluster.
As it turned out there were a few folks that appeared on the forums who had these kinds of misconceptions. After all this time there weren't that many; so I probably overestimated the number of people who signed on based on those misunderstandings.
Having said all that, I guess I should point out that while I do have a very skeptical view of kickstarter as a whole, I don't view Adapteva's kickstarter as being very unethical or as complete failure. They eventually were able to deliver on what I see a bare minimum of what they said they were going to do. Though I do feel we have to be honest and acknowledge that they weren't able to fully live up to all the claims and promises that were made... like you said "more about suggesting dreams than making contracts".
https://www.kickstarter.com/projects/adapteva/parallella-a-s...
Kickstarter goal (first sentence from day 1): Making parallel computing easy to use has been described as "a problem as hard as any that computer science has faced". With such a big challenge ahead, we need to make sure that every programmer has access to cheap and open parallel hardware and development tools. Inspired by great hardware communities like Raspberry Pi and Arduino, we see a critical need for a truly open, high-performance computing platform that will close the knowledge gap in parallel programing. The goal of the Parallella project is to democratize access to parallel computing. If we can pull this off, who knows what kind of breakthrough applications could arise?"
Curious, besides being woefully late, what major promises did the Parallella campaign NOT live up to?
Can GPUs do this. In my view, not a bad result considering 2014 general availability? How long (and how many $billions did it take for CUDA to catch on...)
https://www.parallella.org/2015/05/25/how-the-do-i-program-t...
In terms of cost, not really sure what you are referring to... supercomputer.io is free to researchers, which last time I checked is less than not zero.
As far as major promises did the Parallella campaign NOT live up to (excluding the 64-core version), as far as I know it's still possible to connect multiple Parallella boards together using the Epiphany eLink interfaces. That novel interconnect was to me a key feature of the Epiphany architecture. Without it interfacing multiple Parallella boards together is forced to go through the Ethernet provided by the Zynq... which (assuming the eLink architecture performs as well as is claimed) is impediment & disappointment.
Without fully exploiting the eLink interconnect I think it's really difficult to get a full and accurate understanding of the capabilities and weaknesses of the Epiphany architecture.
Please don't take this in a personal or highly negative way. You guys have accomplished a lot and I really, really don't want to diminish that. You've also dealt with the difficulties that the project faced with a professional, straight forward, and positive demeanour, which I feel is both lacking in many other kickstarter campaigns and which speaks volumes for the quality of character of the developers working at Adaptiva.
I really want to see you guys succeed. I am still interested in buying / building a mini-cluster but I'm holding off until they can be properly interfaced through the eLink fabric. Also, I very much would like to see a next generation Epiphany chip and subsequent nextgen Parallella board, beyond the v2.0 Epiphany III Parallella board that was discussed in the forums.
But how does trying to send fragments over the internet really work?
[1] http://www.nvidia.com/object/jetson-tk1-embedded-dev-kit.htm...
Sequential part and Amdahl’s law will catch you quickly.
The target usage model seems to be where the actual algorithm employed can be pipelined, streaming in-flight results from one core to another. That has even more limited applicability than embarrassingly parallel architectures, and is incredibly difficult to map general problems to keeping the cores busy.
[1]: http://docs.aws.amazon.com/AmazonS3/latest/dev/S3TorrentRetr...
It feels like a big change in dev thinking that I cannot ssh into the board anymore, but also very interesting.
By the way, I see that you'll be supporting the SabreLite soon. Would it mean to be able to support other i.MX6-based boards like the VIA VAB-820 [1] or the UDOO?
We keep adding devices and will soon release a guide on how users can add their own devices to the mix. That said, the primary determinant on whether we can support a device is whether a yocto/openembedded BSP exists and is relatively modern (uses a kernel above 3.8). If that exists, it's almost certain that resin support will be relatively easy. Happy to chat more, email in profile.
And yeah, as a lucky chance, VAB-820 just had a yocto layer released with 3.10.17 [1]. Looking forward to see where this is headed!
What application will be deployed on this for the live test on May 30th?
Basically, there is no scientist in the world today who doesn't have the funding needed to scrape together four used core i7s and network them. Unless you can significantly beat that performance level, you're not "democratizing access to supercomputing for the scientists that need it" in any meaningful way. If you are significantly faster, even for just one particular application, that should be up front and center on you FAQ.
Will you send out any test load before May 30? Would be better to see those cores to work if they are online, otherwise it feels kinda wasted. Wouldn't be surprised if people go offline after a while if there's no work done.
Another question: how are you planning to implement fault tolerance? If you're running across hundreds of nodes via the internet, the probability of one failing while my job is running is high. Are you going to run a fault-tolerant scheduler?
And: how are you going to do file I/O? Does the user have to run the master MPI process on his/her own machine and do I/O there?
Either way, the point of supercomputer.io is that it would be free.(thanks to the contribution of everyone donating cycles). Think BOINC, not AWS. This is not for commercial use.
https://aws.amazon.com/ec2/purchasing-options/spot-instances...
Next!
http://parallellocalypse.s3-website-us-east-1.amazonaws.com/...
http://parallellocalypse.s3-website-us-east-1.amazonaws.com/...