FASTRA II: the world’s most powerful desktop supercomputer
fastra2.ua.ac.be
fastra2.ua.ac.be
If the authors actually tried to measure Linpack scores on their cluster, they'd find that peak performance would be much worse than the 12 teraflops they're claiming as "peak" performance. This link [2] seems to indicate that one S2050 along with a fairly fast CPU gives you about 0.7 teraflops. Even if they get perfect scaling across GPUs, which they won't, and if they could solve all the heating issues with their design, they won't get more than about 0.7*6.5 = 4.55 TFLOPS peak on linpack.
In any case, many (39 out of 500 to be precise) of the most powerful supercomputers in the world also use GPUs to accelerate certain applications so there's simply no way this thing would be faster than those computers.
[1] http://www.top500.org/lists/2011/11/press-release [2] http://hpl-calculator.sourceforge.net/Howto-HPL-GPU.pdf
I'd like more information on the CPUs and CPU/GPU bus of this FASTRA II. CPUs are really gaining back a lot of their performance relative to GPUs - an M2090 is only about 5-10x faster than the newer Intel Xeons in a properly multithreaded program that pays attention to the NUMA nodes. This is efficient enough that for large simulations which do not fit on one GPU, running portions of the simulation on CPU easily overcomes the CPU/GPU communication overhead in terms of performance.
"the graphics cards, which are connected to the motherboard by flexible riser cables." - so this sounds like typical PCIe based design. Usually the bundle of GPU cards on PCIe backplane is placed into separate case which connected back to the CPU host, and in this case they engineered it into one big case thus each card getting it's own x16/x8 pipe.
In general such "supercomputer" is cheaper, hardware-wise, than equivalent CPU-based, yet it is still more expensive human labor (ie. programming) -wise.
Immediately, back in my head, I got this gnarling feeling that I should also use my skills in such life improving projects instead of building the next social _______ life drain.
200 series era Geforce cards were slower per watt and per dollar and per slot (all important here) than 58xx series Radeons on both single precision and double precision math (and integer as well, but integer is only useful for crypto and Radeons have native single cycle instructions common for crypto which makes them about 4-5x faster than Geforces, but not a typical use case).
And if we're comparing series 400/500 to 69xx, the beating is much worse.
http://ewoah.com/technology/a-very-good-guide-to-building-a-...
http://stackoverflow.com/a/4642578/13652
All this is moot if the application they're running only supports CUDA - then they're stuck with Nvidia.
If they're stuck on CUDA, I feel sorry for them. They'd be better off rewriting the app to use OpenCL to take advantage of AMD's superior chip design and not be stuck on Nvidia/Windows (Nvidia Linux support, especially under CUDA, is horrible; rather depressing once you consider most CUDA clusters run Linux).
As a linux developer doing scientific visualization, and having tested (and still testing, btw) how bad still is the AMD driver status on linux, I see basically nvidia as my only choice if I want to get anything done with decent (and stable) performance.
I wouldn't run a cluster on anything but .nix, so I really see no alternative on nvidia by my reasoning here, even if nvidia's architecture is actually inferior.
But I'd really like to have some opinions on that front. I really want alternatives, and the closed nature of nvidia's drivers is a problem for me. So, can you elaborate?
I also programmed on w2k (from around 2001-2004) and I used to target both nvidia and ati, and I still remember the sheer number of bugs I had with GL in general with the radeon drivers. But I cannot really speak anymore on that front.
Nvidia driver quality continues to be subpar on Linux, yet I have zero issues with Catalyst. And even if Catalyst had a problem, as I described in the above URL, AMD is shelling a lot of money out to make the FOSS driver their one and only driver which already is shaping up to be a lot better than Catalyst.
I just don't see a point in doing business with a company that is clearly trying to lock me in with closed source vendor specific APIs if they have no interest in supporting me after the sale of the hardware.
In addition, their OpenCL implementation (which is shared across both Windows and Linux) seems to run code much slower than it should, which I suspect (but have no way of proving) they are doing it to make CUDA look better. The same code ran on AMD hardware has zero issues, and the equivalent code in CUDA runs about as fast as I think it should.
And also, re: Catalyst vs FOSS, in about 2 years the Mesa/Gallium OpenCL stack should be finished, so not only will I be able to quit using Catalyst altogether, Nvidia users will be able to make the better use of their hardware and ditch the broken drivers Nvidia keeps forcing on people.
Now, who knows, maybe Nvidia will change tactics and quit screwing over customers and developers, and release the programming manuals for their hardware so FOSS can develop a better driver faster... I just don't see it _ever_ happening.
Even if AMD hardware had a higher total cost of ownership, the fact AMD actually cares about me as a Linux customer is why for the past 5 video cards I've owned they have gotten my business.
re: Your programming experience, 2004 predates AMD taking over ATI. Driver quality pre-AMD was worse than Nvidia was at the time, but it has greatly changed. AMD has vastly increased the quality of the Windows driver, and the original Linux driver was thrown out in exchange for one that shares almost all the code with the Windows one (day and night kind of change in quality).
AMD didnt finalize their purchase of ATI until 2006, and it wasn't until 2007 or 2008 that driver quality started taking off.
As a statement of driver quality, people who use my software regularly report 30+ day uptimes using Catalyst on Linux. If it was so buggy and unstable, this would probably be impossible.
In terms of GL implementation, the nvidia driver has everything you would expect. If you use GL directly, it is one of the most conforming implementations I've been using (supporting "deprecated" stuff like bitmap operations, line and polygon aliasing, imaging extensions, all GL revisions, pretty much all ARB extensions, etc). Everything works from Cg to CUDA to OpenCL to VDPAU. There is really no feature gap between windows and linux.
I've never compared OpenCL to CUDA performance really, but switching to OpenCL for new projects is something I'm planning to do.
My mayor problem, really, is that CUDA is locked to specific versions of GCC, and that is even of a bigger issue than the closed nature of the driver itself.
I would guess the main attraction of Nvidia cards was probably CUDA, which is still a nicer environment to program for than OpenCL and seems to have a more lively ecosystem of libraries, tools, etc.
Also, floating point & integer math aren't the only relevant performance metrics. For their application domain, memory bandwidth & latency would probably be at least as important. Kernel switching time may be a factor too - there are many things you can't do with a single kernel. I don't know (but would be interested to hear) how the cards stack up against each other in those terms.
I am a software engineer (or whatever the popular term this week is) who uses OpenCL frequently. I couldn't imagine being stuck on something like CUDA, between the C++ obsession and the fact it follows Nvidia Cg's way of syntax design instead of GLSL (OpenGL Shader Language) makes it difficult to use.
In addition, there never will be a native Linux implementation that is open source as Nvidia has basically tried to screw over Mesa/Gallium repeatedly with vague IP rights threats (see also, Nouveau, the X/Gallium driver for Nvidia cards, and how they unofficially threatened to sue if Nouveau continued reverse engineering hardware interfaces; they backpedaled when it got media coverage).
So, I dunno about you, but I don't like depending on companies that continuously screw developers who made the mistake of supporting their hardware.
OTOH, AMD has paid over $2 million in their open source efforts including having their legal department clear the programming manuals for use by X/Mesa/Gallium and also hiring multiple full time developers to work on the open source drivers, and has clearly supported FOSS at every turn.
Re: the original point, I have the impression that Nvidia has been more active in engaging universities and that's probably a better explanation for why their cards were used rather than AMDs.
AMD will support DirectCompute because its the native API on Windows, the same way they continue to support Direct3D. This doesn't mean that Direct3D is more viable of an API than OpenGL, but it means its popular on one platform.
In comparison, OpenCL's compute shaders use the same C-ish that OpenCL does for GLSL.
OpenCL and DirectCompute are pretty high level considering they're running on what equates to massively parallel DSPs.
Oh, and OpenCL already runs on heterogeneous setups. There are OpenCL compilers for three different GPU archs, x86, PPC, Cell SPEs, a few DSPs, and ARM. An OpenCL program can use all of them at the same time if available.
It almost sounds like AMD's project that turns Java bytecode into GPU-able code (although it is limited to what you can do normally, such as no allocation of memory/no object construction, only certain type primitives, and it should be largely branchless/loopless).
See: http://developer.amd.com/zones/java/aparapi/Pages/default.as...
SSE2 integer usage may be fast, but GPUs are faster, typically by several magnitudes.
7 x 5970's would give you just under 35 Teraflops instead of 12.
8 5870s packed into two cases would survive if you keep the cases closed and put 3 Delta AFB1212s in the front pushing all the air through the cards and out their rear vents, this is how most Bitcoin miners build cases.*
If done right, 8 5870s should easily be kept under 85c even when marginally overclocked.
* Large scale Bitcoin miners have a habit of not using cases and use flexible risers to keep the cards away from each other out in the open. I do not recommend this as you lose out on high pressure cooling, a staple of enterprise/high performace computing.
Before adding water cooling - it would lock after 30 minutes of mining at standard 725 Mhz.
This system worked very well for ~4 months (and then I sold it to guy from Alaska, since it became too hot in Texas :))
(http://fastra2.ua.ac.be/?page_id=53)
They claim "slightly over 60 degrees".