World’s fastest supercomputer unveiled in China
wired.co.uk
wired.co.uk
All I know about supercomputers based on GPUs is that they're very cheap, very theoretically fast, and very hard to program for effectively. Any experts in here who can enlighten us?
LINPACK in general is an interesting benchmark. We recently bought a new cluster. We were faced with an interesting decision. Get 4,000 cores with Infiniband, giving us a spot on the top500 list (We've been on there for the past 5 years) or get a 6,000 cores with gigabit ethernet, losing our spot on the top500. (We get %90 of our theoretical max over infiniband. Gigabit drops that down to 50%). We did a survey among our users and very few of them ran multinode jobs. Most used 8 or 12 threads, which runs fine on a dual socket hex core westmere. So we dropped infiniband. So even though we technically have a slower cluster than we could have had, we have more efficient utilization and are able to offer our users more resources.
Since very few comercial applications can currently take advantage of a hybrid system, it seems to me this cluster was built with the top500 in mind, rather than being efficiently used. As far as useful results, I wouldn't be surprised if it only accomplished 2 or 3 times what we do, and we're a 10,000 core shop that probably won't be on the next top500.
One simple question to ask is: how many computations am I going to perform on each element I transfer? For example, if you're doing a matrix multiplication of NxN matrices, each element will be used in 2N computations. (I think. Someone check my mental analysis.) So, matrix multiplication is a pretty good fit. But what if you're just summing a vector of size N? Then each element will be used once. Probably not a good fit.
Then any moronic theoretical physicists/geophysicist/industrial chemist can just use MPI, LINPACK,Nag to do useful work on it
Moreover GPUs have problems with double precision floating point numbers which are a must for most scientific problems (though sometimes you can utilize mixed precision approaches -> estimate in single precision, then get a correction in double).
Another problem is having ECC (error correction). You don't want to run a computation for a week to find out that it crashed because of data corruption. Latest generation of Teslas has ECC, consumer cards don't.
Lawrence Livermore is already planning the next "World's Fastest Computer"
Except one thing. Until recently GPUs have only done single-precision calculations. Double precisions has been added with the latest architectures but double is, I believe, still very slow. So this means that single precision results are sufficient for the top500 tests as well as the particular applications (in the Tianhe’s case: weather and oil exploration)?
1. Any entrant that extensively relies on GPU's for its performance should not even be on the list or at least appear with an asterisk. The "G" in GPU does not stand for general. Consequently this machine is principally applicable to specific problem sets that are designed to take advantage of the processing power provided by the GPU.
2. Primary attention in the supercomputing space has shifted focus from using pure flops as a metric to flops per watt as the primary concern. Beyond that, focus has also shifted towards interconnect technology to keep all those flops well fed with data.
Just like the Chinese to be a few steps behind the thought leaders but front and center for biggest and most easily duplicated.
I'm sure Google could claim to have the largest "computer" if they wanted to, or if the definition of "supercomputer" were stretched a bit.
Yeah, I only use bittorrent to download linux distros as well.
Edit: According to the PCMag article they're Xeon processors.