China Tops U.S. in Supercomputers
eetimes.com
eetimes.com
Every dollar spent on building these computers is a dollar NOT spent in the budget of a research-driven grant. Give these dollars to physicists, mathematicians, genomics scientists, chemists, climate people etc. This way the demand in computing will dictate the type of systems built.
p.s. I've attended the event in person that's listed in the article hosted by the OSTP, which is a part of the White House (National Strategic Computing Initiative). My conclusion was most of these computers aren't needed and a misdirection of funds to the tune of billions of dollars that computing / scientific computing could really benefit from in other ways.
If it's not an MPI job, it's going to be painful.
What are you using for communication between nodes?
To make the tone clear I'm genuinely interested as I'm not sure what running on a Cray prevents you from doing?
We don't need to communicate between the nodes; We don't need MPI. We need distributed databases and we might need Spark. We need machine learning. We need servers. We would like availability and reliability.
We get full batch nodes that are relatively anemic in memory for our workload (4GB/core) for a maximum of 24 hours at a time.
Which just emphasizes the point that most applications would be better served by cheaper, more resilient hardware.
Huh. At the core of almost every problem I've ever encountered in my 20+ year career programming that I would consider "interesting" is the need to allocate work across many compute nodes that have some amount of shared state. You could just as well take A* as a baseline for these types of problems. Your typical cloud computing infra/design is not a good match for these.
It's true that we need shared data, but we don't need shared state (memory) for any of our workloads, and shared data (e.g. disk/db) is a much easier problem to solve than shared state.
There is still coordination and orchestration, but that's an extremely coarse amount of communication compared to the cost of Cray interconnects. That being said, there are cases where we might benefit from the Cray networking, but that comes at the cost of other tradeoffs (no local disk, low memory per core).
So what do we do? Well, we use a handful of Ferraris to get the job done because they are available when cheap bus would have suited us fine. The double whammy is that the Ferraris end up in the shop all the time and occasionally somebody else gets exclusivity to them when they want to get in a Bell prize submission.
Or you can send get the network admins to allow you to send them a email, but the first suggestion sounds more fun :P
1) I have a rail system that I want to optimize train movement through the topology. My search space is the number of running locomotives to the power of the number of switches in the system. On a continental rail system, that's a number with many, many, many zeros after it. To optimize this, I have many cores chewing through the space, but to know where to search they must share a visited node cache otherwise they will search each others space and waste effort.
2) I have a system that predicts part failures by collecting sensor data of the machine in a time series. I must look for patterns of change in these sensor readings and compare that to patterns that I know about, often drilling through data over many weeks looking for feature alignment in the sensor matrices. Same problem: many cores, some amount of space that is known vs. unknown that must be shared, but as a result is a contentious resource.
On the other hand, you can just run N independent simulations, and pool the statistics. In many cases, this is as good as if not better than running one long MD simulation. Here, all the processors share a common input file, but it's a single file, it only has to be communicated once. Any results from the processors are compacted, and sent to a large file server, which in itself is designed with partitioning since all the output files are independent. This approach won't let you work with large systems (unless your individual servers have more RAM), but it scales linearly to far more processors than tightly coupled parallelism, for far cheaper, and often, the scientific results are as good or better.
This shows up in many cases- in my 20+ year career programming, nearly all the problems I saw people run on supercomputers eventually could be run on CPUs with easier code and better performance. There are some irreducible codes for which only supercomputers are capable of meeting the required specs, and those are the only codes which should run on supercomputers. Otherwise you're pissing money on interconnect and hard to program environments).
In short: if you can convert your supercomputer job into a cloud cluster, you should do so, because it will save time and money. If you can't, then you should use a supercomputer.
Note also that cloud interconnects have gotten significantly better over the years- for me, 100Gbit throughput and latency (as long as the cluster has the total parallel bisection bandwidth) is fine for me and most other people.
/I work for Basho/
What's the evidence for this claim? AFAIK the DoE INCITE awards -- which is how you get allocated time on the big supercomputers -- are quite competitive.
Also the supercomputers don't cost billions. Aurora, for example only costs $200 million, and if it lasts as long as its predecessor Mira did, that is $200 million spread over 6 years, which is a pretty small number.
> My conclusion was most of these computers aren't needed
I don't disagree, but I wonder what is behind the investment in China. They are putting a significant chunk of resources on this and I'd like to understand why.5 Preliminary progress of scientific computing applications on the TaihuLight
5.1 Refactoring the community atmospheric model (CAM) on the Sunway TaihuLight
5.2 A fully-implicit nonhydrostatic dynamic solver for cloud-resolving atmospheric simulation on eight million cores
5.3 A highly effective global surface wave numerical simulation with ultra-high resolution
5.4 Peta-scale atomistic simulation of silicon nanowires
5.5 Large-scale phase-field simulation for coarsening dynamics based on Cahn-Hilliard equation with degenerated mobility
5.6 Application: summary and comparison
[1] http://nextbigfuture.com/2016/05/us-supercomputer-chip-ban-d...
http://www.nextplatform.com/2016/06/20/look-inside-chinas-ch...
Summary, it's a RISC architecture that resembles an Alpha but it's not a straight-up copy.
I wasn't aware that China had competitive fabrication plants for processors. If they don't and they built these at say 40 or 45nm then wouldn't this design performance be even more impressive?
> By the time I became NSA director in the late 1990s, however, the calculation was no longer that simple. We still wanted an MTOPS advantage, of course, but we were fast realizing that our preferred limits were undermining the global competitiveness of the U.S. computer industry — the very industry on which we relied for our success. It was becoming clear that the overall health of that industry was more important than any MTOPS advantage against a specific target country. We still insisted on limits with regard to places such as Cuba and North Korea, but we became far more forgiving elsewhere.
This, of course, had a powerful, positive commercial impact, but the NSA didn’t flip its position for commercial reasons. We did it for security reasons. On balance, this change made us stronger, not weaker, over the long haul, since retarding exports would inevitably retard the technological progress that was both our economic and our security lifeblood.
That early lesson has caused me to continue to challenge arguments that technological protectionism furthers national security. It might, but then again, it could have the opposite effect if it freezes development, alienates allies, feeds distrust or invites the creation of similar barriers abroad. I would recommend these broader considerations to those in the U.S. security enterprise with responsibility for evaluating these trade-offs today.
https://www.washingtonpost.com/opinions/dont-let-america-be-...
The whole post is worth reading. The same logic could be applied to US trying to put backdoors in its products for a specific national security goal, which will ultimately end up undermining US technological supremacy and national security as well. And yet the current CIA director just implied that US would be fine with backdoored products, because "where are people going to get their encryption from? The foreigners? Ha!"
This sort of arrogance, which is the same type of arrogance that ended up banning chip sales to China last year, is what will make the US lose out in the long term.
https://www.techdirt.com/articles/20160618/08022234741/cia-d...
What more evidence do you want?
http://nextbigfuture.com/2015/04/us-will-make-180-petaflop-s...
The given reason is all bullshit, and they must think we're all morons if they expect us to buy it. If anything it's retaliation for some IP theft or some silent trade war going on between China and the US. It's certainly not because the US gov is "shocked" that China uses its supercomputers to do nuclear research.
Then comes Machiavellian management strategies that demonstrate large companies are not idiots, and withdrawal of such scheme 2-5 years later, while still holding sales lines in China.
Source: Have worked on both sides of the technology transfer equation.
It might be hard for the U.S. government to effectively prevent the Chinese government from doing these simulations, but it's plausible to me that the U.S. government would still want to see the Chinese government hindered in nuclear weapons simulations.
No, seriously.
1. Intel loses this business.
2. US employees lose their job.
3. China acquires skills designing the system themselves.
4. And they get to have their supercomputer after all.
How can anyone not see how stupid these export restrictions are?
2. Augments slow CPU with impossible to programming Xeon Phi.
3. Spends significant effort achieving heroic linpack benchmark.
4. Researchers are unable to obtain real world fast compute.
5. Chinese programs for stealth, submarines, nuclear are significantly delayed. Expertise diverted into useless PR project.
6. PR backlash as US goaded to reinvest in next generation supercomputing, GPUs, etc.
Mission accomplished.
A while back I recall reading rumors that China didn't even know what to process on their Tianhe-2. That kind of makes a supercomputer look like a pile of hardware, no matter how well it's organized or engraved.
[1] http://www.hpcwire.com/2016/06/19/china-125-petaflops-sunway...
Note that K computer (#5 system, Japan) scores 4.9% for HPCG. So Tianhe-2 and Titan are again 4x worse for HPCG compared to systems which score best for HPCG.
I don't think it's as simple as LINPACK bad, HPCG good. LINPACK is representative for some workloads, when compute dominates. HPCG aims to balance compute and memory. There is Graph500, if memory dominates for your workload.
By the way, K computer is #1 in Graph500.
The interconnect on the Sunway TaihuLight seems to be standard Infiniband, from reading other articles about the system.
Kind of reminds me of when we heard a rumor than a rival ISP had bought a handful of Sun servers (in the year 2000). We thought they had tends of thousands of customers, figured they would be launching DSL, figured that had thousands of web hosting customers...
Turns out they got some weird grant so they spent the cash on those, otherwise they would have lost the money. They went out of business a year later and we kept chugging along with our Pentium III servers.
There are enough high quality universites in China to make use of it. I can not imagine this is true. There may have not been one big algorithm that it was built for, but university researchers always have stuff to compute.
I'd be interested to hear from someone who knows/has experience, whether any of the top 100 super computers ever have idle windows.
Having done exclusively HPC for almost 16 years now, I can state that there is one and only one immutable law of HPC: grad students will find a way to squeeze out every second of available CPU time on a large system. The amount of effort going into taking advantage of those empty spots in the clusters where smaller jobs can run quickly while the scheduler drains resources for a large job (called backfilling) is tremendous.
The thing that sticks with me is what a fantastically complicated problem HPC job scheduling is. I've seen dozens of fresh-faced undergrads or first year grad students come in with a full head of steam and decide that they're going to "solve" HPC job scheduling, but every single one has gotten bogged down in the minutiae and the muck, and I've never seen software that "gets it right". Every scheduler sucks in its own unique way.
Whether all of those jobs need an old-fashioned supercomputer is certainly up for debate, but they certainly get used, at least in my experience.
Or was it a combination of that with "well its there, and we want to do this thing that would happen quicker on the big machine" - IE not necessarily enabling new research, just adding a nice speedup to existing work?
So no, these machines are often used for problems which really do utilize the large-scale nature of the computers.
So lots of the "big data" type problems with lots of uncoordinated jobs aren't going to take advantage of the primary component of these big computers, the node interconnect. The computers really shine with big fluid dynamics simulations where there's lots of chatter between different components of the simulation to communicate boundary conditions and the like.
>> they certainly get used, at least in my experience
As a taxpayer, this made me feel good, thank you for sharing.
That applies to most of supercomputers everywhere. "If you build it, they will come".
Besides the knowledge attained in building supercomputers ( which all major nations should know how to do ) and processor technology, the added benefits and new uses of these machines will come later.
The internet was created before people really knew what to do with it. The same thing with computers. The same thing with everything else.
What matters most might be that China is maturing CPU technology that isn't nearly as congenial for US TLAs to hack in to. It might not matter so much for this supercomputer, but if Chinese suppliers put price ahead of things like opaque management processors, it might lead to tipping the balance away from easily exploitable endpoints.
The following paper compares the EU and US ensembles: http://journals.ametsoc.org/doi/full/10.1175/MWR2905.1
https://en.wikipedia.org/wiki/Tianhe-2
What is interesting about this new supercomputer is that it is apparently built with CPUs that were developed and produced by the Chinese themselves.
New Supercomputer: https://en.wikipedia.org/wiki/Sunway_TaihuLight
CPU architecture: https://en.wikipedia.org/wiki/ShenWei
Sierra and Summit are coming out soon, which are POWER based.
https://www.olcf.ornl.gov/summit/
http://www.cnet.com/news/ibm-nvidia-land-325-million-superco...
According to TOP500 author Jack Dongarra, three scientific simulation codes run
on TaihuLight have been chosen as Gordon Bell Prize finalists, two of which have
managed to reach a sustained performance of 30 to 40 petaflops. The award is
bestowed each year on the most noteworthy HPC application, based on “peak
performance or special achievements in scalability and time-to-solution on
important science and engineering problems.”
Source: http://www.top500.org/news/china-tops-supercomputer-rankings...Is there plans to commercialize these chips? We need accessible low cost ultra-high core count chips widely available. I feel that Intel and AMD have been under performing in this area (growing core count) for the last decade. And Intel has kept the price of CPUs exceptionally high for the last decade as well.
If low cost ultra-high performance chips from China under cut Intel's excessively priced Xeons (+$3K per CPU), it really could change things. If these start to spread, I wouldn't be surprised if Intel tried to prevent the spread of these CPUs outside of China via trade barriers.
Compared to Tianhe-2, the previous top system, 2.7x more flops, and 3.2x more flops per watt.
So that Chinese number seems awfully low to me. I would expect that the number Amazon has to be higher, for instance. Same for Microsoft and Google. Therefore I'd be amazed if the DoD, NSA and even the DoE wouldn't have more capacity available.
When I ran Exacycle, we distributed protein folding, protein design, drug discovery, telescope design, and other problems globally. We never had an issue with latency, because these problems all partition really well. People who claim supercomputers are "necessary" for these problems typically construct problems that are well-matched to supercomputers (for example, running molecular dynamics on huge proteins) but they tend not to have very high scientific value.
In my experience, partitioning to minimize communication has always increased my total scientific throughput, while programming to supercomputers has always reduced it.