Supercomputers: Obama orders world's fastest computer
bbc.co.uk
bbc.co.uk
It seems to me that the aim of taking 10 years to build a supercomputer that is only 20 times faster than the current one might fall a little short if it's aiming to take the top spot.
1. http://www.theplatform.net/2015/07/13/top-500-supercomputer-...
2. The dominance of China is achieved through the use of Xeon Phi accelerators, which may be great for Linpack, but have not made much of a splash for applications yet. GPUs are solidly beating Intel's accelerator offering both on adaptation and performance.
I have have read last 10 years that Processing is a lot more cheaper did by network of computers and clusters, instead of a expensive supercomputer that also demands an appropriate building and infrastructure.
However, the types of computation performed by the top supercomputers are rarely the "embarassingly parallel" programs you can easily distribute via an @Home-style program, or something like Hadoop. They do depend heavily on very reliable, very low latency, high bandwidth networks.
x20 ought to take about five years and a half.
Also the FLOPS measurement is a bit broken: It focuses on dense linear algebra problem, for which GPU or other accelerators boost the results easily. If all you plan to do is running simulations that are easily parallelized on GPU it is fine, for other types of programs it is hard to tell which is the fastest supercomputer.
* France is also ordering a would -be top 10 supercomputer : http://www.hpcwire.com/off-the-wire/the-cea-agency-and-atos-...
Does anybody have a better measurement?
Perhaps it could be the size of the matrix that can be inverted on it in an hour of time, with IEEE double precision floats, using some standard algorithm.
But what if you want to optimize for programs that are communication intensive, or memory intensive?
Should the FLOPS of a very specific linear algebra suite be used as the metric of best computers?
The different performance numbers for top systems on HPCG vs HPL are pretty striking: http://www.hpcg-benchmark.org/custom/index.html?lid=155&slid...
Original proposal to use HPCG as an alternative to HPL for supercomputer rankings: http://www.sandia.gov/~maherou/docs/HPCG-Benchmark.pdf
The reality is that there are many dimensions to supercomputing performance and it's impossible for one number to capture the utility of the machine. Our HPGMG benchmark (https://hpgmg.org) attempts to strike a balance and give useful supplementary information. I do think it's better than any other single benchmark for evaluating today's machines and will also prove to be more durable over time.
With respect to deep learning and other applications using CG or related algorithms, the bottlenecks depend on the scale, and ability to expose locality, and operator/preconditioner representation. If there is no locality, then matrix-vector products require all-to-all communication which tend to dwarf the cost of the reductions in CG. Even with locality in the matrix-vector product, preconditioners often need to communicate globally in a scalable way similar to HPGMG. Operators need not be represented as a table of numbers or a sparse matrix format, but could use a tensor product, fast transform, or other information to compute the action using less storage. If they are represented explicitly (sparse or dense), then matrix-vector product performance (thus CG as a whole) is dominated by memory bandwidth for problem sizes that do not fit in cache. HPGMG tries to strike a balance between memory bandwidth demands and compute using a matrix-free representation. HPGMG also reports dynamic range expressed as Performance versus Time-to-solution as the problem size is varied, which allows applications to see performance barriers that might be relevant to them (e.g., see how Titan cannot do a solve in less than 200 ms while Edison can do 50 ms, and how that relates to climate simulation performance targets; see slide 7 of https://jedbrown.org/files/20150624-Versatility.pdf).
What I meant by asynchronous is that not all terms in a gradient are required to be summed in the same step.
The transpose step in Sibyl is implemented in the Shuffle and Reduce phases. The filesystem is used to hold the temporary data. Nevertheless even for large systems, very few steps are required, and step times are reasonable, even compared to modern supercomputers. This is a tribute primarily to the design of sibyl and the implementation of MapReduce at Google.
This is all explained in online versions of the Sibyl presentation. I really wish more people from DOE who write modern solvers would pay attention to this stuff.
One of the biggest reasons for use of HPL is that many sizing considerations can be based off of the theoretical calculations.
But anyway this is very interesting. I definitely need to check this out.
HPGMG is representative of most structure-exploiting algorithms in that it does not have this abundance of flops, thus theoretical performance is actively constrained by both memory bandwidth and flop/s. We see many active constraints in practice; e.g., improving any of peak flop/s, memory bandwidth, network latency, or network bandwidth produces a tangible improvement in HPGMG performance. Depending on the fidelity of the performance model, these dimensions can be a fairly accurate predictor of performance, but ILP, compiler quality, on-node synchronization latency, cache sizes, and similar factors also matter (more for HPGMG-FE than HPGMG-FV).
I think it is actually quite undesirable for benchmark performance to be trivially computed from one parameter in machine provisioning. No computing center has a mission statement asking for a place on a benchmark ranking list (like Top500). Instead, they have a scientific or engineering mandate. Press releases tend to overemphasize the ranking and I think it is harmful to the science any time the benchmark takes precedence over the expected scientific workload. HPGMG is intended to be representative in the sense that if you build an "HPGMG Machine", you'll get a balanced, versatile machine that scientists and engineers in most disciplines will be happy with. I'd still rather the centers focus on their workload instead of HPGMG.
[1] https://www.whitehouse.gov/sites/default/files/microsites/os...
A huge amount of overall system power is spent in data transport. Plus, double everything for cooling. That brings the total system efficiency way down from what the actual computational components spend.
The best next generation DP GFLOPs/watt from one of the big players will most likely be the 2016 Xeon Phi, at ~10-12GFLOPs/watt... You are also forgetting that GPUs also have a ~100W+ CPU sitting next to it, which brings down total efficiency significantly.
Shameless self promotion: My startup (http://rexcomputing.com) is aiming for 64 double precision GFLOPs/watt, and 128 GFLOPs/watts single precision for its first chip next year.
This is because the supercomputer community has long ignored the Internet-style of computation (MapReduce etc). But most of the new generation of scientists are adapting their codes to this new style, because dollar-for-dollar they can get more throughput than the classic style machines. Classic machines invest heavily in low-latency communication and typically require APIs like MPI to achieve it, while Internet HPC just uses well-designed TCP-based socket communications.
Building dual-design systems like this- especially when the community has little or no skill at building NG Internet HPC systems- is likely to produce a system that is good at few things.
Instead, build two systems. One is the largest (but not necessarily exaflop) you can afford and is a classic supercom[puter. Then, for the second, hire some datacenter designers from Google/Facebook and have them build a modern HPC cloud design.
The biologists will flock to the second one; they have long been underserved by the DOE supercomputing community.
I also helped create the original idea for the Broad collaboration. It was pretty obvious that standard HPC was a waste of money and not designed for high throughput biology, while Google published numerous papers (MapReduce, GFS, Bigtable) that demonstrated they were building infrastructure that was perfect for a wide range of computational biology programs.
I would argue that the community of people who actually have the skills to take advantage of the interconnects in a classic HPC system is vanishingly small, and in consequence we've overbuilt them on an epic scale.
Allow me to vent. I had the good fortune to have a login on a "petascale" HPC system, and access to an allocation of hours.
The /scratch filesystem would fail weekly, which killed everybody's jobs. If you had a big run going when /scratch failed, you lost everything. Scratch failed so much because the models that were being used often did wildly inappropriate amounts of file IO --- debugging print statements, detailed intermediate calculations, excessively verbose output --- that worked all right in development but when run in parallel brought the filesystem to its knees.
Furthermore, the login nodes were almost unusably slow because of all the Python and Perl post-processing scripts running on them. This isn't even a matter of users being cheap with their hours --- post-processing would have been a tiny fraction of their allocations. Instead, it's that many of them gave no thought at all to how the post-processing might be structured and run through the batch scheduler, and saw no downside to abusing the login nodes for that purpose.
In conclusion, I can attest to at least one HPC system that was badly mismatched to its users' needs and level of sophistication, despite allocations of hours being awarded only to a small number of researchers from across the country through a highly competitive process. Building these things serves national and institutional pride far more than any utilitarian interest.
My claim is that the design of classic interconnects is a big waste of money, because only a few codes need it, yet the cost dominates (>50%) of the cluster. I've learned, from years of studying Google's papers, that there are better ways to build code that communicates, and those mechanisms are much easier to teach to scientists and computer scientists than MPI.
Here is my argument: when I worked for DOE, everybody told me I had to run my MD simulations on a super computer using all the processors, and I would judged on my parallel efficiency. This meant using a code that used MPI to communicate at every (or every N) timesteps. I asked, instead, "Why not just run N independent simulations, and pool the results?" In this case, you run an M-thread simulation on each machine (where M = number of cores on the machine) with no internode communication at all except to read input files and write output files.
The short answer is, that approach works just fine, but the DOE supercomputer people won't let you run embarassingly parallel codes because they already spent money on the interconnect to run tightly coupled codes.
In reponse to this, I went to Google, built Exacycle (loosely coupled HPC) and published this well-cited paper: http://www.ncbi.nlm.nih.gov/pubmed/24345941 which in my opinion put the last nail in the coffin of DOE-style physics simulations for molecular dynamics.
That said, there are systems which are so large you can't practically simulate a single instance of the system on a single machine, so you have to partition. Simulating the ribosome is a nice example. However, simulating the ribosome currently provides no valuable scientific data except to tell us that we have major problems with our simulation systems (force field errors, missing QM, electrostatic approximations,e tc).
Eventually it reached the point (~2007) where I could fit the whole simulation on a single 4-core Intel box with similar performance. Then, I ran one "task" per machine, and scaled to the number of available machines. This uses only inter-node communication, which goes over a hub or crossbar on the motherboard. Much faster.
Now, I can fit many copies of DNA on a single machine (one task per core). This is far and away the best, because each processor just accesses its own memory, greatly reducing motherboard traffic, so the problem is basically CPU-bound instead of communication bound (this also now applies to GPUs, such that single GPUs can run one large simulation within its own RAM and not have to spill data back and forth over the CPU/GPU communication path).
This moves the challenge to the IO subsystem- I generate so much simulation data that I need a fat MapReduce cluster to analyze the trajectories.
I'm not just describing strong scaling. I'm describing a cost-effective way to achieve it; that's what really matters.
Why have subsets of nodes for post-simulation cleanup? Why not just run that cleanup on the same nodes you used for simulation? Or other general nodes? Otherwise, you've got two sets of nodes which are used at lower utilization than they would normally be.
This is because 50% of the cost of the machine was the interconnect, and if they let those codes run, it means they wasted budget and will get less next time.
Until I hear that the funders/builders are spending the same amount of budget on machines that let biologists run embarrassingly parallel codes as they spend on TOP500 machines, it's not going to change.
The abuse of /scratch and the login nodes sounds like a classical mixture of not knowing, not caring, and limited time. That isn't something that has a technical solution.
The other problem is that their interconnect is relatively fragile. It's comparatively easy to crash the entire network, at which time your filesystem goes away and processing stops.
But thanks to Lustre, even when it's good, it's bad.
Many argue that an exascale computer can only be cost efficient if the communication capabilities scale highly sublinearly with the computation done in the subsystems [1]. In particular, you can't move the data, and new algorithms are needed that can deal with data that is arbitrarily distributed. This is quite challenging and unfortunately the theoretical computer science community seems to have decided that distributed memory algorithms have been covered since the 90s and are not worth their time. Yet they ignore the progress that has been made in other models of computation since, and many algorithmic improvements of the last decades are not applicable. It is high time to develop communication-efficient algorithms for the basic "toolbox".
I guess what I'm trying to say is that you can't just throw MapReduce at an Exascale machine and expect it to perform well. Instead, you need an environment that is rich in primitives that have been implemented in a communication-efficient way. It's faster and cheaper to spend a little more effort on local communication if that allows for reduced communication volume (and/or the number of connections that need to be established!).
The issue I have with the MapReduce approach is that it doesn't particularly care about data locality. Thus it is very hard to achieve communication volume sublinear in the input size, which is absolutely deadly in an exascale setting.
I also understand the frustration with MPI, it is a very low-level API focused on data movement. It can be rather frustrating to use, but there do exist tools to make it more fun (Boost.MPI with C++11/14 is an excellent example). That said, with a well-engineered set of algorithmic tools, ideally you wouldn't need to use low-level MPI calls at all. However, MPI still remains a useful tool to implement these things.
Exascale computing requires us to rethink a lot of things.
[1] http://www.ipdps.org/ipdps2013/SBorkar_IPDPS_May_2013.pdf Shekhar Borkar (Intel), Keynote presentation at the 2013 IEEE International Parallel & Distributed Processing Symposium
Also, most modern Internet HPC systems dedicate a ton of design and equipment to having very high cross-sectional bandwidth, which enables the locality restrictions to be relaxed.
See also this paper http://research.google.com/pubs/pub36740.html
The main challenge is that because these are built with multistage routers, they have fairly high latency. So much of the effort in modern HPC systems used for Hadoopy workloads goes to latency hiding.
What's important to recognize is you simply cannot buy Infiniband switches that let you contact a lot (10K+) of hosts together. The vendors won't sell you this, they won't do the R&D to make it, and it would cost infinite anyway.
This is a deliberate choice: for most Internet work, it's better to have really fat bisection bandwidth and non-blocking fabrics, and latency is ignored due to the high cost of building a crossbar that supports that with high radix.
Only if you have an algorithm that absolutely requires, and simply cannot be fixed, low latency, you are almost always better off building a cheaper, fatter fabric, and hiring engineers who know how to write applications that are latency tolerant.
Massively distributed systems do not get much benefit from low-latency interconnects. Massively parallel systems do, and in particular, it is a "throw hardware at the problem" kind of solution that helps cover for the fact that virtually no software designers can engineer efficient, non-trivial, massively parallel systems. MapReduce is a distributed model; outside of some trivial cases, it is a poor parallel model. And while the HPC community has a much better understanding of massive parallelism than the Internet-scale systems community, the HPC community largely doesn't grok massively distributed systems in the way that someone working on Google's infrastructure would.
I benefitted from having spent several years designing software for both HPC and Internet-scale systems. They are not fungible, and both communities grok things that the other is oblivious to. Even within the HPC community though, the number of people skilled at the design of massively parallel software systems is quite small, much smaller than people that know massively distributed systems.
You do not need two systems, you need one system and more people that have figured out how to design massively parallel software -- the real problem. It is difficult to overstate just how rare this skill is even within the HPC community.
So given Moore's Law, by the time it's finished in 2025, it will be 50 times slower than 2025's fastest?
Yes, yes, Moore's Law is slowing, transistors on a chip =/= flops, etc. Still seems like they'd want to aim higher than 20x in 10 years.
Therefore, even if the raw performance in terms of FLOPS sound similar, the two systems will have widely differing performance on real workloads.
Capturing and indexing the entire web is certainly a real workload, even if it is massively psrallelizable, so it would probably run equally well on Google's infrastructure as on a supercomputer because those fast interconnects wouldn't provide much advantage, right?
However,when simulating a nuclear explosion or a weather system (maybe that's what you mean by "real" workloads?), the heavy node-to-node communication makes the supercomputer much, much better suited.
The NSA's currently being sued over its metadata collection. I wrote that comment slightly tongue-in-cheek, but repurposing that supercomputer for civilian use would actually be a great way to recoup some of your losses.
Edit: also should say it is amusing to hear Obama talking about exascale conputing by 2025 when the NSA's goal (read the Wired article) is to get there by 2018.
Now I am curious what is the fastest single processor?
Mechanical CAD geometry kernels are one such application. Recently, I had a use case that demanded the peak single threaded performance.
In PC land, it's this chip clocked at 4.x Ghz: Intel® Core™ i7-4790K. It's multi-core performance is pretty great too, so it's not that big of a trade-off to maximize single thread performance.
I would be very interested in knowing what faster solutions exist. Are there any, regardless of instruction set?
I assume that when most people think of 'single core' performance, they really mean 'single thread'. I think Intel's CPUs win out there, based on your linked benchmarks.
What does your setup look like? Cooling, RAM, etc...?
And is it good for long term, like say crunch on it for a week type problems?
No problem with stability, it ran a few times over weekend at full load. Maximal temperature about 90C. I have put quite high voltage. Over summer it goes down to 4.8GHz.
Exascale is power-hungry so power must go way down and efficiency of calculation way up.
http://hwbot.org/submission/%202615355
I happened to catch an overclocking competition being streamed on Twitch one late night many months ago. It was really interesting to see more about the methods and techniques involved and how the competitions work.
edit: Actually, in the scenarios you'd use a supercomputer for, the added latency and overhead (shoddy servers, network, etc.) would most likely make the run time orders of magnitude higher.
There are many existing valuable codes written in FORTRAN. They work, it's not worth the investment to replace them with something else.
Second, many of the codes are in C++, not FORTRAN. Not clear that's any less of a problem.
By codes they mean -- at the minimum -- pretty much anything that requires frequent communication between any or all nodes as a necessary part of computation. (For example, simulations across a large 3D space, where the changing states of particles on node A directly impacts the states of particles on adjacent nodes.)
Also, there is a wide range of literature about communication patterns for supercomputer apps; my argument is that often times, to solve the problem that matters, you may not actually need to run the simulation you think you do. It's more that people are just used to running that way.
For example, with MD, you can run 1 sim parallelized over 100 machines using tightly coupled communication (doesn't necessarily mean the forces and positions of every particle have to be shared between node decompositions) or run 100 sims over 100 machines, with no communication except for input and output files. The latter can often answer the same question far more cheaply.
I don't want to drag this out, but where do you see the language constraint? You need an MPI binding, sure, but what else?
Supercomputers aren't built so that people can squander the resource (desktop PCs, closest clusters, and phones fulfill that role).
Anyway, the issue with JVMs is that they don't have predictable performance, not that the compilers can't be ported.
There are still some fortran libraries in large scale use for this sort of thing. They are still in use because they are very good, and replacing them would be very expensive for little gain.
U.S. and other countries have been in a race for exascale. The thing holding us back isn't funding or political will: exascale is so ridiculously hard that it requires fundamentally different architectures. The main issues are making our CPU's do more work, eliminating memory bottlenecks, and dramatically improving energy efficiency of both. It's just very tough, technical challenges that might also have to operate on process nodes that are themselves tough.
Rexx Computing is one attempt whose founder posts here a lot [except in one thread dedicated to it lol]. I'm curious if any other exascale researchers read HN and can post their concepts as it's probably interesting stuff. Here's some links for readers interested in this stuff.
LLNL gives data on exascale and its challenges https://asc.llnl.gov/content/assets/docs/exascale-white.pdf
Also describes problems but skip to Venray's TOMI approach http://www.edn.com/design/systems-design/4368705/The-future-...
Rexx Computing's approach http://www.theplatform.net/2015/03/12/the-little-chip-that-c...
Intel's relatively conventional approach http://www.exascale-computing.eu/wp-content/uploads/2012/02/...
Architecture from Univ of Texas and NVIDIA https://www.cs.utexas.edu/users/skeckler/pubs/SC_2014_Exasca...
Boise exploring non-Von-Neuman with ParalleX http://cswarm.nd.edu/news-events/assets/PSAAP_II_Kick-off_CS...
Same group enlightens on details that all fight with http://sites.ieee.org/boise-cs/files/2015/04/Thomas-Sterling...
Bonus: 1,000 core, cache-coherent, optical interconnect. Sort of thing might be useful in exascale. http://dspace.mit.edu/openaccess-disseminate/1721.1/67490
Have fun with these. Submit a link if I left out any chip architecture in exascale race that's pretty cool.
Notice step 3 implies international stability and hence profit. NNSA stresses computation so heavily because stockpile stewardship cannot be done by noncomputational means.
Thanks.
I checked the list of projects running on ALCF and I see some pretty obvious nuclear weapon stockpile stewardship and weapon design projects, such as "Validation Simulations of Macroscopic Burning-Plasma Dynamics"
(I know a guy that Los Alamos is trying to recruit to do some dark work for them, and he does nucleosynthesis in neutron star mergers.)
I think the folks simulating supernova and neutron stars have a lot of physics overlaps, but I don't think that data is used directly for stockpile stewardship.
Mostly non embarrassingly parallel problems, where its high speed interconnect pays off.
Otherwise data parallel is an other phrase that is used for the same concept.
Of course in an ideal world those cycles would be used to help cure cancer, but given that these warheads exist, it's probably a good idea to invest resources into getting an idea of what shape they're in.
We went to the moon because of the Cold War. The military paid for the development of the Internet. GPS? Supersonic flight? Nuclear energy? Autonomous vehicles?
I wish private industry would do more. Every company should have a Bell Labs.
He parts with: "I wish private industry would do more." That seems to suggest his main point is large R&D spending, not the virtues of the Cold War.
http://airandspace.si.edu/exhibitions/wright-brothers/online...
Rather than tell us how much you hate the military and capitalism maybe you can figure out a better way?
I'd love to see more non-military funding. Perhaps medical research could be used to increase public interest in funding supercomputer research?
Well, knowing a few scientists, I can say that I know well the creative and inquiring spirits that advance us. The world is structured so that the people who do the actual development and discovery have to sell their talents to the logic of property and capital, either directly, or to the government that enforces that structure.
So, I think we're probably in agreement on the forces wherefrom those technologies came. I think we're also in agreement on whose behalf those forces act. I think we're probably in disagreement that capital is any better a master than the government it is in collusion with.
Being conceptually slippery enough to get out of any problem comes naturally to skillful people in their fields. The situation at hand provides all the impetus for theory and practice there is. Conflict thrives off of superfluity: superfluous methods, superfluous justifications, and superfluous issues. Methodology and structure gets in the way of skillfulness.
I know that real human power scales on its own without armies and guns forcing it into a certain shape. People collaborate and collude very easily. People form groups very easily. People associate with people. It's probably the only culturally universal thing people do (well of course; culture only makes sense when there are associating people).
If we want association to work well, groups need to be able to disintegrate as easily as they come together. People need to be slippery, too. Without what makes groups cohere, they disintegrate on their own. When you see violence and hierarchy used to keep groups coherent, it means they've lost the essential power that makes them useful.
People like achieving status. Status is a social signal, and it communicates both ways. Status is not held like a title. Titles wax and wane in the status they confer just like every other human object and activity. When we pretend status can be held, that it can be concentrated and preserved, we get problems. If we let status come and go by its own logic, we'd have no problems with status. A society where status is maintained by a legal system will have hierarchies, and it will certainly have problems.
So ways that our current systems-- not just capitalism, fail: Labor is not free. Association is not free. Status is not free.
Even skillful people who believe in property as an organizing force wish to be locally free in these ways. Skillful people want to work unimpeded by the politics of labor, associating with likeminded people unimpeded by the politics of social groups, praised for the inherent virtue of their actions and unimpeded by the politics of status. They only believe in property because it helps them manage the world that is beyond their control, beyond the means of their skill. They use property to create a bubble in which they can live that is free from its control.
People who believe in property because they enjoy its logic, enjoy the wheeling and dealing, enjoy the frantic rush to get more of it, who feel more worthwhile the more property they have are nervous people, who can never be fulfilled, because property's logic does not lead to fullness, the only conclusion is 'not enough', the only purpose is 'more'.
I don't know how many people in the latter category actually exist. I suspect enough to cause a great deal of problems. In any case they require the assent of everyone else. I think all I have to say to everyone else is this:
The world without property is still full and whole, is still just, is still full of vitality, is still nourishing, and receptive to human power.
Yes, a world without nculear weapons would be nice. Unfortunately that ship has sailed.