Cray, AMD to Extend DOE’s Exascale Frontier
hpcwire.com
hpcwire.com
I wonder if that will be the case on Frontier.
I feel these things always go back and forth in cycles in the industry.
"fully coherent single address space with all memory available to [all processors]" is a lot easier if there is no virtual memory, no caching, and all processors are on a single shared bus to the same memory chips.
> Single-socket nodes will consist of one CPU and four GPUs, connected by AMD’s custom high bandwidth, low latency coherent Infinity fabric.
The article, at least, only describes coherency in the context of single-socket nodes. I suppose it's ambiguous whether Infinity fabric only connects components on the node or connects nodes to each other, but in fact Infinity fabric is the interconnect used with Zen CPUs, so I think it's pretty clear that in this case it only connects and provides memory coherency for 1xCPU and 4xGPUs.
Performant memory coherency across nodes would be quite an achievement, but not one they seem to have achieved. Rather, it seems they've simply revived the PS4 model.
Would you mind pointing me in the right direction to learn about how to do this? (something equivalent to slide 2 of https://www.olcf.ornl.gov/wp-content/uploads/2018/12/summit_... but for mobile/APU)?
Mostly my point was the unified system/gpu memory is pretty common outside of the desktop space.
However, maybe this is just me, but I don't completely trust this to work without losing a certain amount of peak memory performance. I hope they at least leave the option to turn it off, so we can verify the impact it has on a per-application basis.
So 30 megawatts of computing, plus cooling and other supporting services. How do you power something like this? Does ORNL have their own power station (given they have reactor(s) on site)? If power comes from an external station do they coordinate with the station operator when bringing a system like this online?
I could find the numbers for two German super computers I have used in the past. SuperMUC (Phase 1 and 2 combined) had "above the desired 85%" utilization in 2017 and Hazel Hen at HLRS in Stuttgart reported a utilization "between 92% and 98%" in 2017.
https://blogs.mathworks.com/cleve/2013/06/24/the-linpack-ben...
Different programs consume dramatically different amounts of power.
We are definitely "doing something wrong" when it comes to artificial neural networking; even though are models are much simpler, it still takes an enormous amount of computing power, both in terms of raw CPU as well as actual electrical needs, just to be able to simulate things at a small scale (and if we use more accurate models, based on what we know about the brain and neurons, then at best our simulations can only be run to simulate, over our actual-time, what would be in actually fractions of a second in real-time).
That our brains can do so much using so little power (wattage), with such a high number of nodes and interconnections that dwarf anything we've so far have managed to simulate - it's a bit mind-boggling and humbling.
I just wonder where and what the issue actually is.
Why do our current practical models of a neuron, which are vastly simplified, require so much power to run at scale?
Is the issue related to the fact that they are simplified models, and actual neurons with their complexity are able to do things we don't yet understand or know about?
All of this is also related to back-propagation; such a thing doesn't seem to exist in nature (jury is still out on the theory, though) - so how do biological neural nets "learn"?
If we could eliminate or reduce the need for backpropagation, would that lower our power requirements for artificial network implementations?
As someone who has merely dabbled with artificial neural networks, these questions and conundrums fascinate me, and cause me to attempt to think up potential solutions, however far-fetched.
I highly doubt I will be the one to solve the issue, but I do hope to see it solved within my lifetime.
Someone might also find that kind of network can run a KHz like human brain cells instead of MHz, GHz. The power usage will go down even more.
At any rate, it is quite likely that the neocortex of the brain simply computes a function(s) recursively upon sensory input (see Chomsky's minimalist program for suggestions on what it could be). What is unclear is what this function is, and how it comes about - for this reason the approach by some has been to attempt to simulate an entire brain to see what it does. But without the necessary abstractions, this will be inherently wasteful and generates nothing new other than validating your experimental data.
neural networks are only a small component of what we use supercomputers for.
TVA recently completed a 210 MW substation on ORNL's campus to better serve our needs. We do not need to coordinate with them for large runs on the machines.
In reference to AI, NVidia has things "locked up" with CUDA, versus 2nd cousin AMD's OpenCL.
From what I understand, it is possible to recompile TensorFlow (for instance - not that ORNL will be using TF) for OpenCL - but I don't know how well it works. Personally, I've only used TF with CUDA.
Does this mean we might see greater/better support for OpenCL in the AI realm? Might we seem it become on-par with CUDA because of this collaboration for this HPC?
Or will things stay as-is, at least "down here" in the consumer/business realm of AI hardware and applications? Do things like this trickle down, or are things so customized and/or proprietary for the needs of HPC at ORNL (or elsewhere) that anything to do with AI on this machine will have little to no bearing outside of the lab?
Ultimately, I'd just like to see another choice (a lower cost choice!) for GPU in the world of consumer/enthusiast/hobbyist AI/DL/ML - while today's higher-end GPUs, no matter the manufacturer, tend to be fairly expensive, AMD still has an edge here that make them attractive to users (not to mention the fact that their Linux drivers are open-source, which is also a plus).
"Exascale Deep Learning for Climate Analytics" won the 2018 Gordon bell using summit.
There should be more information at https://www.olcf.ornl.gov/leadership-science/
From your link:
"New Frontiers for Material Modeling via Machine Learning Techniques"
"Large scale deep neural network optimization for neutrino physics"
"High-Fidelity Simulations of Gas Turbine Stages for Model Development using Machine Learning"
"HPC4mfg – Reinforcement Learning-based Traffic Control to Optimize Energy Usage and Throughput"
"Advances in Machine Learning to Improve Scientific Discovery"
ML (not AI) will probably end up driving a lot of supercomputer workloads because distributed ML training resembles modern HPC simulation codes.
http://www.doeleadershipcomputing.org/proposal/call-for-prop...
In reality, most HPC sites use high performance interconnects, such as infiniband or a proprietary high performance ethernet. They are also less likely to use virtualization, and every node looks 'bare metal'. The software stacks are very different, from everything between the distributed memory model, compilers, to the system schedulers and diagnostic software.
Nevertheless, there is a great deal of convergence happening between HPC and hyperscale data centers, particularly as hyperscale uses more machine learning, which has a similar flavor to HPC. Many believe that the FANG companies have exaflop capabilities already, but they just aren't well optimized for scientific workloads.
https://www.anandtech.com/show/14302/us-dept-of-energy-annou...
What interconnects do these sorts of machines use? I assume even 100GbE isn't enough?
Just curious. It's interesting what exists in the "so far beyond my price range as to be ludicrous" category.
~ 37 GFLOPS/W is probably a better projection if we assume (out of nowhere) that the theoretical/achieved ratio of Frontier is comparable to Summit (75%). Still very impressive.
Jaguar | 2.3 PF XT5 (CPU-only) | 7 MW for HPL
Titan | 27 PF XK7 (1:1 CPU to GPU) | 8.2 MW
Summit | 186 PF (2:6 CPU to GPU | 8.8 MW
Overall, substantial changes in computing power and 10%-20% increase in power.