LUMI to become the world's fastest supercomputer
lumi-supercomputer.eu
lumi-supercomputer.eu
> When LUMI’s operations start next year, it will be one of the world’s fastest supercomputers.
> The peak performance of LUMI is an astonishing 552 petaflop/s meaning 552 *10^15 floating point operations per second. This figure makes LUMI one of the world’s fastest supercomputers. For comparison, the world’s fastest computer today (Fugaku in Japan) reaches 513 petaflop/s and the second fastest (Summit in the US) 200 petaflop/s
But good to see heat used for district heating, also that's probably the first supercomputer which will not be used for calculating aging atomic weapon yields.
edit: first time I heard about not calculating aging atomic weapon yields. Thanks.
Users would need to be incentivised to install distributed computing software, but I think it has promise.
[0]: https://archive.vn/20200412111010/https://stats.foldingathom...
LUMI has 200Gbit connections between nodes, or roughly 25GBytes/sec: faster than PCIe 3.0 x16 (15.8 GBps).
In effect: supercomputers can share "remote memory" as if it were local (RDMA protocol). As such, you can treat the entire RAM-space as if it were unified (your 64-bit pointers can be unified across the entire supercomputer, your data-structures can be distributed and always accessed through a 200Gbit-pipeline).
--------
As it turns out: you need a very high I/O connection to truly sustain supercomputing workloads. A lot of these things turn out to be just crazy big matrix-multiplication problems that require a fair amount of coordination between all nodes.
You can't share a problem like that on distributed compute. At best, you can only share problems that can fit on one machine (under 32GBs of RAM). In contrast, these supercomputers can work on 100+ TBs of shared-RAM problems with 100,000+ TBs of shared storage (such as simulating quantum effects). The shared storage is accessed at 2TB/s speeds, and is accelerated with Flash SSD cache layers.
---------
As some people say: the job of a supercomputer is to turn everything into an I/O constrained problem. As such, a HUGE amount of money is poured into making I/O as fast as reasonably possible. You don't want your PFlop-scale machine to be throttled by slow storage or communications.