UK to invest £900M in supercomputer in bid to build own 'BritGPT'
theguardian.com
theguardian.com
>Write a short story about springtime in England. Use British English spellings when doing so.
GPT-4 went off and wrote me a book, and I don't actually know that many Brit vs US spellings, but I do see things like
...fragrant blooms and preparing colourful bouquets for the villagers....
Feels very “me too” without much of a plan frankly
Last year it was a flagship yacht, this year BritGPT. Just stick a Union Jack and 'Brit' somewhere in it, and call it a day.
The country is going through an economic and existential crisis, with amongst the lowest productivity stats in Europe, the shitshow of Brexit impacting everyone, healthcare crisis and national strikes due to high inflation.
Truthfully the UK economy is held together by a handful of dodgy banks in London, and if these left, we'd be left as a borderline third world country. The country essentially makes nothing and is fuelled entirely by debt and consumption. We've become a nation of hairdressers and "luxury" cafes. Electorate is generally stupid, so none of this will change anytime soon.
For example, a typical DGX supercomputer system from NVidia is pushing 7.2TBps (https://www.nvidia.com/en-us/data-center/dgx-h100/) GPU-to-GPU communications.
In contrast, a typical DDR4 RAM on your typical desktop is 0.05 TBps, so yeah, 7.2TBps external bandwidth between GPUs is quite a lot.
-----------
For Frontier, the Slingshot NICs are 100GBps each. So each node can communicate with more bandwidth to each other than your typical Desktop computer has RAM-bandwidth.
The diagrams imply that there's 4x Slingshot NICs per node on Frontier, suggesting 400GBps bandwidth to the interconnect. (https://www.olcf.ornl.gov/wp-content/uploads/2020/02/frontie...)
You mixed up GBps (gigaBYTES per second) with Gbps (gigaBITs per second)- divide your numbers by ten to get a good idea of gigabytes for comparing to RAM.
RAM is about 50GBytes/sec, fast individual nics are 400Gbps (or about 40GB/sec). Unless you have special caches or very large ram, you will run out of bits to send very quickly on a network like that.
Typically supercomputers don't have RAM or busses that are more than 2X faster than conventional machines because it's not economic.
> 7.2 terabytes per second of bidirectional GPU-to-GPU bandwidth, 1.5X more than previous generation
And...
https://www.olcf.ornl.gov/frontier/
> Multiple Slingshot NICs providing 100 GB/s network bandwidth. Slingshot network which provides adaptive routing, congestion management and quality of service.
That's not "little-b bits", that's big-B Bytes.
> Typically supercomputers don't have RAM or busses that are more than 2X faster than conventional machines because it's not economic.
These are top-end GPUs equipped with the latest-and-greatest HBM RAM, with 2TBytes/sec RAM bandwidth or more. They put the super in "supercomputer".
When you have 1TByte to 2TBytes of RAM bandwidth, you need 400GByte+ network connections to keep up with them.
But "multiple slingshot NICs providing 100GB/sec"- that's ganged NICs, each one only does 200Gbps. It looks like you're conflating what's available to a single GPU, versus cross sectional bandwidth? I had to spend some time looking at the Slingshot docs because they throw around a bunch of numbers aggregated over different dimensions, leaving out the number of items they're aggregating over (common technique to make things look "bigger").
Another big issue that threw me off is that you're comparing the GPU network on the hosts to desktop computers. But in thinking, I think that's fair since the GPUs are the (primary) processing elements.
The cluster that MSFT build for OpenAI is definitely a supercomputer.
UK had its chance to be a tech powerhouse, and wasted it. Nobody ever cared about why.