Comparing the 1970's Cray-1 supercomputer against the Raspberry Pi
blog.adafruit.com
blog.adafruit.com
[1] http://www.roylongbottom.org.uk/Cray%201%20Supercomputer%20P...
Cray-1 vs Raspberry Pi - https://news.ycombinator.com/item?id=38758355 - Dec 2023 (180 comments)
[1] http://www.roylongbottom.org.uk/Cray%201%20Supercomputer%20P...
A 4090 today is roughly 500,000 times faster, which means we now have achieved one Cray per pixel (!) for an 800x600 image (smaller than images today, but maybe a bit larger than the average image size in the late 70s).
Does anybody have a reliable link to the Whitted quote?
https://www.youtube.com/live/LUFp6sjKbkE?si=8vcxo-Vp8oeRUnob
Scrub to 3:52:19 for the Turner Whitted story.
BTW, mostly unrelated, but scrub that video to 5:47:10 for an amazing talk by Ivan Sutherland (“father of computer graphics”) that is not about graphics (he politely refuses to talk about graphics anymore :P), but about the active research he’s been doing (at 85 years old) into Single Quantum Flux circuits (an alternative to CMOS).
Also- Jim Clark at 4:48:40
Jim’s idea you mention is still correct & compatible with Turner’s idea. Jim’s point is that we only need to render finite pixels. You might need an x-factor more triangles than pixels because of sampling and depth complexity and secondary lighting, so 100M polys is probably in the ball park, as long as we can quickly pick the right 100M polys in real time…
In a way, Turner was talking about a lower bound, while Jim is talking about an upper bound, albeit slightly different things but they are similar, both relate to how much compute is needed for real time rendering.
> This means that the Cray could run it
Gladly, we had better things to do than that :D
But seriously, while we could have run it maybe speedwise, it definitely lacked the memory, not? And if one tried to train it he wouldn't be finished today. But would make a fun backwards sci-fi story imagining a time traveller that brought the 80ies an LLM from today, what would the world say and do with that slow oracle?
What value does an LLM hold intrinsically.
Lets say "brought an LLM from today"
Does that mean just a multi gig file? What is INSIDE the LLM that would be of value? How does one speak to an LLM WRT 80's tech, and what could one glean from it....
ELI5 an LLM;
BARD: https://i.imgur.com/ahRVECz.png
OpenAI: https://i.imgur.com/Rbk5BD6.png
Bing: https://i.imgur.com/zVJ1tu6.png
--
So, how would one explain 80s folks what even an LLM is when we cant even ELI5 2024?
An LLM is a very highly compressed store of knowledge combined with an advanced parser than understands questions in plain English. A consequence of the compression is that sometimes the answers lose some accuracy, which is a deliberate trade-off to make it work at all.
LLM could explain itself what it is.. (if there are not more important questions to ask, contention would ensue).
I think the process of collecting and storing all the data would be more mind blowing to them—of course they were at the beginning of Moore’s law, so they could see the trajectory if they looked for it, but it is one thing to stand on the coast with waves lapping at your ankles and imagine how the ocean gets deeper as you keep going and another to get chucked out of a helicopter in the middle of the Pacific.
The hard part of LLMs (and current AI in general) is training, which is orders of magnitude harder than inference.
If somehow we had a way to travel to the future in the 70s, train the models and then come back, we would be in Star Trek right now
But... lets look at the availability of DATA in the 80s..
Frankly, this is how hacking/phreaking was invented.
Dumpster-diving for line-printer discards in dumpsters to understand what their systems did.
(This is an actual story; people were bin dipping (at&t?) dumpsters and finding exploits (social or electronic) in the discarded line-printer outputs....
Can someone validate that comment?
--
Brian Roemmele says they've been dumpster diving for decades salvaging huge collections of microfilm/microfiche that's been thrown out by libraries, research institutions, etc.
Now that LLMs are here, they're taking that collection and training an LLM against it (instead of the internet): https://twitter.com/BrianRoemmele/status/1746945969533665422
I wonder if the Raspberry Pi is faster on all tasks, or is there some type of computation the old Cray is still competitive?
But you can emulate a Cray on an FPGA: https://www.chrisfenton.com/homebrew-cray-1a/ so I suspect that while it could still do "real work" you can also beat the pants off it if you setup your code as designed to run on modern GPUs.
[QUOTE] Comparison - The three 700 MHz Pi 1 main measurements (Loops, Linpack and Whetstone) were 55, 42 and 94 MFLOPS, with the four gains over Cray 1 being 8.8 times for MHz and 4.6, 1.6, 15.7 times for MFLOPS.
The 2020 1800 MHz Pi 400 provided 819, 1147 and 498 MFLOPS, with MHz speed gains of 23 times and 69, 42 and 83 times for MFLOPS. With more advanced SIMD options, the 64 bit compilation produced Cray 1 MFLOPS gains of 78.8, 49.5 and 95.5 times.[/QUOTE]
"Raspberry Pi ARM CPUs - The comment above was for the 2012 Pi 1. In 2020, the Pi 400 average Livermore Loops, Linpack and Whetstone MFLOPS reached 78.8, 49.5 and 95.5 times faster than the Cray 1." http://www.roylongbottom.org.uk/Cray%201%20Supercomputer%20P...
A Pi 4 can infer ~0.8 tokens/sec with some of the more optimized configs (as per https://www.dfrobot.com/blog-13498.html). So the Cray would have needed ~2 minutes per token, so ~2.5 hours to generate one sentence... if hypothetically it had enough RAM (it didn't).
In 1978 RAM cost about $25k per megabyte (https://jcmit.net/memoryprice.htm). Assuming you needed 4GB for inference, RAM would have cost $100M in 1978 dollars, or $470M in today's dollars.
For comparison, the Cray cost $7M in 1978 which is $32M in today's dollars. So once you buy a Cray you would have had to spend 14 times that amount on building a custom RAM device extension of 4GB, somehow hooked to the Cray, to finally be able to generate one sentence every 2.5 hours...
But in 1978, even if RAM was available to do LLM inference, it would have been impossible to train the model, as vastly more compute power is needed than for inference.
Besides, it's way further behind in basically every respect but compute.
The applications have grown as well.
The Cray 1 was used for mundane tasks like "large-scale scientific applications, such as simulating complex physical phenomena, and was sold to government and university laboratories." [1] But the power of the Raspberry Pi allows for cutting edge computing tasks like "watering plants, monitoring the birds in your yard, or for a smart doorbell!" [2]
[1] https://www.britannica.com/topic/Cray-1
[2] https://picockpit.com/raspberry-pi/the-7-most-common-uses-fo...
I've always taken them with a grain of salt, but even if they were only an order of magnitude off, a Pi is loads faster than a sidekick. And sure the Cray is loads faster than the Apollo computers, but I wouldn't have thought it was THAT much faster.
I am amazed.
Also, fun fact, it didn't have a CPU. It used all discrete logic chips, and was wired by hand, with lots and lots of wire. IIRC Seymour Cray liked to hire women to do the wiring job, because they had an easier time fitting inside the computer core to wire it, and doing detailed work because they had smaller hands.
That's how CPUs were built back then. What it didn't have was a single-chip CPU, or a microprocessor.
Next up: "The Model T Ford didn't have an engine, it had this gasoline-burning device to provide motive power."
Take a look at all the PCB traces near RAM and CPU. You can see them 'squiggle' looking for matching-space at either end of the RAM/CPU connection.
PCBs make things like consistent length matching easier. All circuit boards have near identical lengths, controlled impedances and other features that support ... Well... Effectively 6-billion baudrates for DDR5.
6000Mhz DDR5 with 64-bits per half-clock (3000Mhz clock so 6000 transfers per second) has a bit every 5 centimeters or so.
Err...
That term is used to describe modern vector instructions that are NOT vector instructions in the classical Cray-sense. Modern CPUs use wide registers that can be regarded as vectors of 2/4/8/... values. Cray used memory-to-memory variable-length vectors.
I'm not saying that "SIMD vector instructions" is an incorrect description of what the Cray machines did, I'm saying that the term usually means something else today.
edit: some searching later:
You get something completely different if failure is not an option.
But also yes, it is a big leap. NASA bootstrapped the semiconductor industry by buying up most of the world's supply. Without the Apollo program we may only just now have gotten smartphones. (And some people still think it's a waste of money. pff!)
The cray was 10,000 lb and 115 kilowatts. Not payload friendly.
I'm imagining Kevin Bacon doing his Apollo 13 power budgeting scene, but needs another ~114kw.
Some of those people are right here. The number of times I've seen the 'whitey on the moon' nonsense is quite large.
https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
And probably most of those people don't realize they wouldn't be posting anything at all if not for 'whitey on the moon', the Apollo program era started with sliderules in 1962 and ended in 1972 with functional microcomputers (8008, 1972, shortly followed by the 6502, 1975) and a very short while later we had programmable pocket calculators (Ti59, 1977). The whole semi conductor industry was jumpstarted in those years.
When I first started out with electronics transistors were an absolute rarity and tubes were normal, in the space of a decade that changed completely.
I strongly disagree with that. The semiconductor industry was thriving before NASA even existed, none of the central enabling inventions (MOSFET transistors, semiconductor manufacturing techniques) were made there or even related to the Apollo program, and "computers" as in turing complete machines already existed before NASA and had plenty of applications apart from space travel.
NASA/Apollo giving todays smartphones a 10 year technology boost is just pure fiction and not even remotely supported by facts.
> By the mid-1960s, according to the PBS documentary, NASA was buying 60 percent of the integrated circuits produced in the United States. Fairchild was a major supplier, shipping about 100,000 devices for the Apollo space program in 1964 alone.
Radiation hardened electronics are fascinating to me. To this day, electronics that function in outer space are much slower than datacenter or consumer-oriented components.
The rPi 4 is over 10x faster than the 1.
The original article points out that the raspberry pi 400, with the benchmark targeting 64 bits, was about 80x Cray for the same benchmark which was 4.5x for the rPi 1.
Raspberry Pi 5 advertises a 2-3x performance gain over earlier versions, though this wasn’t benchmarked in these tests.
As to ARM in general, here's a post about how ARM-1 compared to the 387: https://retrocomputing.stackexchange.com/questions/24826/did...
80486
I think the meaning of my comment was clear: the 80486 was the first Intel CPU guaranteed to have an FPU. I didn't say all the Intel CPUs which came afterward were guaranteed to have them...
Looking over some old info about the "i486" (I swear I had completely forgotten about the i) was a treat. It's only barely believable but I think the sheer annoyingness of Intel's marketing and market segmentation might have actually peaked three decades ago.
It looks like the Cray gets a bit of a boost from the linpack scores (the pi is only 1.6x faster!), which is a good test for the Cray (understatement!).
Well, alright, you can sit on one once.
8MB was the main memory, but it was made from really high speed (and super expensive!) static RAM chips instead of the dynamic RAMs most other machines used. Other machines used SRAMs for caches, so I guess you could consider the Cray 1's main memory to be a cache.
The Cray had 8 "A" address registers but also 64 "B" registers. In the same way you had 8 "S" scalar registers plus 64 "T" registers. The main memory was highly interleaved so you could very quickly load and save blocks of B and T registers as well as the vector registers. You can think of B as a sort of cache for A and T as a sort of cache for S, but you had to explicitly handle this in your program.
Thanks for all your informations.
Seriously though, I think the fact that the PI is general purpose makes it even more impressive.