> This means that the Cray could run it
Gladly, we had better things to do than that :D
But seriously, while we could have run it maybe speedwise, it definitely lacked the memory, not? And if one tried to train it he wouldn't be finished today. But would make a fun backwards sci-fi story imagining a time traveller that brought the 80ies an LLM from today, what would the world say and do with that slow oracle?
What value does an LLM hold intrinsically.
Lets say "brought an LLM from today"
Does that mean just a multi gig file? What is INSIDE the LLM that would be of value? How does one speak to an LLM WRT 80's tech, and what could one glean from it....
ELI5 an LLM;
BARD: https://i.imgur.com/ahRVECz.png
OpenAI: https://i.imgur.com/Rbk5BD6.png
Bing: https://i.imgur.com/zVJ1tu6.png
--
So, how would one explain 80s folks what even an LLM is when we cant even ELI5 2024?
An LLM is a very highly compressed store of knowledge combined with an advanced parser than understands questions in plain English. A consequence of the compression is that sometimes the answers lose some accuracy, which is a deliberate trade-off to make it work at all.
LLM could explain itself what it is.. (if there are not more important questions to ask, contention would ensue).
I think the process of collecting and storing all the data would be more mind blowing to them—of course they were at the beginning of Moore’s law, so they could see the trajectory if they looked for it, but it is one thing to stand on the coast with waves lapping at your ankles and imagine how the ocean gets deeper as you keep going and another to get chucked out of a helicopter in the middle of the Pacific.
The hard part of LLMs (and current AI in general) is training, which is orders of magnitude harder than inference.
If somehow we had a way to travel to the future in the 70s, train the models and then come back, we would be in Star Trek right now
"Raspberry Pi ARM CPUs - The comment above was for the 2012 Pi 1. In 2020, the Pi 400 average Livermore Loops, Linpack and Whetstone MFLOPS reached 78.8, 49.5 and 95.5 times faster than the Cray 1." http://www.roylongbottom.org.uk/Cray%201%20Supercomputer%20P...
A Pi 4 can infer ~0.8 tokens/sec with some of the more optimized configs (as per https://www.dfrobot.com/blog-13498.html). So the Cray would have needed ~2 minutes per token, so ~2.5 hours to generate one sentence... if hypothetically it had enough RAM (it didn't).
In 1978 RAM cost about $25k per megabyte (https://jcmit.net/memoryprice.htm). Assuming you needed 4GB for inference, RAM would have cost $100M in 1978 dollars, or $470M in today's dollars.
For comparison, the Cray cost $7M in 1978 which is $32M in today's dollars. So once you buy a Cray you would have had to spend 14 times that amount on building a custom RAM device extension of 4GB, somehow hooked to the Cray, to finally be able to generate one sentence every 2.5 hours...
But in 1978, even if RAM was available to do LLM inference, it would have been impossible to train the model, as vastly more compute power is needed than for inference.
[QUOTE] Comparison - The three 700 MHz Pi 1 main measurements (Loops, Linpack and Whetstone) were 55, 42 and 94 MFLOPS, with the four gains over Cray 1 being 8.8 times for MHz and 4.6, 1.6, 15.7 times for MFLOPS.
The 2020 1800 MHz Pi 400 provided 819, 1147 and 498 MFLOPS, with MHz speed gains of 23 times and 69, 42 and 83 times for MFLOPS. With more advanced SIMD options, the 64 bit compilation produced Cray 1 MFLOPS gains of 78.8, 49.5 and 95.5 times.[/QUOTE]
But... lets look at the availability of DATA in the 80s..
Frankly, this is how hacking/phreaking was invented.
Dumpster-diving for line-printer discards in dumpsters to understand what their systems did.
(This is an actual story; people were bin dipping (at&t?) dumpsters and finding exploits (social or electronic) in the discarded line-printer outputs....
Can someone validate that comment?
--
Brian Roemmele says they've been dumpster diving for decades salvaging huge collections of microfilm/microfiche that's been thrown out by libraries, research institutions, etc.
Now that LLMs are here, they're taking that collection and training an LLM against it (instead of the internet): https://twitter.com/BrianRoemmele/status/1746945969533665422
I wonder if the Raspberry Pi is faster on all tasks, or is there some type of computation the old Cray is still competitive?
But you can emulate a Cray on an FPGA: https://www.chrisfenton.com/homebrew-cray-1a/ so I suspect that while it could still do "real work" you can also beat the pants off it if you setup your code as designed to run on modern GPUs.
Besides, it's way further behind in basically every respect but compute.