Cray-1 vs Raspberry Pi
roylongbottom.org.uk
roylongbottom.org.uk
The fidelity would almost certainly be super low compared to modern FEA software, but it would be a fun exercise to try.
in europe
( There was a seperate dedicated weather computer, this one was used for 'other' jobs like speculative weather modelling, monster group algebraic fun, et al.)
https://en.wikipedia.org/wiki/CDC_Cyber
The UK was the first customer:
In 1980, the successor to the Cyber 203, the Cyber 205 was announced. The UK Meteorological Office at Bracknell, England was the first customer and they received their Cyber 205 in 1981.If you really want a challenge, do it using pen, paper and a slide rule, like in the old days[1]. Just make sure to apply appropriate smoothing of the input data first[2].
[1]: https://www.smithsonianmag.com/history/how-world-war-i-chang...
the first machine went to Los Alamos
https://www.theatlantic.com/technology/archive/2014/01/these...
https://ahf.nuclearmuseum.org/ahf/history/human-computers-lo...
"The staff in the T-5 group included recruited women who had degrees in mathematics or physics, as well as, wives of scientists and other workers at Los Alamos. According to Their Day in the Sun: Women of the Manhattan Project, some of the human computers were Mary Frankel, Josephine Elliot, Beatrice “Bea” Langer, Augusta “Mici” Teller, Jean Bacher, and Kay Manley. While some of the computers worked full time, others, especially those who had young children, only worked part time.
General Leslie R. Groves, the Director of the Manhattan Project, pressured the wives of Los Alamos to work because he felt that it was a waste of resources to accommodate civilians. As told by Kay Manley, the wife of Los Alamos physicist John Manley, the recruitment of wives can also be traced to a desire to limit the housing of “any more people than was absolutely necessary.” This reason makes sense given the secretive nature of Los Alamos and the Manhattan Project. SEDs, a group of drafted men who were to serve domestically using their scientific and engineering backgrounds, also worked in the T division."
These are incredibly expensive even on today’s hardware. If you look through some of the unclassified ASCI reports from the early 2000s, 3D calculations of this equation set were implied to be leadership-class computations. At the time of the Cray, it must’ve been coarse-grid 1D as the standard, with 2D as the dream.
Oh cool they got CrayOS working.
But still 1MB RAM, I remember getting the slow RAM 512KB update for my Amiga 500 in the early 90s.
I wish I had access to the code I wrote back then - what took minutes or hours on the Cray could probably run in seconds on a RPi now…
Just before a big tour group of VIPs he knew would come by, he hid inside the Cray-1, and waited patiently for them to arrive.
Then he casually strolled out from inside the Cray-1, pulling up the zipper of his jeans, with a sheepish relieved expression on his face, looked up startled at the tour group gaping at him through the window, and scurried off!
I understand that the exercise may still be theoretical as any sw used by Met now will be designed for million time fast computers. But there should exist some software that would have required Cray to run then.
The Cray-1 had memory management hardware similar to that on DEC PDP-10s that had KA-10 processors and numerous 68K Unix workstations in the '80s, all of which served as interactive time sharing systems just fine.
The memory management hardware in those systems would take the addresses generated by the instructions, compare them to a limit register, and if below the limit would then add them to a base register to get the final memory address. Some systems would have more than one pair of base/limit registers, such as a pair for code and pair for data.
Some people do call such systems virtual memory, because the addresses specified in the code are not the addresses that the memory sees (unless the base register contains 0). But I think it was more common for systems where all the MMU did was simply relocate segments and all of the program's code and data had to be in memory for the program to run to not be considered to be virtual memory. Virtual memory usually meant systems that allowed you to run programs that used address spaces larger than the available memory.
If you had two processes that each needed 100% of the RAM that was available for user program the operating system would have to keep one in RAM and one swapped entirely out to disk or drum. On a task switch it would have entirely swap out the current process, and entirely swap in the other process.
Even if each process was actually spending most of its time in just 10% of its code and most of the time was just actively using 10% of its data, it had to have 100% of its code and 100% of its data present.
The systems people usually gave the name virtual memory to had more sophisticated memory management hardware that would let them map a continuous process address space to a discontinuous physical address space, and you could have gaps in the mapping. Attempts to access gaps would interrupt the process in a resumable way, so the OS could handle that interrupt, map memory into that gap, and resume the process. With these systems they could run both of those 100% memory using programs at the same time while only having to actually allocate enough RAM to cover the 10% of the memory that the two programs were actively using. Context switches between the two did not have to touch drum or disk and so were much faster.
When a program switched to using a different 10% of its code or data, that OS could then map that to real RAM. Since both programs do over time do use 100% of their logical RAM eventually the OS will have to start moving data back and forth between RAM and the drum or disk. But it is not having to do that on every context switch like those systems that needed programs to be 100% in RAM did. Also since it only has to load regions of the program that are actively being used instead of the whole thing, when it does have to use the disk or drum it is for a smaller amount of data and so is faster.
Some TOPS-10 memory trivia. The "core" command allocates memory. On the KA-10 memory was allocated in multiples of 1k words, so "core 1k" was the minimal amount you could allocate. On the later KL-10 model (and I think the KI-10 model but am not sure) the hardware supported allocations in pages, which were half that size. On those you could use "core 1p" to allocate one page.
I saw that in the manual, and was curious what the error message would be if your tried to allocate 1 page on a KA-10. So I typed that command to see what it would say. The command never returned.
I then noticed there was silence in the room and some cursing, and it was apparent the system had crashed. I didn't think much of it because the system crashed fairly often. This was at Caltech and that PDP-10 was the computer that most undergraduate accounts were on, and people were always trying to push limits and explore interesting edge cases. So I figured this was another one of those.
I waited for the admins and operators to bring it back up, and then typed "core 1p" again hoping to this time get to see the error message. It crashed again!
When it came back up, and before I could try a third time, a broadcast message from the admins was sent to all terminals. It said something like "If you want to know why the system keeps crashing, ask tzs who is currently sitting at terminal 7".
Oops. Oh well, at least now I knew that there was not an error message for "tried to allocate 1 page on a system that does not support pages".
Think about mechanical engineering: Back then they might have simulated how cars deform in a crash. Now we can perform similar simulations in real-time for fun in our video games. Afaik it's hardly ever done because no one actually needs physically accurate models in games, but it could be done.
Same goes for rendering back then they rendered each frame of toy story for a good few hours, now we achieve arguably better graphics in real time.
BeamNG is basically that.
It all feels very mundane when I can do it on my slow commodity laptop in under a second.
Because a vector machine is what it was.
With x86 SIMD the standard solution is to compile the same code multiple times for different SIMD widths (using different instructions) and then detect the CPU at runtime. Though that is such a pain that it's only really done in explicitly numerical libraries (Numpy, Eigen, etc.). In theory with Vector you can compile once, run anywhere.
I guess part of my thinking here is that the ISA designers at intel, arm, aren’t stupid, but they ended up with fixed widths for sse, neon, knights landing, avx, avx-512. Presumably they had reasons to prefer that to the dynamic risc-v style thing. So I wonder: are there some risc-v constraints the push this design (eg maybe low-power environments presumably pushed neon to have a small width and this made higher-power environments suffer; having a dynamic length might allow both to use the same machine code), or were there some reasons intel preferred to stick with fixed widths, eg making something that could only work on more expensive chips and thereby having something people can pay more for? Is there something reasonable written about why risc-v went with this design.
[1] what do you even pass to vsetvl in this case as you don’t know your string length.
>what do you even pass to vsetvl in this case as you don’t know your string length.
I'm not sure what you are trying to say here. You must know the length of the buffer, if you don't know the length of the buffer, then processing the string is inherently sequential, just like reading from a linked list, since accessing even a single byte beyond the null terminator risks a buffer overflow. Why pick an example that can't be vectorized by definition?
It’s pretty simple to imagine how to write an unoptimized version: read a vector from the start of the string, compare it to 0, convert that to a bitvector, test for equal to zero, then loop or clz and finish.
I would call this vectorized because it operates on 16 bytes (sse) at a time.
There are a few issues:
1. You’re still spending a lot of time in the scalar code checking loop conditions.
2. You’re doing unaligned reads which are slower on old processors
3. You may read across a cache line forcing you to pull a second line into cache even if the string ends before then.
4. You may read across a page boundary which could cause a segfault if the next page is not accessible
So the fixes are to do 64-byte (ie cache line) aligned accesses which also means page-aligned (so you won’t read from a page until you know the string doesn’t end in the previous page). That deals with alignment problems. You read four vector registers at a time but this doesn’t really cost much more if the string is shorter as it all comes from one cache line. Another trick in the linked code is that it first finds the cache line by reading the first 16 bytes then merging in the next 3 groups with unsigned-min, so it only requires one test against a zero vector instead of 4. Then it finds the zero in the cache line. You need to do a bit of work in the first iteration to become aligned. With AVX, you can use mask registers on reads to handle that first step instead.
To the maximum, of course, `vsetvl x<not 0>, x0, ...` will do that for you.
You might read over a page boundery, but there is an instruction for that `vle8ff.v`, it's a fault-only-first uni-stride load. That is, it doesn't fault when one of the later elements goes outside our page, and adjusts the vector length accordingly. [0]
> With the various simd extensions on x86 or arm, such a routine could be carefully written to align with cache lines and so avoid depending on reading the next line when the string ends before it.
In practice, on current very early hardware, doing just that is faster, and definitely possible. [1]
> I also worry about various simd tricks that seem to rely on the width. Eg I think there’s some instruction to interpret each half-byte of one vector as an index into some vector of 16 things
I'm not aware of such an instruction, but rvv has vrgather.vv and vrgatherei16.vv to do something similar. Note that you can always return to "fixed-size" implementations if you really need to, buy just setting the vl accordingly. But for the most part I don't think this will be necessary. Do you have any specific problem in mind that may seem hard to do without fixed size SIMD?
> I guess part of my thinking here is that the ISA designers at intel, arm, aren’t stupid, but they ended up with fixed widths for sse, neon, knights landing, avx, avx-512.
I mean, ARM now has something similar with SVE, and x86 has a metric ton of legacy to work with and the market size to get adoption of whatever new instruction prefix they add. Edit: Also, didn't AVX10/AVX512vl go into a similar direction as SVE?
> Is there something reasonable written about why risc-v went with this design.
I'm not quite sure, but I remember that it was one thing from very early in the design. I think it's to have a mostly unified ecosystem for the binary app market. It also makes working with mixed precision easier, because it allows for the LMUL model.
But RISC-V is extendable, the P (Packed SIMD) extension is currently in the works and aimed at the embedded market, for in GPR SIMD operations for DSP type applications. [2]
[0] https://github.com/riscv/riscv-v-spec/blob/master/example/st...
[1] https://camel-cdr.github.io/rvv-bench-results/canmv_k230/str...
The instruction I was thinking of is pshufb. An example ‘weird’ use can be found for detecting white space in simdjson: https://github.com/simdjson/simdjson/blob/24b44309fb52c3e2c5...
This works as follows:
1. Observe that each ascii whitespace character ends with a different nibble.
2. Make some vector of 16 bytes which has the white space character whose final nibble is the index of the byte, or some other character with a different final nibble from the byte (eg first element is space =0x20, next could be eg 0xff but not 0xf1 as that ends in the same nibble as index)
3. For each block where you want to find white space, compute pcmpeqb(pshufb(whitespace, input), input). The rules of pshufb mean (a) non-ascii (ie bit 7 set) characters go to 0 so will compare false, (b) other characters are replaced with an element of whitespace according to their last nibble so will compare equal only if they are that whitespace character.
I’m not sure how easy it would be to do such tricks with vgather.vv. In particular, the length of the input doesn’t matter (could be longer) but the length of white space must be 16 bytes. I’m not sure how the whole vlen stuff interacts with tricks like this where you (a) require certain fixed lengths and (b) may have different lengths for tables and input vectors. (and indeed there might just be better ways, eg you could imagine an operation with a 256-bit register where you permute some vector of bytes by sign-extending the nth bit of the 256-bit register into the result where the input byte is n).
For utf8 validation there is a clever algorithm that uses three 4-bit look-ups to detect utf8 errors: https://github.com/simdutf/simdutf/blob/master/src/icelake/i...
Aside on LMUL, if you haven't encountered it yet: rvv allows you to group vector registers when configuring the vector configuration with vsetvl such that vector instruction operate on multiple vector registers at once. That is, with LMUL=1 you have v0,v1...v31. With LMUL=2 you effectively have v0,v2,...v30, where each vector register is twice as large. with LMUL=4 v0,v4,...v28, with LMUL=8 v0,v8,...v24.
In my code, I happen to read the data with LMUL=2. The trivial implementation would just call vrgather.vv with LMUL=2, but since we only need a lookup table with 128 bits, LMUL=1 would be enough to store the lookup table (V requires a minimum VLEN of 128 bits).
So instead I do six LMUL=1 vrgather.vv's instead of three LMUL=2 vrgather.vv's because there is no lane crossing required and this will run faster in hardware: (see [0] for a relevant mico benchmark)
# codegen for equivalent of that function
vsetvli a1, zero, e16, m2, ta, ma
vsrl.vi v16, v10, 4
vsrl.vi v12, v12, 4
vsetvli zero, a0, e8, m2, ta, ma
vand.vi v16, v16, 15
vand.vi v10, v10, 15
vand.vi v12, v12, 15
vsetvli a1, zero, e8, m1, ta, ma
vrgather.vv v18, v8, v16
vrgather.vv v19, v8, v17
vrgather.vv v16, v9, v10
vrgather.vv v17, v9, v11
vrgather.vv v8, v14, v12
vrgather.vv v9, v14, v13
vsetvli zero, a0, e8, m2, ta, ma
vand.vv v10, v18, v16
vand.vv v8, v10, v8
This works for every VLEN greater than 128 bits, but an implementation with larger VLENs do have to do a theoretically more complex operation.I don't think this will be much of a problem in practice though, as I predict most implementations with a smaller VLEN (128,256,512 bits) will have a fast LMUL=1 vrgather.vv. Implementations with very long VLENs (e.g. 4096 bits, like ara) could have a special fast path optimizations for smaller lookup ranges, although it remains to be seen what the hardware ecosystem will converge to.
I'm still contemplating whether or not to add a non vrgather version and runtime dispatch based on large VLENs or quick performance measurements. In my case this would require almost >30 instructions when done trivially. Your example would require about about 8 eq + 8 and vs 4 shuffle + 4 eq, that isn't that bad.
vrgather.vv is probably the most decisive instruction when it comes to scaling to larger vector lengths.
[0] https://camel-cdr.github.io/rvv-bench-results/canmv_k230/byt...
PS: I just looked over my optimized strlen implementation and realized it had a bug. That's fixed now, and the hot path didn't change, just the setup didn't work correctly.
Some of them (Kendryte K230, a MCU) have already shipped to people.
Years ago, some chips shipped, with 0.7.1 (incompatible, pre-ratification). One of them is the TH1520, SoC in some SBCs released earlier this year.
Few have adequate cooling. (flirc is good, the ones with fans are just annoying)
I'd love to have a pi case that had a built-in breadboard.
...or a case with comfortable seating.
"I'm calculating to see when my PC would have been the fastest on Earth. It looks like in 1992 it would be able to out-compute the latest Dept of Defense $90m supercomputer that filled an entire room, would you believe?"
"That's lovely. How will that help us pay our credit bills?"
Jesting aside. There is a bunch of data for this, like this set here:
https://en.wikipedia.org/wiki/TOP500
And if you extrapolate backwards or find older data, like I did, I came to the conclusion that if I took my PC back to 1981 it would actually be faster than every computer on Earth combined, or some insane statistic like that.
https://www.top500.org/system/173736/
When it was commissioned in 2004, this array of 1100x Apple PowerPC 970 systems was the 7th most powerful computer on the list.
It's Linpack Performance was 12,250.00 GFlop/s.
https://www.google.com/amp/s/phys.org/news/2010-12-air-plays...
(Great line in the movie Shooter"
Last capture - https://web.archive.org/web/20070606024231/http://www.tcf.vt...
An FAQ from somewhere in the middle - https://web.archive.org/web/20060708113430/http://www.tcf.vt...
An original RPi is who it 6-7x the performance of a SPARCstation 20 according to the benchmarks
In the meantime plenty of colleges and companies were running entire departments on a PDP-11 that had a fraction of the power.
Raspberry Pi faster than a Cray-1 is cool benchmark of how far we have come! The Cray had built-in seating though, which the Pi doesn't! :-)
The best outcome for plastics, would be to bury them very deep (like nuclear waste), where they could eventually become some new oil like substance. No carbon escape.
The computer had 2048 words of erasable magnetic-core memory and 36,864 words of read-only core rope memory. Both had cycle times of 11.72 microseconds. The memory word length was 16 bits: 15 bits of data and one odd-parity bit. The CPU-internal 16-bit word format was 14 bits of data, one overflow bit, and one sign bit (ones' complement representation). [1]
Or when economics of fabricating such structures just aren't worth it.
Moores ‘law’ is a human driven law.
Computing is basically the absolute center of our society.
As long as our civilization exists we will spend massive resources on this.
Thus as long as it’s physically possible we’ll have progress.
Three semi manufacturers are telling their investors they'll be at 2nm (or something) in late 2024 or 2025. So no, Moore's law has not seen its end, despite ~40 years of predictions to the contrary.
It doesn't have an fpu, and not much ram so it might actually be a close race for some of these tests.
Rpi4s are nice, in a sense, because you can only rarely honestly claim that the speed of the system is holding you back. Most times, presumably, it's the efficiency of the operations you are telling it to execute.
The simpler point I was aiming at was just that: the amount of computational power at the fingertips of so many of us is huge, and it's important to appreciate that.
As someone who uses them for a variety of purposes, I gotta note that they have pretty huge limitations. Like, the moment graphics enter the picture (no pun intended) you’re moving an order of magnitude slower than most desktops or laptops. Not to mention that support for hardware video encode/decode (which, especially decode, we generally take for granted) aren’t always available depending on the library or tool you’re working with.
Like yes, you can totally run a serviceable web server on a Pi and serve a blog or a small web app, but let’s not get carried away here.
Although they did remove the h264 decoder and encoder which is a bummer, like you say it's hard to get working support for it anyway. Vulkan + regular GPU acceleration might be easier. And it still only has 4 cores which is crap for desktop multitasking.
> "Jesus," David said. "That's a Cray 2!"
> "Ten of them." McKittrick said.
> "I didn't know they were out yet."
> McKittrick almost preened. "Only ten. Come on, I want to show you something."
And then you don’t get anything done because there is always a better computer just around the corner. Most of the time, proposals are written for hardware that already exist and don’t need the absolute best. If you have some CFD or MHD calculations to do for a rocket engine or a nuclear reactor, you don’t care about the computer on which it ran, just that it ran on time and did not hold the whole project back. Even cutting edge science does not require cutting edge hardware most of the time.
Just like buying a desktop next year won’t help you play games today, at some point you have to settle and accept that your hardware will be outdated by the time it comes online (it’s a bit better now, but leading HPC clusters still get obsolesced quite quickly).
> maybe this was before moores law though
The exponential character of available CPU time on larger computers was apparent before Moore’s law.
If you compared integer operations it would be a lot closer, but that's not really what the cray was designed for. (The rp2040 at 125mhz * 2 cores is in a pretty similar range)
https://www.top500.org/news/frontier-remains-no-1-in-the-top...
So, how does the Raspberry Pi stack-up against today's computers?
edit: thank you for the Christmas present, yc algorithm. God bless us every one.
Apart from the Cray-1, that whole section is also worth reading for some interesting insights into relative speed differences between various modern CPUs as well. (Though I do wish it was presented in table rather than narrative form, it’d be a lot easier to follow that way; there are also more detailed tables further down the page.)
Trivia: Seymour was user U0100 on our in-house systems.
> Cray has been credited with creating the supercomputer industry. Joel S. Birnbaum, then chief technology officer of Hewlett-Packard, said of him: "It seems impossible to exaggerate the effect he had on the industry; many of the things that high performance computers now do routinely were at the farthest edge of credibility when Seymour envisioned them.
> One story has it that when Cray was asked by management to provide detailed one-year and five-year plans for his next machine, he simply wrote, "Five-year goal: Build the biggest computer in the world. One year goal: One-fifth of the above." And another time, when expected to write a multi-page detailed status report for the company executives, Cray's two sentence report read: "Activity is progressing satisfactorily as outlined under the June plan. There have been no significant changes or deviations from the June plan."
> Cray avoided publicity, and there are a number of unusual tales about his life away from work, termed "Rollwagenisms", from then-CEO of Cray Research, John A. Rollwagen. He enjoyed skiing, windsurfing, tennis, and other sports. Another favorite pastime was digging a tunnel under his home; he attributed the secret of his success to "visits by elves" while he worked in the tunnel: "While I'm digging in the tunnel, the elves will often come to me with solutions to my problem."
Well Seymour you are an odd fellow, but I must say you design a good mainframe.
It supports valves, pumps, schedules, etc. I programmed mine once a few years ago with YAML (no code!) and now I just power cycle them once every month or so. Been running great.
Indeed. I work in both MCUs and full-featured Linux environments and there is zero value to using the former for non-safety critical, non-power constrained, low precision applications. Running a sprinkler system on an RPi is an entirely reasonable choice: you have ample storage for history, trivially simple remote control using a variety of protocols and media and ample compute to operate high level languages and run easily maintained programs, including nice-to-haves like continuous integration of public weather data to optimize your schedule against prevailing rainfall.
Can you shoehorn all that into a ESP32 + micropython or whatever? Sure. I'll bet someone already has. And they spent 10x the time it would have taken otherwise. At least.
Plenty fast for a couple of VMs with web servers. It's also a lot better for 24/7 than RPi, as I was always struggling with SD card wear