Nobody ever ported Doom to run on a Cray 1
twitter.com
twitter.com
A guy is making a replica. He had some issues finding the operating system and systems software but he overcame it
Persistence pays off.
I am sure it would be fun to part doom to it. Easier too since the new versions draws Considerable less power.
https://gigaom.com/2014/01/14/the-search-for-the-lost-cray-s...
"Andy Gelme, an Australian software developer who once worked for Cray. He too had a disk pack [containing the OS]."
"For the greater part of the last year, [Tantos, a Microsoft electrical engineer] arduously reverse engineered the OS from the [corrupted / incomplete] image. Despite a few remaining bugs, the Cray OS now works."
Can you elaborate?
It is described in detail in the manual: http://www.bitsavers.org/pdf/cray/Disk/HR-0031_SolidStateSto...
That's even faster than many modern SSDs.
https://web.archive.org/web/20201112040756/http://www.s100co...
An uncle that worked in that area in the 80s/90s had stories about people who still argued these new fangled journaling file systems were pointless.
There used to be big caveats to using XFS on non-SGI hardware. Though, I think those caveats were roughly "XFS guarantees aren't much better than ext3/ext4 guarantees on power loss, on non-SGI hardware".
I'd imagine it would be practical to build and use something similar for a multi-million dollar supercomputer where cost is secondary to performance.
Because there's absolutely no scenario after launch, in which the constant trickle of wattage flow from the RTG to the onboard DC distribution buses and computers would ever be interrupted.
If the power from the RTG were to ever be cut off it would be a catastrophic mission failure anyways, entirely aside from the power to the onboard computing systems being interrupted.
I don't buy it. They specifically design these systems with duplication in nature in case RAM gets corrupted by radiation in space. So you can't treat it as storage, but for very different reasons.
The SSD wasn't new by any means. A decade earlier the Cray-1M optionally could have had an SSD.
The gigabytes of SRAM used as main RAM you could stick on a Cray was a key differentiator for the platform (and would be still today).
These days with M2 SSDs it isn't really necessary and (as was the case back then) you are usually better off using any spare RAM for programs.
The biggest problems will be finding a C compiler for the Cray and adding a framebuffer output, I think... but this is a very nice challenge :).
I suppose you could run Doom as a batch job: input is a file with the control input and output is a gameplay video.
https://doomwiki.org/wiki/Demo
https://doomwiki.org/wiki/Doom_Wiki:Departure_from_Wikia
“LMP-to-video converter” is any Doom port which output you can record by common (or less common) means. PrBoom+ has a built-in demo playback recording (if you consider running the console encoders and feeding them sound and image data “built-in”).
https://github.com/coelckers/prboom-plus/blob/master/prboom2...
Seems possible that this aspect of the platform might drive the need some rework though [1]:
sizeof(unsigned int) = 8; UINT_MAX = 18446744073709551615
sizeof(unsigned long) = 8; ULONG_MAX = 18446744073709551615
sizeof(unsigned char) = 1; UCHAR_MAX = 255
sizeof(unsigned short) = 8; USHRT_MAX = 4294967295
i.e. ints are 64-bit longs, shorts are 32-bit but use 64-bits[1] http://www.modularcircuits.com/blog/articles/the-return-of-t...
My favorite anecdote of 25 years on those machines is what we put into the 64-bit word at address 0 on the Cray-2. In ASCII, it read ~Z~E~R~O. If you jumped to it, it worked as four no-op ("PASS") instructions, and at word 1 was a jump to the library's routine that dumped registers and said "hey you jumped to a null function pointer". If you ever saw ~Z~E~R~O in a dump, you knew that a load from null had taken place.
I still wish a modern ISA would implement real vectors; SIMD is still a distant second-best.
There's this newly refound idea that maybe the Seymour Cray guy knew what was up.
As a 'practical' example, I was able to write an N-body simulator of Jupiter and 63 of its moons (using the vector registers) orbiting one another in only 127 total instructions!
This idea has been resurrected recently with RISC-V and ARM's scalable vector instructions. There the general idea is an instruction that assigns the minimum of an argument value and the hardware vector register length to a register, and sets the masking appropriately if the argument is smaller. This makes for a very straightforward strip mined loop without a branch to check for and handle the remainder in the last iteration.
The Cray-1 line could "chain" the results of one vector operation into operand(s) of another without waiting for the first to complete. (On the Cray-1, the later operation had to issue at the exact "chain slot" cycle at which the first result element appeared, so scheduling was fun; on the X-MP and later, "flexible chaining" was possible). Scheduling vector code involved grouping operations into "chimes" that would run as parallel chained operations, and so long as you could pack more vector instructions into a chime without causing synchronization due to register use or blocking on a functional unit busy, you won. Getting a 3-chime loop down to 2 chimes was fun puzzle solving, and if the loop used (say) the floating adder twice, you knew you could stop optimizing.
The Cray-2 didn't chain, but the Cray-3 had "tailgating", which was kind of the opposite -- a new vector result could start writing to a vector register that was in use as an operand without having to wait for that operand use to complete.
It helps to think about these vector machines as being pipelined (which they were). A single chime sequence was basically flowing data from memory to functional units and back to memory without really needing to use the vector registers per se for anything unless an interrupt arrived in the middle of the sequence.
- vector length register, to avoid loop epilogues.
- scatter-gather and strided memory ops
- support for predication / masking
- looser alignment restrictions, ideally as small as the element size rather than the entire vector width.
Look into Arm SVE and RISC-V V extension for modern incantations of this. Though the x86 world is slowly getting closer too.
Incidentally, a Cray-1 uses 115 kW of power when running. Would that be a new power consumption record for a device that's just running Doom?
They looked more like abstract art pieces or 70ties science fiction props.
A museum piece now. Would be great to see it running again. Not likely.
(It draws a mere 100kW at the wall, I'm guessing no)
So where did the rest of the power go, into cooling? Perhaps it was a linear power supply instead of a switcher? That would make for lots of inefficiency.
Interestingly, the power supplies themselves are unregulated; regulation is done by the motor-generator unit.
http://www.chrisfenton.com/cray-1-digital-archeology/
http://www.chrisfenton.com/cos-recovery/
I just wish he'd made higher resolution pictures.
He's also seriously hardcore (or maybe I feel out of depth because I'm a software guy):
> I built a robot that would manually move the head forward 1/5200th of an inch at a time (there are 400 data tracks per inch, so this gives me a whopping 13 steps per data track!), while a high-speed analog-to-digital converter would take the analog signal straight from the drive’s read amplifier and buffer it into an FPGA at a blistering 80 million samples-per-second (like I said earlier, the theme here was overkill…the data was only changing at ~10 MHz or so)
Related: http://www.modularcircuits.com/blog/articles/the-cray-files/
The FPGA thing isn't that inaccessible to normies like us (i.e. some FPGA Dev boards will easily cost thousands and thousands), but the ecosystem is a complete clusterfuck so getting started is the hard bit. Obviously you need to know electronics too but 80MSPS will let you get away with a lot (The term is signal integrity) of bad circuitry (as I can attest to...).
Verilog isn't too bad as long as you treat it as a description of a circuit (in a language from the past, not even mentioning VHDL...)
EDIT for those who may not have understood my comment. You would be hard-pressed to find a Cray that performed division and would instead need to perform a reciprocal multiplication in constant time.
1) Make a guess at the reciprocal X(0) (this can come from a small LUT, but a static "guess" can also work)
2) Calculate an improved reciprocal X(n) iteratively using the relation:
X(i+1) = X(i)*(2-X(i)*D)
Repeat (2) until X(i) is accurate "enough". This is relatively fast, since the # of correct bits in X() doubles for each iteration.There are several other methods: https://en.wikipedia.org/wiki/Division_algorithm#Fast_divisi...
Has anyone been able to emulate Tandem systems and their OS? That was a interesting high-reliability system, still worth attention. The hardware just cost too much back when it was a product.
HPE still sells the Tandem OS, NonStop, and they have ported it to run on x86. Unfortunately, unlike OpenVMS, they’ve never (to the best of my knowledge) run a hobbyist program for it.
Also I'm pretty sure ints were 64-bits.. it's possible that shorts were also 64-bits, but I don't remember.
the guy at the computer museum was happy
"This is E1M1. I know this!!"