Homebrew Cray-1A
chrisfenton.com
chrisfenton.com
It was the first computer I had ever seen, and it didn't disappoint. In 1976, the idea of getting to see a computer at all was the ultimate marvel to a little kid whose granddad had already started reading Asimov and Heinlein to him at night. The Cray fit every imagination I had about what a computer should be—sleek, huge, and like nothing I'd ever seen.
I don't remember anything else about the visit. I don't remember what else was in the room, or what I thought all the people might be doing. All I remember is that it was the first spark of fascination that led me to seek out any chance I could get to be around any computer. Seeing that same machine reproduced in miniature, right down to the bench, warms my heart. I wish my grandfather were around to make it with me. I can't wait to give this a try.
What important computations were made on those machines, i.e., what did they make possible? (Honest question)
That's actually a common optimization, e.g, Itanium doesn't have divide; just reciprocal.
I haven't spent time doing assembly optimization for a very long time, but it used to be the case that, even in x86, you were often better off using various tricks to avoid having to divide. I'd be interested in hearing if that's still the case, from someone who does that sort of thing today.
The actual design was implemented in a Xilinx Spartan-3E 1600 development board.
I love that hardware has gotten so cheap. When I was a teenager, I implemented a Sega on an FPGA, and it took Virtex II board (very high end, at the time) to handle everything. Now, an entire supercomputer fits on one of Xilinx's Spartan (i.e., budget) boards.
It's a bit more interesting than that: The Itanium doesn't have a reciprocal either, just an approximation of the reciprocal accurate to about 9 bits, which can then be refined if more accuracy is needed.
The Cray has no variable time operations at all. All operations have fixed timing. This makes the implementation of vector pipelining much easier.
The Itanic on the other hand relies on the compiler to instructions for optimal instruction pipelining. To do this the compiler needs to know the latency of the operations.
For a x86 a reciprocal instruction makes less sense (except for vectors of course) since most x86 schedule instructions dynamically by hardware. It does make sense multiply with a constant instead of dividing, etc. But a reciprocal doesnt.