Which Machines Do Computer Architects Admire? (2013)
people.cs.clemson.edu
people.cs.clemson.edu
At the time, it seemed (to me at least) that it really only died because the backwards compatibility mode was slow. (I think some of the current perception of Itanium is revisionist history.) It's tough to say what it could've become if AMD64 hadn't eaten it's lunch by running precompiled software better. It would've been interesting if Intel and compiler writers could've kept focus on it.
Nowdays, it's obvious GPUs are the winners for horsepower, and it's telling that we're willing to use new languages and strategies to get that win. However, GPU programming really feels like you're locked outside of the box - you shuffle the data back and forth to it. I like to imagine a C-like language (analogous to CUDA) that would pump a lot of instructions to the "Explicitly Parallel" architecture.
Now we're all stuck with the AMD64 ISA for our compatibility processor, and it seems like another example where the computing world isn't as good as it should be.
The FPGA projects I've seen (using very high end FPAGs, not commodity/cheap ones) seem like they're always bumping up against clock rates and making timing as soon as they try to do anything approaching what you can do on a CPU or GPU. Of course there are exceptions where the FPGA does really simple and parallel things, but FPAGs aren't a panacea.
Ways it is not c-like:
- begin...end instead of curly braces
- parallelism with assign, always, initial
- tasks and functions and their subtle differences
- non-blocking assignment
- bit-oriented variables and operations
- 4-state logic (0, 1, X, Z)
- other constructs for modeling hardware like time delays, tri-state wires, drive strengths, etc.
Not to mention that SystemVerilog has taken over Verilog and adds OOP with classes, a complicated (sorry, powerful) assertion mini-language, constrained randomization, a streaming operator, and so on.
It’s hard to compete with the scale of x86. Like software I feel the industry tends toward one architecture (the more people use the architecture the better the compilers the more users ...) Even Apple abandoned PowerPC chips.
Linux Torvald has an interesting take on this (whether or not you agree): https://yarchive.net/comp/linux/x86.html
> the people who actually buy Intel kit basically want a faster 8088/80386/Pentium
My company is small fries compared to most folks, but what we really wanted at the time was cheaper DEC Alphas.
I think he'd be the first to agree with that.
PAE, which is pretty invisible to use userland folk
Perhaps, but the original conversation is about architecture, and PAE was a pretty grungy architectural wart, and the sort of thing that's very visible to the OS folks (as you say, what Linus cares about).
we really wanted at the time was cheaper DEC Alphas
There were 'cheap' Alphas: the 21066/21068 like in the Multia. But they were dogs; to be cheap, they had to give up big cache and wide paths to main memory. Expensive system support level stuff was required for fast Alphas (complex support chipsets, 128-bit wide (later 256-bit) memory buses). More commodity inertia would have fixed that over time, but they never got there. Intel on the other hand was way down the road reaping commodity benefits, and it ran the software commodity folks wanted.
I suspect if he was into writing compilers (or graphics, or numerics, or ...), some of the other grungy architectural warts of x86 might annoy him too.
Itanic was the safer bet in 64-bit computing. It just sucked. Intel didn't switch to AMD64 until underdog AMD was already eating their lunch.
Today there are probably more aarch64 CPUs being sold every month than amd64 CPUs (including Intel's).
> Author: AgnerDate: 2015-12-28 01:46
> Ethan wrote:
> > Agner, what's your opinion on the Itanium instruction set in isolation, assuming a compiler is written and backwards compatibility do not matter?
> The advantage of the Itanium instruction set was of course that decoding was easy. The biggest problem with the Itanium instruction set was indeed that it was almost impossible to write a good compiler for it. It is quite inflexible because the compiler always has to schedule instructions 3 at a time, whether this fits the actual amount of parallelism in the code or not. Branching is messy when all instructions are organized into triplets. The instruction size is fixed at 41 bits and 5 bits are wasted on a template. If you need more bits and make an 82 bit instruction then it has to be paired with a 41 bit instruction.
(https://www.agner.org/optimize/blog/read.php?i=425)
Besides, the memory consistency model of Itanium is also a brain teaser used in interviews as counterexamples to poorly-synchonized solutions.
If I'm doing a million point FFT, I can easily give you 2 million operations in a row without a loop/branch. Maybe 1000 of those at a time could be run in parallel before the results needed to commit for the next 1000. I'd be willing to pay for 1 or 2 nops in the last bundle of every 1000 operations. I admit, the idea might not be awesome for a word processor or spreadsheet, but I did specify signal processing and machine learning.
Nowdays, it's pretty clear GPUs are the winner, but like I said, they just don't feel like you actually live and breath inside of them the way you do with CPU code (you shuffle your data over and shuffle it back), and I'm kind of just imagining an alternative timeline where Intel and compiler writers got a chance to run with the EPIC idea.
Or a web server or browser. In fact, it pretty much only helps for your use case. Which is why people are converging on specialised hardware for it, and why Itanium was a commercial failure.
I know it's all a fantasy now, but I wonder what the world might have been like if there wasn't such a split between what you can only do on a CPU and what you can only do on a GPU. Maybe you like heterogeneous computing and gluing C++ to CUDA, but I think it's ugly. Stretch outside of the current box a little bit and imagine a hybrid somewhere in the middle ground of CPUs and GPUs. I think a variable sized VLIW could've gotten there if the market had any more imagination than it does. It's Ford's "faster horses" problem.
Related discussion: https://news.ycombinator.com/item?id=17606037
GPUs showed two things: one, you can relegate kernels to accelerators instead of having to maximize performance in the CPU core; and two, you can convince people to rewrite their code, if the gains are sufficiently compelling.
(I ended up running the OS group at Multiflow before bailing right before they hit the wall.)
Erlang's OTP captures this pretty okay-ish with their concept of processes (preemptively scheduled green threads that get aggressively multiplexed on all CPU cores) but I feel we can go a little bit further than that and have some sort of shorter markers in the code, say `actor { ... }` or `supervises(a1, a2, a3) { ... }` or something.
https://www.anandtech.com/show/10025/examining-soft-machines...
First, compilers just could never seem to find enough (static) ILP in programs to fill up all the instructions in a VLIW bundle. Integer and pointer-chasing programs are just too full of branches and loops can't be unrolled enough before register pressure kills you (which, btw, is why Itanium had a stupidly huge register file).
Second, it exposes microarchitectural details that can (and maybe should) change quite rapidly. The width of the VLIW is baked into the ISA. Processors these days have 8 or even 10 execution ports; no way one could even have space for that many instructions in a bundle.
Third, all those wasted slots in VLIW words and huge 6-bit register indices take up a lot of space in instruction encodings. That means I-cache problems, fetch bandwidth problems, etc. Fetch bandwidth is one of the big bottlenecks these days, which is why processors now have big u-op caches and loop stream detectors.
Fourth, there are just too many dynamic data dependencies through memory and too many cache misses to statically schedule code. Code in VLIW is scheduled for the best case, which means a cache miss completely stalls out the carefully constructed static schedule. So the processor fundamentally needs to go out of order to find some work (from the future) to do right now, otherwise all those execution units are idle. If you are going out of order with a huge number of execution ports, there is almost no point in bothering with static instruction scheduling at all. (E.g. our advice from Intel in deploying instruction scheduling for TurboFan was to not bother for big cores--it only makes sense on Core and Atom that don't have (as) fancy OOO engines).
There is one exception though, and that is floating point code. There, kernels are so much different from integer/pointer programs that one can do lots of tricks from SIMD to vectors to lots of loop transforms. The code is dense with operations and far easier to parallelize. The Itanium was a real superstar for floating point performance. But even there I think a lot of the scheduling was done by hand with hand-written assembly.
I wish I had read to the end before replying (and then deleting) responses to each of your items. :-)
I think we're mostly in agreement, except for the following minor tidbits:
> The width of the VLIW is baked into the ISA [...] no way one could even have space for that many instructions in a bundle.
My understanding was that your software/compiler could indicate as many instructions as possible in parallel, and then stuff a "stop" into the 5 bit "template" to indicate the previous batch needed to commit before proceeding. So, you wouldn't be limited to 3 instructions per bundle, and if done well (hopefull), your software would automatically run faster as the next generation comes out and has more parallel execution units.
> all those wasted slots in VLIW words and huge 6-bit register indices take up a lot of space in instruction encodings
128 registers, so I'd think it'd be 7 bit indices. Each instruction was 41 bits, which is roughly 5 bytes. Most SSE/AVX instructions end up being 4-5 bytes (assuming no immediates or displacements), and that's for just 4 bit register indices. So it doesn't seem much worse than we have now.
No. Itanium was never a superstar - it was merely competitive when it was at best, and even that was only if you ignored the price/performance, and some of its versions were pretty bad in absolute performance and nowhere near competitive, and it was plain abysmal if you consider price/performance. Also, majority of practical, important numerical computations were memory bandwidth bound, and thus it didn't matter as much whether you can pack the loop perfectly. And Itanium was almost never the highest memory bandwidth machine during its lifetime, partially due to many of its delays.
Itanium would be my choice for the worst architecture, as it successfully killed other cpus, and produced a lot of not very useful research.
MIPS (and Hennessy and Patterson) would be my first choice, for upending the architecture design. Honorable mentions from me would be IBM 801 (lead to many research), Intel iAXP 432 (for capability architecture, which I think will come back at some point), z80 and 8501 for ushering the computing power everywhere, x86-64 for "the best enduring hack".
Anyway, the article itself is great, and I wish I had asked the same question to some of those folks mentioned in the article, and many other architects and researchers when I met them...
I agree that it was a revolutionary design for the time. From today's perspective, some of the choices made did not age all that well (in particular branch and load delay slots).
For assembly level programming/debugging, my favorite architecture by far was POWER/PowerPC. Looking at x86 code vs PowerPC after Apple's switch made me almost cry, although the performance benefits were undeniable.
It kinda didn't work that way though.
In practice, all of your ALUs, including your extra ones, were waiting on cache fetches or latencies from previous ALU instructions.
Modern x86 CPUs have 2-4 ALUs which are dispatched to in parallel 4-5 instructions wide, and these dispatches are aware of cache fetches and previous latencies in real time. VLIW can't compete here.
VLIW made sense when main memory was as fast as CPU and all instructions shared the same latency. History hasn't been kind to these assumptions. I doubt we'll see another VLIW arch anytime soon.
I accept the idea that x86 is a local minimum, but it's a deep, wide one. Itanium or other VLIW architectures like it were never deep enough to disrupt it.
I think if anyone could get around memory bandwidth problems they would, but for some very interesting and useful algorithms, I can tell you way in advance exactly when I'll need each piece of memory. For these problems, VLIW/EPIC with prefetch instructions would be a win over all the speculation and cleverness.
> Itanium or other VLIW architectures like it were never deep enough to disrupt it.
History is what it is, but I'm just imagining an alternative timeline where all the effort spent making Pentium/AMD64 fast was pumped into Itanium instead, and compiler writers and language creators got to target an architecture that didn't act like a 64 bit PDP-11.
Anyway, you increase throughput by just adding more hardware. That's easy and widely done.
In Java... not so much. I believe that "." in Java is the same as "->" in C/C++, unless what's on the right is a primitive data type rather than an object.
https://en.wikipedia.org/wiki/Content-addressable_memory (CAM), as used in network switches, isn’t under the same constraints as regular RAM. The requests you make to CAM are CISC—effectively search queries—putting the whole memory-cell array to work at 100% utilization on each bus cycle.
But even CAM is still slower than the CPU. Even when it’s on the same SoC package as the CPU, it’s still clocked in such a way that it takes multiple CPU cycles to answer a query. So, at least in this case, bus bandwidth is not “the” constraint.
There is nothing that prevents you from making SRAM/CAM array running at same or even higher clock speed than CPU made with same semiconductor technology except cost of the thing. And in fact, n-way associative L1 cache (for n>1) is exactly such an CAM array.
Of course, the joke is that cheap x86 processors did outperform Itanium (and every other architectures, eventually).
I'll admit I had a limited worldview, but not running Excel as quickly seemed like the kind of criticisms I saw in the trade rags at the time.
> the idea that a massive array of cheap x86 processors will outperform enterprise-class servers simply hadn't occurred to most people yet
Oh, I don't know. There was a really common Slashdot cliche running around at that time: "Can you imagine a Beowulf cluster of these?"
HOT GRITS!
The Mozilla open source announcement, the Microsoft anti-trust case, Linux exploding in popularity, it all felt like we were changing the world for the better.
HN is also less than it used to be, but I think that the mods have kept it better than /. became - so far, at least.
When do you mean? In 2001?
It occurred to Yahoo, whose site had been run that way since almost the beginning, on FreeBSD. It occurred to Google, whose site was run that way since the beginning. It occurred to anyone who was watching the Top500 list, which was already crawling with Beowulfs — admittedly, not at the top of the list yet. It should have occurred to Intel, who were presumably the ones selling those servers to Yahoo and Penguin Computing and VA Research (who had IPOed in 1999 under the symbol LNUX). It had occurred to Intergraph, who had switched from their own high-performance graphics chips to Intel's by 1998. It had occurred to Jim Clark, who had jumped off the sinking MIPS ship at SGI. In 1994. It occurred to the rest of SGI by 1998, when they launched the SGI Visual Workstation, then announced they were going to give up on MIPS and board the Itanic.
I mean, yes, it hadn't occurred to most people yet. Because most people are stupid, and most of the ones who weren't stupid weren't paying attention. But it hadn't occurred to most people at Intel? You'd think they'd have a pretty good handle on how much ass they were already kicking.
> the joke is that cheap x86 processors did outperform Itanium (and every other architectures, eventually).
Even on my Intel laptop, more of the computrons come from the Intel integrated GPU. In machines with ATI (cough) and NVIDIA cards, it's no contest; the GPU is an order of magnitude beefier.
The idea was certainly around in the early-to-mid 1980s, when some former Intel engineers founded Sequent.
The Balance 8000, released in 1984, supported up to 12 processors on dual-CPU boards, while the Balance 21000, released in 1986, supported up to 30.
I interviewed the founder, Casey Powell, and he was explicit about multiple Intel microprocessors replacing large systems. He was targeting minicomputers at the time, of course, but we all anticipated that bigger sets of more powerful CPUs would eventually surpass even the biggest "big iron".
Powell was a great guy. However, his company got taken over by IBM. In the end, he didn't get to change the world.
"It's hard to be the little guy on the block and have really great technology and get beaten, just because the other guy is big." https://www.cnet.com/news/sequent-was-overmatched-ceo-says/
An interesting question is: what are the structural advantages of bigness? When Control Data produced the world's fastest computer, some people at IBM wondered how it could happen that a much smaller company could beat them to the punch that way; others believed that that smallness was precisely the reason.
The advantages of bigness are that you can use scale to make the same thing less expensive, and that you can make at least one mistake without it killing you, and that you can chase more than one "next big things" at once.
It did. At that time Linux and open source were not the clear winners in the server space they are now and people were not used (or able) to recompile their code.
Windows took a long time to support Itanium and companies wouldn't buy it because they had nothing to run on those machines. They got x86 machines instead and amd64 when it became available.
Regarding GPUs, Fujitsu may not agree (for Fugaku and spin-offs) depending on the value of "horsepower" relevant for HPC systems, even if an A64FX doesn't have the peak performance of a V100. They have form from the K Computer, and if they basically did it themselves again, there was presumably "co-design" for the hardware and software which may be relevant here; I haven't seen anything written about that, though.
Looking at this chip, it seems to me that almost all the innovations Intel brought to the Pentium lines of CPU over many years were basically reimplementing features pioneered by the DEC Alpha, just over a decade later, and bringing these innovations to consumer-grade CPUs.
> it seems to me that almost all the innovations Intel brought to the Pentium lines of CPU over many years were basically reimplementing features pioneered by the DEC Alpha
I can't find a strong source to link, but I thought most of the Alpha team ended up at Intel. If so, that would explain the trickling in of re-implementations.
DEC was bought by Compaq, who sold the team to Intel at the same time it was bought by HP. Intel's Massachusetts site is a former DEC facility.
Compaq abandoned DEC's Alpha for the HP/Intel Itanic, so Intel didn't get an Alpha business. However, Compaq sold the Alpha IP to Intel in 2001, before the HP takeover in 2002.
I'd be interested to know what happened to the Alpha architects, Richard L. Sites and Richard T. Witek. A quick search doesn't find anything interesting.
I remember there was a breakaway of DEC engineers founding a small chip design company, but can't remember what it was called.
Jim Keller and other folks from PA Semi worked at DEC earlier in their careers.
Always loved Alpha's, their influence cast a long shadow.
I have had very unconfirmed rumour that K8 involved so much of DEC stuff (starting with it's obviously EV7-derived design) that at one point it still had support for VAX floating point...
I used the DEC as my personal machine. Hardly anyone ever touched the other two, except to verify bug reports.
There where patent disputes and such.
According to this article which describes the sale, windows NT ran on alpha.
https://www.latimes.com/archives/la-xpm-1997-oct-28-fi-47463...
[1] https://web.archive.org/web/20170408085131/https://katzentie...
Hellish price!
DEC had bragged about the features of the cpu useful for parallelization. Cray engineering complained about the features missing for parallelization. It was all described in a glossy Cray monthly magazine description of the new machine but I've been unable to locate a copy.
I disliked the Alpha floating point, it was always signalling exceptions for underflow. Otherwise a fine set of machines.
Alpha floating point wasn't the problem.
The problem was that a whole bunch of "clever" folks used "underflow" for all manner of weird reasons on x86. So, whenever Alpha either ran ported or emulated x86 code, it ran into underflows with far greater frequency than any actual numerical applications ever would.
I remember DEC Alphas absolutely stomped all over the x86 stuff that everyone else was using, but the flexibility and price of the commodity PCs was just too attractive. Pity, really.
Maybe these are the machines bad computer architects, like Alpert, admire. Alpert is notable mostly for leading the computer industry's most expensive and embarrassing failure, the Itanic (formally known as the Itanium), despite the presence on his team of many of the world's best CPU designers, who had just come from designing the HP-PA --- a niche CPU architecture nevertheless so successful that HP's workstation competitors, such as NeXT, started using it. Earlier in his career he sunk the already-struggling 32000, the machine that by rights should have been the 68000. (And maybe if they'd funded GCC it could have been.)
What about the Tera MTA, with its massive hardware multithreading and its packet-switched RAM, which was gorgeous and prefigured significant features of the GPU explosion?
What about the DG Nova, with its bitslice ALU chips and horizontal-microcode instructions? What about the MuP21, with its radical on-chip dual circular stacks?
What about the HP 9100, with its dual stacks and PCB-inductance microcode, where the instruction set was the user interface?
What about the LGP-30, which managed to deliver a usable von Neumann computer with only 113 vacuum tubes (for amplification, inversion, and sequencing)?
What about the 26-bit ARM, with its conditional execution on every instruction, and packing the program status register into the program counter so it automatically gets restored by subroutine return, and, more importantly, interrupt return?
What about Thumb-2 with its unequaled code density?
What about the CM-1? Anyone can see that AVX-512 (or for that matter modern timing-attack-resistant AES implementations!) owe everything to the CM-1.
And the conspicuous omission of the Burroughs 5000 has already been noted by others.
I mean, there are some good designs on the list! But it hardly seems like a very comprehensive list of admirable designs.
Alpert is a bad architect...funny.
It came out in 2003, and most of the people were queried for their opinions in 2001.
I'd add the Tandem NonStop to my personal list. I don't know why I overlooked the LGP-30 [1], I'll have to find a schematic. 113 vacuum tubes is really impressive, I wonder if there is any overlap with this design and System Hyper Pipelining [2]. Do you know of other architectures that use time multiplexing to reduce part count?
What bit serial computers do you like?
Ahh, it is the Story of Mel computer, awesome.
Delay-line and drum computers (like the LGP-30, the HP 9100, and the grandmama of them all, the Pilot ACE) all sort of had to do a sort of time multiplexing; the Tera MTA I mentioned, as well as the CDC 6600's PPs (FEPs), worked that way too, time-sharing a single ALU and control unit among many register sets. That's also one of the things going in modern GPUs, but it's hard to say it's to reduce part count. Still, they'd need a lot more parts to do the same thing if they didn't do it.
This CSR/SHP thing sounds really interesting! Thank you!
I got to tour the Tera offices in Seattle in the late 90s, about all I remember is that it was a torus and it used some finicky silicon process that was leading to manufacturing delays. I was all into Beowulf and Mosix [2] at the time using Alpha or x86, so I wasn't drawn to it that much.
Tera’s first machine was bipolar (ECL I assume) and they finally squeezed out a CMOS successor with a lot of assistance from their EDA vendor. Never knew the story of why moving to CMOS was so urgent.
Amusingly, it was the Beowulf list where someone converted me to the Tera religion (rgb I think). I was convinced that was the way all computers would work soon. And, well, it's how GPUs work, kind of. But mostly I was wrong.
Not sure I agree about modularity. Galaxies, mammalian bodies, trees, bacterial films, cars, books, and river systems are modular. It would be surprising if we could make software non-modular. But we could make it only as modular as a tree.
Composition is great, scale free self similarity is probably the basis for the universe.
Modularity is a great design technique, it can also make things weaker and force other (unknowable) design choices because the module boundary prevents the flow of information/force. Overly constrained modular systems encourage globals, under constrained modular systems are asymptotic to mud/clay.
I don't want to use K8S as a strawman to attack modularity, but I think it is an example of using this powerful design tool to solve the wrong problem using mis-applied methods all the while being more complex and using more resources. In the case of designing systems, modules/objects/processes (Erlang sense) are critical, but not so much in building/engineering them. Demodularizing or fusing a design can make it more robust and more efficient.
I don't dislike modularity, I just think it is a bigger more complex topic than most give it credit for. Unix is highly-non modular and very poor composition. It sits on a molehill of a local maximum, itself sitting in the bottom of a Caldera, a sort of Wizard Mt on Wizard Island.
Other things you might like is the research around "Collapsing Towers of Interpreters" [1]
Or Dave Ackley's T2 Tile Project and Robust First Computing [2]
Would love to chat more, but internet access is spotty for the next week, non-replies are not ignores.
[1] https://lobste.rs/s/yj31ty/collapsing_towers_interpreters
[2] https://www.youtube.com/watch?v=7hwO8Q_TyCA https://www.youtube.com/watch?v=Z5RUVyPKkUg
I had a really good mentor when working on this product and I learn the importance of modularity in design. I was just off my internship doing a D2K/Oracle implementation for an airline in-cabin inventory application, so working on Spooler was a breath of fresh air.
[1] https://support.hpe.com/hpsc/doc/public/display?docId=emr_na...
The last high-performance design I actually liked was the DEC Alpha. You could write a useful JIT compiler in a couple hundred lines.
I suspect that nVidia's recent GPUs are wonderfully clever inside, but they don't publish their ISA and the drivers are super-clunky. So I can't admire them.
I appreciate the performance of intel Core chips, but there's so much to dislike. The ISA is literally too big to fully document. The kernel needs 1000s of workarounds for CPU weirdnesses. You have to apply security patches to microcode, FFS.
RISC-V would be great if we had fast servers and laptops.
The average consumer doesn't buy a fantastically well designed CPU if it doesn't run the software they care about. x86, externally, is horrifically ugly primarily because of backwards compatibility (I've legitimately had a nightmare once from writing an x86 JIT compiler). Internally, I'm almost certain it's an incredible feat of engineering. People who admire architecture aren't a powerful market force, I'm sad to say.
On the other end, pure CPU only machines are kind of interesting as a study in economy, like the ZX Spectrum, a horrible, limited architecture that managed to hit the market at an unreasonably cheap price, make money, and end up with tens of thousands of games.
Overlooked by many (but not all) was its built-in MIDI ports and its abilty to control all those early model beatboxs and synths...unfortunately I forget the name of the rather crappy software that I used to get things talking and synced up, but it did work and the bitmapped color graphics were way ahead of its time.
Too bad Atari self destructed with the cartrage business, and who knows what other poor business decisions it made, but that computer was one of my favorite things of my late teens.
The Burroughs B5000 was the commercial fountainhead of this philosophy (High-Level-Language Computer Architectures), but today there is no significant commercial descendant of this 1960s radical.
Larrabee was cute, but to this day I still have no idea what their target workload was.
I haven't fired up my Cell dev board (Mercury) in a while. Prolly should do that. :-)
It is also notable that quite a bit of R&D was done in Chippewa Falls, WI, which is just a regular old town in America's Dairyland.
One of the MUD admins was astounded when he noticed I was playing from what he thought was a real Cray Y/MP. If I was smarter I would have played along. Alas...
Funny enough the Y/MP-48 was running Unicos (Cray's mostly-Unix) and someone had compiled nethack or rouge or something similar and was playing it until he had sucked up all the funny-money the department had budgeted for the year (yeah...we got accounted for CPU time, and yeah...someone screwed up the quotas). There was a kerfuffle....
The funny thing was...no one told me the secret and I musta spent 5 long hours pulling out my hair.
If I remember right, one of the RAs running the computer center finally had pity on me and showed me what was going on...
I'm not really sure I can call them the "good ole days of programming", but that was how things were done back then.
As total systems I loved the HP 41CX, the Sinclair ZX Spectrum, the Symbolics Lisp Machine and the Apple Mac IIcx (or really just any Mac before the PowerPC debacle).
After that era, I just started home-building x86 machines, and while there was the odd preferred component, it never went beyond the 'A is better than B' stage.
Still, cult chip! I mean, something like the following shows obsessive dedication to the thing:
Is there an equivalent pitfall in designing the ISA to support a specific Virtual Machine?
For example, wouldn't the performance of a server processor when running the Java Virtual Machine be a key factor in determining its commercial success? I've always wondered whether the failure of Itanium wasn't at least partly caused by the shift from binary executables to bytecode with the contemporary success of the Java language. Even when JIT compilers were used, they were probably too simple to take advantage of the VLIW architecture.
Not sure if that's the exact case for Itanium but your argument fired a neuron. :)
He was a little strange, but I loved seeing his workshop and attempts at building such computers. It's a pity I'm not in touch any more, I would like to learn a bit more about what he was actually doing and if it was effective at all.
- CDC-6600 and 7600 - listed by Fisher, Sites, Smith, Worley
- Cray-1 - listed by Hill, Patterson, Sites, Smith, Sohi, Wallach (also Bell, sorta)
- IBM S/360 and S/370 - listed by Alpert, Hill, Patterson, Sites (also Bell)
- MIPS - listed by Alpert, Hill, Patterson, Sohi, Worley
Special mention:
- 6502 - only listed by Wilson, but she was the chief architect of ARM so i think her choice is important to note
- Itanium - mentioned in the top-ranked comment in this HN discussion
- DEC Alpha - mentioned in the second-ranked comment in this HN discussion
And the M68K powered my Atari ST, my friends Amigas, and so much more.
Makes me wish I could get a small ZX Spectrum, Atari ST, and Amiga hardware kit that would interface with USB HID stuff and output displayport/hdmi. (rather than software emulation on a Raspberry Pi)
32 bit system from the late 1970s.
I believe the design was influenced by Djikstra's Structured Programming book but have no evidence.
My epiphany on the issues with the isa came when I discovered that the checksum calculation used by the VMS backup utility was faster when done in a short instruction loop over the microcoded instruction. MicroVAX II. Microcodes were a huge barrier between the speed potential of the electronics and the actual visible isa. Duh!
Cray knew this, but he didn't build product lines, just single point products. Sun built product lines with RISC and ate Digital Equipment's lunch.
If you haven't yet do it now, check k8s out, get knee-deep into it and not from a devops perspective but as one who admire computers.
It's a little sad to see my previous career disappear, I'd love nothing more than to manage some VMware clusters running Linux VMs for the rest of my life, but technology moves on. It's forcing me to become a coder and manager rather than sysadmin.