The Soul of an Old Machine: Revisiting the Timeless von Neumann Architecture
ankush.dev
ankush.dev
https://www.goodreads.com/book/show/7090.The_Soul_of_a_New_M...
which was one of the first computer books I ever read --- I believe in an abbreviated form in _Reader's Digest_ or in a condensed version published by them (can anyone confirm that?)
EDIT: or, maybe I got a copy from a book club --- if not that, must have gotten it from the local college after prevailing upon a parent to drive me 26 miles to a nearby town....
One of my favorite quotes in the book - when an overworked engineer resigns from his job at DG. The engineer, coming off a death march, leaves behind a note on his terminal as his letter of resignation. The incident occurs during a period when the microcode and logic were glitching at the nanosecond level.
[1] https://www.nybooks.com/articles/1982/02/18/man-and-supermin...
[1] http://www.mirrorservice.org/sites/www.bitsavers.org/compone...
I also highly recommend the TV show Halt and Catch Fire. It's not related to the book but very similar spiritually.
[0] https://bits.ashleyblewer.com/halt-and-catch-fire-syllabus/
The most surprising thing so far is how advanced the hardware was. I wasn't expecting to hear about pipelining, branch prediction, SIMD, microcode, instruction and data caches, etc. in the context of an early-80s minicomputer.
"He traveled to a city, which was located, he would only say, somewhere in America. He walked into a building, just as though he belonged there, went down a hallway, and let himself quietly into a windowless room. The floor was torn up; a sort of trench filled with fat power cables traversed it. Along the far wall, at the end of the trench, stood a brand-new example of DEC’s VAX, enclosed in several large cabinets that vaguely resembled refrigerators. But to West’s surprise, one of the cabinets stood open and a man with tools was standing in front of it. A technician from DEC, still installing the machine, West figured.
Although West’s purposes were not illegal, they were sly, and he had no intention of embarrassing the friend who had given him permission to visit this room. If the technician had asked West to identify himself, West would not have lied, and he wouldn’t have answered the question either. But the moment went by. The technician didn’t inquire. West stood around and watched him work, and in a little while, the technician packed up his tools and left.
Then West closed the door, went back across the room to the computer, which was now all but fully assembled, and began to take it apart.
The cabinet he opened contained the VAX’s Central Processing Unit, known as the CPU—the heart of the physical machine. In the VAX, twenty-seven printed-circuit boards, arranged like books on a shelf, made up this thing of things. West spent most of the rest of the morning pulling out boards; he’d examine each one, then put it back.
..He examined the outside of the VAX’s chips—some had numbers on them that were like familiar names to him—and he counted the various types and the quantities of each. Later on, he looked at other pieces of the machine. He identified them generally too. He did more counting. And when he was all done, he added everything together and decided that it probably cost $22,500 to manufacture the essential hardware that comprised a VAX (which DEC was selling for somewhat more than $100,000). He left the machine exactly as he had found it."
Perhaps the most important bit-flip of this paper’s time (and perhaps first fully realized in it) might be summarized as ‘instructions are data.’
This got me thinking: today, we are going through a bit-flip that might be seen as a follow-on to the above: after von Neumann, programs were seen to be data, but different from problem/input data, in that the result/output depends on the latter, but only through channels explicitly set out by the programmer in the program.
This is still true with machine learning, but to conclude that an LLM is just another program would miss something significant, I think - it is training, not programming, that is responsible for their significant features and capabilities. A computer programmed with an untrained LLM is more closely analogous to an unprogrammed von Neumann computer than it is to one running any program from the 20th. century (to pick a conservative tipping point.)
One could argue that, with things like microcode and virtual machines, this has been going on for a long time, but again, I personally feel that this view is missing something important - but only time will tell, just as with the von Neumann paper.
This view puts a spin on the quote from Leslie Lamport in the prologue: maybe the future of a significant part of computing will be more like biology than logic?
is an insight I've often seen attributed to von Neumann, but isn't it just the basic idea of a universal Turing machine? - one basic machine whose data encode the instructions+data of an arbitrary machine. What was von Neumann's innovation here?
> 6.8.5 ... There should be some means by which the computer can signal to the operator when a computation has been concluded. Hence an order is needed ... to flash a light or ring a bell.
(later operators would discover than an AM radio placed near their CPU would also provide audible indication of process status)
There's also audible noise that can sometimes be heard from singing capacitors[1] and coil whine[2], as mentioned in a sibling comment.
[1]: https://product.tdk.com/system/files/contents/faq/capacitors...
[2]: https://en.wikipedia.org/wiki/Electromagnetically_induced_ac...
Thus they're very prone to the effects mentioned.
The other two docking stations were built differently and had the wires run somewhere else.
EDIT: Typo
Sure enough - when I went onsite to assist, that "green" CRT was so incredibly wavy, I could not understand how she could do her job at all. First thing I tried was moving it away from the wall, in-case I had to unplug it to replace.
It stopped shimmering and shifting immediately. Moved it back - it started again.
That's when we realised that her desk was against a wall hiding major electrical connectivity to the regional Bell Canada switching data-centre on the floors above her.
I asked politely if she could have her desk moved - but no... that was not an option...
... So - I built my first (but not last) solution using some tin-foil and a cardboard box that covered the back and sides of her monitor - allowing for airflow...
It was ugly - but it worked, we never heard from her again.
My 2nd favourite was with GSM mobile devices - and my car radio - inevitably immediately before getting a call, or TXT message (or email on my Blackberry), if I was in my car and had the radio going, I would get a little "dit-dit-dit" preamble and know that something was incoming...
(hahahaha - I read this and realise that this is the new "old man story", and see what the next 20-30 years of my life will be, boring the younglings with ancient tales of tech uselessness...)
I remember this happening all the time in meetings. Every few minutes, the conversation would stop because a phone call was coming in and all the conference room phones started buzzing. One of those things that just fades away so you don't notice it coming to an end.
I haven’t thought about that in a long time.
Also, during this era I was on a flight and the pilot came over the PA right before pushing from the gate saying exasperatedly "Please turn your damn phones off!" (I assumed the same RF noise was leaking into his then-unsheilded headset).
The solution I picked was to put the machine in Text/Graphics mode (instead of normal character rom text mode, this was back in the MS-DOS days), so the vertical sync rate then matched up with the swirling EM fields, and there was almost zero apparent motion.
(Please feel free to reference the Monty Python “Four Yorkshiremen” sketch. But this really happened.)
The. I had to walk all the way back to the terminal room, and was a lot more careful the next time.
For some time I've been drawing parallels between ML/AI and how biology "solves problems" - evolution. And I also am bit disappointed by the fact the future might lead us in a different direction than mathematical elegance of solving problems.
You’re right though, we’re basically there with AI/ML aren’t we. I mean I guess we know why it does the things it does in general, but the specific “reasoning“ on any single question is pretty much unanswerable.
Neural network: created via training, does amazing things but (this is the area my knowledge is too small) it's bloody hard to decipher why is it doing the thing it does.
"Traditional" programming: you can (as in article) explain it down to the single most basic unit.
"The attribution of the invention of the architecture to von Neumann is controversial, not least because Eckert and Mauchly had done a lot of the required design work and claim to have had the idea for stored programs long before discussing the ideas with von Neumann and Herman Goldstine[3]"
The events are covered in great detail in Jean Jennings Bartik's autobiography "Pioneer Programmer", according to her von Neumann really wasn't that instrumental to this particular project, nor did he mean to take credit for things -- it was others that were big fans of his that hyped up his accomplishments.
I attended a lecture by Mauchly's son, Bill, "History of the ENIAC", he explains how eniac was a dataflow computer that, when it was just switches and patch cables, could do operations in parallel. There's a DH Lehmer quote, "This was a highly parallel machine, before von Neumann spoiled it." https://youtu.be/EcWsNdyl264
In this quote from Leslie Lamport, I took "If we don't" to mean "If we don't accept this". But the rest of the sentence made no sense then.
Could it be a really awkward way to say: If we don't "not accept this", i.e. if we accept this?
That is rather awkward isn’t it.
I really liked that whole block quote though.
It's my birthday today(61)... I'm up early to get a tooth pulled, and I read this wonderful story, and all I have is a tangent I hope some of you think is interesting. It would be nice if someone who isn't brain-foggy could run with the idea and make the Billion dollars or so I think can be wrung out of it. You get a bowl of spaghetti, that could contain the key to Petaflops, secure computing, and a new universal solvent of computing like the Turing machine, as an instructional model.
The BitGrid is an FPGA without all the fancy routing hardware, that being replace with a grid of latches... if fact, the whole chip would consist almost entirely of D flip-flops and 16:1 multiplexers. (I lack the funds to do a TinyTapeout, but started going there should the money appear)
All the computation happens in cells that are 4 bit in, 4 bit out Look up tables. (Mostly so signals can cross without XOR tricks, etc) For the times you need a chunk of RAM, the closest I got is using one of the tables as a 16 bit wide shift register, which I've decided to call isolinear memory[6]
You can try the online React emulator I'm still working on [1], and see the source[2]. Or the older one I wrote in Pascal, that's MUCH faster[3]. There's a writeup from someone else about it as an Esoteric Language[4]. I've even written a blog about it over the years.[5] It's all out in public, for decades up to the present... so it should be easy to contest any patents, and keep it fun.
I'm sorry it's all a big incoherent mess.... completely replacing an architecture is freaking hard, especially when confronted with 79 years of optimization in another direction. I do think it's possible to target it with LLVM, if you're feeling frisky.
[1] https://mikewarot.github.io/bitgrid-react-app/
[2] https://github.com/mikewarot/bitgrid-react-app
[3] https://github.com/mikewarot/Bitgrid
[4] https://esolangs.org/wiki/Bitgrid
[5] https://bitgrid.blogspot.com/
[6] https://bitgrid.blogspot.com/2024/09/bitgrid-and-isolinear-m...
[1] https://cloud.google.com/blog/products/ai-machine-learning/a...
[2] https://aws.amazon.com/blogs/aws/amazon-ecs-now-supports-ec2...
The goal of an FPGA is to get a signal through the chip in as few nanoSeconds as possible. It would be insane to rip out that hardware in that case.
However.... I care about throughput, and latency usually doesn't matter much. In many cases, you could tolerate startup delays in the milliseconds. Because everything is latched, clock skew isn't really an issue. All of the lines carrying data are short, VERY short, thus lower capacitance. I believe it'll be possible to bring the clock rates up into the Gigahertz. I think it'll be possible to push 500 Mhz on the Skywater 130 process that Tiny Tapeout uses. (Not outside of the chip, of course, but internally)
[edit/append]
Once you've got a working BitGrid chip rated for a given clock rate, it doesn't matter how you program it, because all logic is latched, the clock skew will always be within bounds. With an FPGA, you've got to run simulations, adjust timings, etc and wait for things to compile and optimize to fit the chip, this can sometimes take a day.
Low latency links are just byproduct, added to improve FPGA performance (which is definitely bad if compare to ASIC), you don't have to use them in your design.
If you will make working FPGA prototype, it would be not hard to convert them to ASIC.
Happy birthday!
I have no idea what you’re talking about but I’m gonna follow these links.
Actually, what was described is currently called the "Control Unit". We divide CPUs into a Datapath (which includes registers, busses, ALU...) and a Control Unit which is either hardwired or microcoded. Decoding instructions is the job of the Control Unit though it might get a little help from the Datapath in some designs.
The block diagram for a modern processor at the end of the article is just the Datapath - the Control Unit is not shown at all which is the most popular current style. It must actually exist in the chip and HDL sources, of course, if the processor is to actually do anything.
It is reasonable amount for CISC CPU cache, as for CISC normal to have small number of registers (RISC is definitely multiple-register machine, as memory operations are expensive with RISC paradigm).
A very common fallacy. For CISC it's normal to have small number of programmer-visible (architectural) registers. But nothing stops them from having large register files having about a hundred or two of physical registers which are renamed during the execution. IBM has been doing this sort of thing since the late sixties.
Modern x64 have IIRC 180 integer general-purpose registers, and I imagine this number can be increased further: say, AMD's Am29000 from 1985 had 192 registers... and honestly, having more than 16 programmer-visible registers is not that useful when 99.999% of the code is written for the high-level languages with separate-compilation schemes: any function call requires you to spill all your live data from the registers to the memory; having call-preserved registers means that you can do such a spill only once, at the start of the outer function, but it still has to happen. And outside of the function calls, in the real world programs most of the expressions that use only built-in operators either don't require more than 4 registers calculate, or can be done much more efficiently using vector/SIMD instructions and registers anyway.
You compare apples with carrots.
Question is what you name CISC/RISC. If under CISC considered early single-chip stubs like 8080 or even x86 before AMD64, this is one history, but if CISC are IBM/360,/370,/390 and DEC PDP-11 or later, this is totally other drama. Same is with RISC.
What really important, genuine CISC typically have human-friendly instruction set, and even /360 assembler is comparable to C language; many RISCs are load/store architectures, don't have rich memory addressing modes, so by definition tied to work inside registers. Also, when compare similar level of manufacture and similar timeline, will easy see, 68k series have just 8 registers, but same era RISCs have at least 16.
This analogy, I swear. Yes, I do, and this is an entirely reasonable comparison. Both are foodstuffs, and you can make pies out of both. But which pie would be healthier, and/or more nutritious? Now take the relative prices (or comparative difficulty of growing your own) of the ingredients into the account.
> 68k series have just 8 registers
Didn't they have 16?
And Western Electric 32000, a CISC processor from the same period, also had 16 32-bit registers.
It is your own opinion. Wikipedia have other definition.
> The CPU has eight 32-bit general-purpose data registers (D0-D7), and eight address registers (A0-A7).
And to the best of my knowledge, moving data from a data register to an address register and back did not destroy the top 8 bits.
Yes, you could use address registers for data storage, but you cannot make general computations on them, and because of this wiki definition don't account address registers in list of general purpose registers.
(There may be an effect because of instruction density, but that’s not very big, and will have less of an effect the larger your working set)
Ideally, your working set fits into the registers, and that was most win feature or RISC, having relatively simple hardware (if compared to genuine CISC, like IBM-360 mainframes) they saved significant space on chip for more registers.
- Register-register operations are definitely avoid memory wall, so data cache become insignificant.
Sure, if we compare comparable, not apples vs carrots.
I'd argue that MCUs, the highest volume of which are largely based on Harvard architecture, far outnumber von Neumanns.
As a native English speaker, it was understandable but feels rather foreign because you never hear parts of computers referred to that way these days.