I'm fascinated by RTX2010 [0], a radiation-hardened processor which uses Forth and has been deployed on several space missions.
I'm fascinated by RTX2010 [0], a radiation-hardened processor which uses Forth and has been deployed on several space missions.
This is a good read on the subject: https://uwspace.uwaterloo.ca/bitstream/handle/10012/10810/La...
Also, check out the GA144: https://www.greenarraychips.com/
Also, it's not clear a super scalar stack machine would be possible. So it might not be suitable for replacing x86 soon ;)
The "Boost: Berkeley's Out-of-Order Stack Thingy" (Steve Sinha, Satrajit Chatterjee and Kaushik Ravindran) [1] effectively translates stack code into RISC as part of instruction decoding and issue. This is similar to how "Design of a Superscalar Processor Based on Queue Machine Computation Model" (Shusuke Okamoto, Hitoshi Suzuki, Atusi Maeda, Masahiro Sowa) [2] does the same thing but for a queue machine instead of a stack machine.
[1]: https://citeseerx.ist.psu.edu/document?repid=rep1&type=pdf&d...
quite aside from the new results it presents, it seems to be a much more comprehensive answer to the question of why stack machines fell out of favor than i've seen before
The size was a concern because of the limited size (16kbyes) of bram modules on an fpga.
I can't say why it performs so much better though, maybe just because it's simpler & 16 bit so it's smaller and can be clocked faster on the fpga? That is just a guess though.
I'm not sure if it is fair to compare the J1 to the BOOM, the latter does a lot more I'd think. That being said, it shows that for such use cases a stack machine can be better suited and much less complex to design, maintain and work with.
I too found Charle's Moore's cpu designs interesting. One possibility I thought about for a bit was a VLIW stack hybrid, where you'd have bundles of say 8 stack instructions each executing concurrently in their own lane, and some instructions for moving values between lanes. It'd be a weird thing to program for, but potentially quite fast for how minimal it'd be.
That's an incredible design I would love to get my paws on.
I was musing about this stuff back when reading Hennessy and Patterson in the early 00's. Moore's homepage was really interesting and inspiring, that someone could work idiosyncratically and largely individually and actually get chips to tape out. Cool to hear he's still going strong.
Seems to me registers are a fancy hardware implementation detail, that software at higher abstraction can and should avoid unless maximum performance is required. It was surprising for me to learn recently that on modern x86 CPUs, named registers actually reference hidden, dynamically allocated physical registers, a bit like virtual memory is an indirection layer over physical memory.
It's levels of indirection all the way down...
If you wanna learn about how register renaming works, look into Tomasulo's algorithm. State of the art chips are a lot more complex, but that's a good historic introduction to the fundamental ideas.
Personally, I’d say it’s at least as much of a register machine as a stack machine and ways in which it is one vs the other at times sure had me running in circles.
I’ve seen plenty of quotes about how it’s not supposed to be a real assembly, how you’re not supposed to deal with its text format or bytecode, and how it and its documentation are for compiler writers not typical programmers.
Which is exactly the take that I think is so wrong with our status quo of opaque, overly abstracted technologies that are highly optimized in their parts yet so disappointing in the whole.
As a Forth substitute, I was sorely disappointed. I did learn a lot in my wanderings, and what WebAssembly is mostly is a an implementation to run a C architecture. Which is pretty much what it came from. Not that I’m against that, but I sure ran in circles trying to make it more Forth-like.
And given it’s raison d’etre to be more performant vs Javascript it seems to be faster in some benchmarks while slower in others. I think this is more a testament about how optimized V8 and its brethren are then a slam on WebAssembly.
It's closeness to Javascript and the fact that it's basically managed by companies that want to keep Web as complicated as possible so they can enjoy their monopoly on it, turns me off big time.
And lastly, the only thing it has in common with Forth, is being stack-based. Computing has become an unsustainable pile of buggy abstractions on top of one another, and WASM is simply another layer on top this pile of dung, rather than a return to the basics. I get why it wants to give programmers nothing more than a rounded pair of scissors, but then for a glorified VM it's incredibly over-engineered. Art has never been created by a committee of commercial entities.
I had in mind to write a OS that runs WASM natively, but these days I'd rather put a persistent, programmable Forth interpreter that gives you full access to the hardware at the lowest level, and build up from there.
https://news.ycombinator.com/item?id=34374057
>This is a great article! I love his description of Forth as "a weird backwards lisp with no parentheses".
>Reading the source code of "WAForth" (Forth for WebAssembly) really helped me learn about how WebAssembly works deep down, from the ground up.
>It demonstrates the first step of what the article says about bootstrapping a Forth system, and it has some beautiful hand written WebAssembly code implementing the primitives and even the compiler and JavaScript interop plumbing. We discussed the possibility of developing a metacompiler in the reddit discussion.
[...]
Check out this beautiful hand written wasm code, a dynamic wasm based Forth compiler:
https://github.com/remko/waforth/blob/master/src/waforth.wat
It compiles new Forth words into teenie tiny little wasm modules, and hooks out through JavaScript in the browser to create and link them in to the main module, so it can call the code it just compiled. Wasm doesn't let you modify compiled code, but you can compile new code and link it in and execute it!
> a stack-based VM is relatively easy to optimise for a register CPU.
You get SSA for almost free (just phi the stacks at control joins), but after making SSA everything is the same regardless of stack- or register- VM.
[0] https://academic.oup.com/comjnl/article/6/4/308/375725?login...
https://news.ycombinator.com/item?id=8860786
https://news.ycombinator.com/item?id=14469113
https://news.ycombinator.com/item?id=34561910
Rudy Rucker writes about his CAM-6 in the CelLab manual:
http://www.fourmilab.ch/cellab/manual/chap5.html
Computer science is still so new that many of the people at the cutting edge have come from other fields. Though Toffoli holds degrees in physics and computer science, Bennett's Ph.D. is in physical chemistry. And twenty-nine year old Margolus is still a graduate student in physics, his dissertation delayed by the work of inventing, with Toffoli, the CAM-6 Cellular Automaton Machine.
After watching the CAM in operation at Margolus's office, I am sure the thing will be a hit. Just as the Moog synthesizer changed the sound of music, cellular automata will change the look of video.
I tell this to Toffoli and Margolus, and they look unconcerned. What they care most deeply about is science, about Edward Fredkin's vision of explaining the world in terms of cellular automata and information mechanics. Margolus talks about computer hackers, and how a successful program is called “a good hack.” As the unbelievably bizarre cellular automata images flash by on his screen, Margolus leans back in his chair and smiles slyly. And then he tells me his conception of the world we live in.
“The universe is a good hack.”
[...]
Margolus and Toffoli's CAM-6 board was finally coming into production around then, and I got the Department to order one. The company making the boards was Systems Concepts of San Francisco; I think they cost $1500. We put our order in, and I started phoning Systems Concepts up and asking them when I was going to get my board. By then I'd gotten a copy of Margolus and Toffoli's book, Cellular Automata Machines, and I was itching to start playing with the board. And still it didn't come. Finally I told System Concepts that SJSU was going to have to cancel the purchase order. The next week they sent the board. By now it was August, 1987.
The packaging of the board was kind of incredible. It came naked, all by itself, in a plastic bag in a small box of styrofoam peanuts. No cables, no software, no documentation. Just a three inch by twelve inch rectangle of plastic—actually two rectangles one on top of the other—completely covered with computer chips. There were two sockets at one end. I called Systems Concepts again, and they sent me a few pages of documentation. You were supposed to put a cable running your graphics card's output into the CAM-6 board, and then plug your monitor cable into the CAM-6's other socket. No, Systems Concepts didn't have any cables, they were waiting for a special kind of cable from Asia. So Steve Ware, one of the SJSU Math&CS Department techs, made me a cable. All I needed then was the software to drive the board, and as soon as I phoned Toffoli he sent me a copy.
Starting to write programs for the CAM-6 took a little bit of time because the language it uses is Forth. This is an offbeat computer language that uses reverse Polish notation. Once you get used to it, Forth is very clean and nice, but it makes you worry about things you shouldn't really have to worry about. But, hey, if I needed to know Forth to see cellular automata, then by God I'd know Forth. I picked it up fast and spent the next four or five months hacking the CAM-6.
The big turning point came in October, when I was invited to Hackers 3.0, the 1987 edition of the great annual Hackers' conference held at a camp near Saratoga, CA. I got invited thanks to James Blinn, a graphics wizard who also happens to be a fan of my science fiction books. As a relative novice to computing, I felt a little diffident showing up at Hackers, but everyone there was really nice. It was like, “Come on in! The more the merrier! We're having fun, yeeeeee-haw!”
I brought my AT along with the CAM-6 in it, and did demos all night long. People were blown away by the images, though not too many of them sounded like they were ready to a) cough up $1500, b) beg Systems Concepts for delivery, and c) learn Forth in order to use a CAM-6 themselves. A bunch of the hackers made me take the board out of my computer and let them look at it. Not knowing too much about hardware, I'd imagined all along that the CAM-6 had some special processors on it. But the hackers informed me that all it really had was a few latches and a lot of fast RAM memory chips.
(i think your text means that rucker tried the cam-6, not that you did, but i'm not entirely sure)
I played with the CAM-6 in Norman Margolus's office at MIT, and a friend of mine who worked for him brought one to a science fiction convention where we tripped out on it all night in a hotel room! I saved a copy of the floppies full of Forth code. (linked above)
Flickercladding:
https://www.fourmilab.ch/cellab/manual/rules.html#Flick
>Flick is named after “flickercladding,” the CA skin which covers the robots in my books Software and Wetware. In Flick, we see an AutoShade®d office whose rug is made of flickercladding that runs the TimeTun rule. You can tell which parts of the picture are “rug” because these cells have their bit #7 set to 1.
Flickercladding Interior Decoration
Conceived by Rudy Rucker
Drawn by Gary Wells
Modeled with AutoCAD
Rendered by AutoShade
Perpetrated by Kelvin R. Throop.
In this rule, we only change the cells whose high
bits are on. These cells are updated according to
the TimeTun rule.
http://www.technovelgy.com/ct/content.asp?Bnum=299>Some looked humanoid, some looked like spiders, some looked like snakes ...All were covered with flickercladding, a microwired imipolex compound that could absorb and emit light.
https://en.wikipedia.org/wiki/Wetware_(novel)
>The plot goes disastrously awry, and a human corporation called ISDN retaliates against the boppers by infecting them with a genetically modified organism called chipmold. The artificial disease succeeds in killing off the boppers, but when it infects the boppers' outer coating, a kind of smart plastic known as flickercladding, it creates a new race of intelligent symbiotes known as moldies — thus fulfilling Berenice's dream of an organic/synthetic hybrid.
https://news.ycombinator.com/item?id=15546769
DonHopkins on Oct 25, 2017 | parent | context | favorite | on: Boustrophedon
The Floyd Steinberg error diffusion dithering algorithm can use a boustrophedonous scan order to eliminate the diagonal geometric artifacts you get by scanning each row the same direction.
https://en.wikipedia.org/wiki/Floyd%E2%80%93Steinberg_dither...
"In some implementations, the horizontal direction of scan alternates between lines; this is called "serpentine scanning" or boustrophedon transform dithering."
I implemented some eight bit cellular automata heat diffusion rules with error diffusion, which accumulated an unfortunate drift up and to the right because of the scan order.
Rudy Rucker pointed out the problem:
https://web.archive.org/web/20180909074032/http://donhopkins...
"Rudy Rucker: I feel like you might have some kind of bug in your update code, an off-by-one thing or a problem with the buffer flipping. My reason is that I see persistent upward drift in the action, like if I mouse drag a blob it generally moves up. Also the patterns appearing in the blob aren't uniform. I mean...this IS supposed to be the 2D Rug rule, isn't it?"
So instead of scanning back and forth boustrophedoniously (which wouldn't eliminate the vertical drift, just the horizontal drift), I rotated the direction of scanning 90 degrees each frame ("spinning scan") to spread the drift out evenly in all directions over time.
https://github.com/SimHacker/CAM6/blob/master/javascript/CAM...
// Rotate the direction of scanning 90 degrees every step,
// to cancel out the dithering artifacts that would cause the
// heat to drift up and to the right.
That totally canceled out the unwanted drifting and geometric dithering artifacts! That made it possible to cultivate much more subtle (or not-so-subtle) effects, like dynamically switching per-cell between different anisotropic convolution kernels (see the "Twistier Marble" rule for an extreme example).http://donhopkins.com/home/CAM6
CAM6 Demo:
https://www.youtube.com/watch?v=LyLMHxRNuck
Demo of Don Hopkins' CAM6 Cellular Automata Machine simulator.
Live App: https://donhopkins.com/home/CAM6
Github Repo: https://github.com/SimHacker/CAM6
Javacript Source Code: https://github.com/SimHacker/CAM6/blob/master/javascript/CAM...
Comments from the code:
// This code originally started life as a CAM6 simulator written in C
// and Forth, based on the original CAM6 hardware and compatible with
// the brilliant Forth software developed by Toffoli and Margolus. But
// then it took on a life of its own (not to mention a lot of other CA
// rules), and evolved into supporting many other cellular automata
// rules and image processing effects. Eventually it was translated to
// C++ and Python, and then more recently it has finally been
// rewritten from the ground up in JavaScript.
// The CAM6 hardware and Forth software for defining rules and
// orchestrating simulations is thoroughly described in this wonderful
// book by Tommaso Toffoli and Norman Margolus of MIT.
// Cellular Automata Machines: A New Environment for Modeling
// Published April 1987 by MIT Press. ISBN: 9780262200608.
// http://mitpress.mit.edu/9780262526319/https://donhopkins.com/home/cam-book.pdf
CAM6 Simulator Demo:
https://www.youtube.com/watch?v=LyLMHxRNuck
Forth source code for CAM-6 hardware:
This isn't a CAM simulator, but I once wrote a little hack of a CA playground inspired by it. Table-driven in a similar way, user-scriptable with JS, includes neighborhoods like the Margolus neighborhood. https://github.com/darius/js-playground/blob/master/ca.js and the associated ca.html.
for out-of-order and superscalar processing, register machines have an advantage in that you can more easily have successive instructions that don't interfere, simply by not using the same registers. generally in a stack machine instruction it's hard to avoid using the result of the previous instruction; you need an extra instruction to do it (`over` or `swap` or something)
but i suspect it's partly just path-dependence: koopman's new wave of stack machines were already struggling upwind against 30 years of ibm 360, intel 8080, dec pdp-11, motorola 68000, dg nova, sun/fujitsu sparc, etc., all register machines, and so a rich store of knowledge had built up about them in compiler backends, assembly programmers, and cpu designers. maybe if chuck had designed the rtx2010 in 01973 instead of 01988 the story would have been different
maybe worth noting that all of the jvm, microsoft's clone of it (the cil), and wasm (as sph pointed out) are stack-based, and arm for a while had a 'jazelle' instruction set which interpreted many jvm bytecodes directly (trapping to software for more complex ones), so in a sense stack-based programs are more widespread now than they ever were before; that's what you see inside a .class file. they just are usually translated into some other instruction set before execution
also several people have pointed out that the stack machines most commonly used in practice, including the ones i mentioned above, are hybrids; they have architectural registers (often called local variables) but you have to copy them onto the stack to operate on them