CISC-y RISC-ness
tedium.co
tedium.co
On single core microcontrollers, all that carefully timed assembly ends up interleaved with other logic in ways that make it nearly impossible to untangle for reuse.
The article makes a big deal about CISC vs. RISC, but in the end all of the CISC architectures ended up being RISC under the hood with front end instruction translation for the more complex CISC instructions. The transition to 64 bit should offer chip manufacturers a golden opportunity to ditch old disused complexity, like multiple addressing modes and variable length instructions.
> The article makes it pretty clear why Linux can't run directly on the Crusoe: Linux expects the hardware to have a virtual memory manager, which the Crusoe doesn't have. Consequently, any port of Linux will need to be running on an emulated memory manager.
As a side note, the Crusoe is also missing native support for certain other helpful features: Memory protection -- without that, a segfault can take out the entire OS. Running code from user memory -- without this, any application code will need to be piped through the OS to the CPU.
[0] https://developers.slashdot.org/comments.pl?sid=96231&cid=82...
Some people did do a bit of dive into the raw architecture (https://www.realworldtech.com/crusoe-exposed/) but I don't think anyone ever really tried to create a compiler for it, and it would be a moving target anyway as the different Transmeta CPUs seem to use different instruction sets widths (128 bit wide on Crusoe, 256 bit on Efficeon).
An interesting fact was that CMS could occasionally (not always, but not rarely either) do a better job than the ahead-of-time compiler as CMS benefits from: statistics collected at runtime and can generate code under assumption without compensation code. When the assumptions are found to be violated, it can just recompile the code again.
The challenge is always having the compilation overhead be [more than] paid off by the runtime gain. Sometimes it worked, others it didn't. Unfortunately, people really notice when it doesn't work which added to a negative perception of Crusoe. The conventional processors (even out-of-order) have far more predictable performance.
vendor_id : GenuineTMx86
model name : Transmeta(tm) Crusoe(tm) Processor TM5800
cpu MHz : 800.023
flags : fpu vme de pse tsc msr cx8 sep cmov mmx longrun lrti constant_tsc cpuid
bugs : cpu_meltdown spectre_v1 spectre_v2 spec_store_bypass l1tf mds swapgs itlb_multihit mmio_unknown
But you can't do much with such a machine nowadays.TM5800 might be subject to Spectre, but I posit that it would be very tricky to exploit as the x86 code you write isn't what's being run and in fact the code execution will look vastly different.
A lot of the IP that people had to create involved getting around Intel's patents, some of which were both basic and likely bogus (prior art etc) but Intel was widely seen as having far more lawyers than anyone could afford so taking them on directly was best avoided
Not really mentioned in this article was TM's biggest score at the time, which was getting Linus to come work for them
https://www.realworldtech.com/crusoe-intro/
To my knowledge the foreshadowed followup was never released.
If Transmeta had survived, perhaps they could have added support for different instruction sets to run e.g. Java or .NET "natively".
I never heard from them again.
https://millcomputing.com/topic/what-is-your-roadmap-for-202...
i don't have any complaint about people writing about things they don't know very much about; it's a great way to learn. i just don't think they should put on this fake authoritative tone, because they can mislead other people. i mean that's how we got the cda and dmca
We did look at handwriting assembly code, since we were already intimately familiar with the classic TI 320C30s and C40s (fantastic DSPs!) but it was just so crazy. Even TI told us not to bother and wait for their super cool compilers.
* Crusoe/Efficeon/... used a software JIT (called CMS)
* Ran on a very custom architecture (P95/P2000) designed for JITted code with speculation (software controllable checkpoints, "assert" instructions, speculative cache, small lookup tables for translating x86 addresses to P95 etc).
* It had no dynamic scheduling, reorder buffer, renaming, etc.
The thesis was to eliminate the power/area hungry parts of microprocessors by having the "compiler" (= JIT) do it up front.
Some things worked well (very power efficient, small die). Others less so: huge cold code penalty, limited ability to deal with large cache misses (like IA-64).
Transmeta had a huge impact on the industry; Intel has officially credited Transmeta and Transmeta's LongRun with getting Intel focused on power. Both NVIDIA and "another company" have explored the CMS idea.
There is so much to the Transmeta story and the failure was more of a business issue than a technical one. This was a completely new approach and Transmeta needed more iterations to refine the idea.
ADD: A VLIW "instruction" (called a "molecule" in TM terminology) is compiled, fetched, issued, and (mostly) executed as a (parallel) unit. That is very very different from what modern superscalar microprocessors do. They issue instructions from many issue queues. The instructions that execute in the same cycle usually do so mostly as a consequence of data-flow and the current state of the pipeline. Large x86 instruction translated to microcode typically leads to many µops that do not execute in parallel, but are intermixed with other µops.
Honestly, if you look at recent parts, the internal designs are _super_ wide multi-issue like a VLIW, it's just a difference of sophistication vs. coupling in your JITy thing.
Let's use Intel Sunny Cove as an example because it's recent and there is a good diagram on Wikpedia: https://en.wikipedia.org/wiki/Sunny_Cove_(microarchitecture)
Because they're wide multiple issue (SC: 8 execution units wide), the internal execution is fairly VLIW-like. Unlike an exposed VLIW, it's plausible possible to actually fill all those pipes because they're doing dynamic out-of-order issue on a window of (sort of hard to count, SC:something like 50) instructions coupled to all the register renaming and memory access scheduling to keep instructions out of each others way.
Transmeta's proposition was not really super wild, Multiflow (and very briefly Apollo) were building VLIWs in the 80s, so that wasn't crazy. Their core trick was splitting the difference between the dynamic decomposition driven out of order/multi issue/superscalar designs that used relatively dumb, shallow heuristics but were very dynamic and close to the hardware (like above), and the compiler-driven RISC/VLIW designs that tried to do fancier scheduling and optimization on larger units but were much more static and more removed from the execution process.
The fast dumb heuristics close to the hardware won in a _big_ way over all competitors.
https://joyoftech.com/geekycomics/Aftery2k/y2Karchives/125.h...
https://joyoftech.com/geekycomics/Aftery2k/y2Karchives/158b....
By the way, it's interesting to read the iAPX 432 architecture paper along with the Case for RISC paper, since the 432 paper argues the exact opposite. Essentially, due to increasing software costs, you should put as much as possible into hardware. The semantic gap between hardware and programming languages should be minimized by implementing high-level things such as objects and garbage collection in hardware, which they call the Silicon Operating System. The instruction set should be as complete as possible, with lots of data types.
I should also mention that the 432's instructions were not byte-aligned: they were anywhere from 6 to 321 bits long. Just decoding the instructions took a complex chip, with a second chip to run them. This is the opposite of the RISC idea to make instruction decoding as simple as possible.
The paper says, "The iAPX 432 represents one of the most significant advances in computer architecture since the 1950s."
http://www.bitsavers.org/components/intel/iAPX_432/171821-00...
https://en.wikipedia.org/wiki/IBM_AS/400
Of course Java provided a similar architecture in many ways on a mainstream architecture the same way that Common Lisp made Lisp machines obsolete. Java never really got a universal architecture for serialization and persistence the way the AS/400 did.
That said, I find AVR-8 pretty interesting in that it is the last 8-bit architecture and is, from the viewpoint of the assembly programamer, technically superior to all those machines I had or craved in the 1980s.
RISC-I probably hadn't been written about publicly when he submitted that paper.
'RISC I: A Reduced Instruction Set VLSI Computer' appeared in May 81
https://dl.acm.org/doi/10.5555/800052.801895
(and of course the original RISC paper was even earlier in 1980)
The iAPX 432 paper in June 82:
https://dl.acm.org/doi/10.1145/641542.641545
I should add that I only skimmed the 432 paper for references so I might be wrong on that!
It's pretty ballsy of Intel's marketing department to so blatantly lie to your face like that. Maybe they were big believers of the "Big Lie" rhetorical device? If you say something so outrageous people are less likely to question it because you must be a genius with loads of inside knowledge to assert something so completely crazy on the face of it.
It was just a low end, low power CPU. It ran standard X86 code, and I had Linux running on it. I'd say it was still not ideal, as the system I had had a CPU fan and the inside of the case got pretty hot. I recall the CPU supported crypto and RNG acceleration, which I think was a nice and fairly rare feature at the time.
Other than that, VIA's hardware sucked. The CPU is unimpressive and the board I got had an embedded NIC that corrupted network packets. That one was fun to figure out.
Edit: also the "cool chip in a can" packaging from one of the contemporary computer conventions. And possibly the mini-ITX era was also kickstarted with a C3-powered motherboard, I think...?
Unfortunately I developed a searing hatred of VIA hardware back then. Besides the slow CPU and buggy board, I also managed to run into the SB Live + VIA chipset = disk corruption issue.
My original Macbook Air ran its Core2Duo at the rated 1.6 GHz for about 5 seconds before dropping to 1.2 GHz. Thirty seconds later it dropped to 1.0 GHz, and a couple of minutes later to 800 MHz where it could run all day.
In 2019 I had at the same time two machines with i7-8650U CPU: a NUC and a Thinkpad X1 Carbon. The NUC could build riscv-gnu-toolchain (which I did many times a day, as I was helping develop new RISC-V instructions) in 20 minutes with mild throttling (to 3.2 GHz) while the Thinkpad took over 30 minutes with extreme throttling.
Both amd64's "extreme instruction size variability" and arm64's "extreme only one instruction size" philosophies give worse results than riscv64's "two instructions lengths make small code that is still easy to parse in a wide machine". You'd have thought Arm of all people would have known that, but oh well...
Source: looking at the output of "size" on binaries in the the last couple of years of Fedora / Debian / Ubuntu etc versions that support all three.