The SuperH-3, part 1: Introduction (2019)
devblogs.microsoft.com
devblogs.microsoft.com
> The SH-4 is probably most famous for being the processor behind the Sega Dreamcast.
It was always amusing to visit Sega HQ back in those days and see how all-in they were on Hitachi: from the Hitachi elevator in the building to the conference rooms with Hitachi OHP, Hitachi wall clock, Hitachi conference table, Hitachi pens, hitachi computers and of course Hitachi chips... Those huge conglomerates (Keiretsu) are crazy!
Super hot like they're great chips, or super hot like they're space heaters that incidentally do math?
Hitachi apparently had a trademark on "Cool Engine" for the SuperH (https://www.cpushack.com/CIC/embed/announce/HitachiSH7709.ht...).
One thing I remember was how good Hitachi's documentation is! It was clear and concise and any quirk was well explained. The pseudo-code they used to describe each instruction was also good, C-like to the point that I was able to copy-paste the snippets with minor changes to write an emulator.
However the architecture itself was a bit annoying to work with because of the small opcodes. Doing something as simple as loading a 32bit immediate to a register ended up being 7-9 instructions. More realistically the compiler would put the values in LTORG and we'd load them from there in 1 or 2 instructions. Relative jumps were also an issue for the same reasons. It also inherited the infamous delay slots where half the instructions are illegal and the other half is quirky so we just NOPed them but I know that smarter compilers do useful work there.
That being said I do admire the elegance of how they fitted their 100 instructions with 16/24 registers in a 16bit opcode, it's a beautiful application of engineering trade-offs :).
mov r0, r3 <-result from a subroutine
mov r3, @r4
mov @r4, r3
do something with value.
And many many more examples like this. Wasting cycles and space.
The SH has addressing modes for postincrement reads and predecrement writes, so something like
var = *ptr++;
can be compiled down to a single instruction. But GCC almost always does something moronic. I've seen it generate code equivalent to: ptr++; tmp = ptr; tmp -= 16; var = tmp[15];
Insane.Older versions of GCC (3.x) are generally much better at low level instruction selection, but newer versions are better at big-picture stuff like optimizing across function calls, so speed-wise, across a whole program, they aren't too dissimilar. Newer versions are still worth it for the feature set (better warnings, modern standards like C23/C++23, etc).
Its not obvious what they trade off and got in return.
If you read through all 10 parts of the series, Raymond later mentions that it might look like a lot of code but instructions are half as big as other RISC contemporaries so it's better than it looks.
That's a bit of a gimmick, tbh. Yes you can just about make it work if you have a really barebones insn set, as the early SuperH chips had. (You do need to start from a 2-register insn format of course.) But modern chips - even quite small chips nowadays - benefit from having more than that, so you're left with a really limited niche.
(I do think that the calculus changes if you have a strictly Harvard architecture chip/core, since these might be able to afford a "weird" instruction length. A core with, e.g. 20-bit or 24-bit fixed-length instructions designed along the same lines gains some much needed space for ISA extension.)
There's a j32 & j32smp core now, but I'm not sure where-ish that would slot into the old roadmap, and it seems like development isnt active. https://j-core.org/roadmap.html
Meanwhile support keeps bitrotting. Someone managed to get a Debian port going in 2015, & it sounded semi adventurous to make go. https://lwn.net/Articles/647636/
Chipmaker Renesas - who kind of was super-h for a while - meanwhile seems to be starting down the RISC-V path, alongside a variety of ARM cores & their own (small microcontrollers) RL78 architecture that I don't know much about. https://www.cnx-software.com/news/renesas/
http://wiki.netbsd.org/ports/dreamcast/
GXemul is what I used at the time.
https://gavare.se/gxemul/gxemul-stable/doc/machine_dreamcast...
There are also lots of interesting things going on with NAOMI emulation on the gaming side. All of it is very hackable and j-core would have made a great addition to that.
If you’re interested in SuperH Linux, feel free to join #linux-sh on Libera IRC.
Had fun almost 20 years ago with it during my internship at Ricoh : they were porting Linux on it (mostly as a research, the device was already nearing its EOL, but it had a touch screen with a stylus and 2 PCMCIA ports, which made it possible to put a WIFI card on it) and a made some small demonstration programs. Spent a lot of my time fighting to manage to make libs compile on it ^^;
I actually got interested in SH3 again recently as one of my viewers donated me some hardware to see if I could port Gentoo to run on it. It's more a passion project however I did manage to get a PoC image to semi boot over a weekend so I do hope I'll manage to get this into a state that I can merge this into Gentoo to officially support it again.
I’m one of the kernel maintainers for SuperH Linux.