SuperH
en.wikipedia.org
en.wikipedia.org
I cut my embedded development teeth writing device drivers and custom firmware targeting it.
The reason we used SH4 is because the Dreamcast had failed, so there was a huge surplus of them available on the market at the time.
It was easily one of the most interesting and rewarding projects of my life to work on.
It was obsoleted but we could get enough stock for the project. What we didn’t realize at the start was that the previous project had lower standards for development and testing and didn’t have the special debugger or great code. The debugger had to be special ordered for a ton of money and no vendors wanted to work on it.
Had a chance to revision up and changed to a pretty obscure atmega micro that was a perfect replacement. Did radiation screening again, development and testing..
That part was obsoleted too and I can’t use it for my next project. Will probably buy a space grade cpu for $20k this time. Radiation screening alone for a few components is $50k.
I've never sent any hardware into space, but some former colleagues of mine were working on OWL around the time I left: https://asd.gsfc.nasa.gov/archive/owl/science.html
I have a good friend however who is a great writer, and who has written a lot about her experiences, and who is working there again: https://www.jamiezvirzdin.com/
She didn't work much on the software/hardware side of things, however.
Maybe one day I'll sit down and put some of it to paper, I definitely have a lot of stories to tell.
It was extremely expensive, and commercial, I'm honestly surprised I could even find a website for it.
It was extremely cool to basically scroll through memory, our hardware had an on-board FPGA that gathered waveforms from a photomultiplier tube attached to a scintillator, and when I was writing the driver for it, I could watch the memory in real-time with the debugger changing.
Setting initial values of memory to something like 0xDEADBEEF, and using this thing, was an incredibly powerful debugging strategy, especially when dealing with hardware.
J2 open processor: an open source processor using the SuperH ISA - https://news.ycombinator.com/item?id=26866065 - April 2021 (45 comments)
The SuperH-3, part 15: Code walkthrough - https://news.ycombinator.com/item?id=20779622 - Aug 2019 (1 comment)
The SuperH-3, part 1: Introduction - https://news.ycombinator.com/item?id=20622921 - Aug 2019 (2 comments)
Building a SuperH-compatible CPU from scratch [video] - https://news.ycombinator.com/item?id=11886079 - June 2016 (24 comments)
Resurrecting the SuperH architecture - https://news.ycombinator.com/item?id=9812010 - July 2015 (15 comments)
https://m.youtube.com/watch?v=dVD1Yws__v0
(open source GPS!)
J2 open processor: an open source processor using the SuperH ISA - https://news.ycombinator.com/item?id=26866065 - April 2021 (45 comments)
Why the J-core open processor is cool - https://news.ycombinator.com/item?id=24163584 - Aug 2020 (1 comment)
J-Core Open Processor - https://news.ycombinator.com/item?id=20658584 - Aug 2019 (31 comments)
J-core Open Processor - https://news.ycombinator.com/item?id=12105913 - July 2016 (27 comments)
Building a CPU from Scratch: Jcore Design Walkthrough [video] - https://news.ycombinator.com/item?id=12101908 - July 2016 (8 comments)
A little ascetic for these days though, IMO.
You and everyone else, I feel. I tried to find some of my early projects from two decades back, and despite having backups floating around on various hard-drives, I could not find them to save my life. I wish I'd accepted the good word of CVS/SVN back then!
(Found this a while back, mild impostor syndrome moment much)
https://reviews.llvm.org/D94928
Other tools to do this exist (or have existed, Intel IACA rest in peace), but I'm slightly sceptical how useful they are in practice.
They can't model anything transient in the processor (can't easily, at least) like branch prediction and the memory hierarchy. That and there's only so much detail one can fit inside a model.
You know what would be better? If you had a decompression routine built into the CPU that would convert division and modulus commands into this loop at the micro-code level.
That way, a divide / modulus cycle (probably taking 20 instructions taking 50+ bytes) can be compressed into a singular instruction (1 instruction taking 4 bytes)... using less L1 cache.
I agree though in the general case, hence why I end with it being a little ascetic these days.
Was not a fan of the name SH, thought it was boring. Tried to get Hitachi to call it "Sonic" instead. As in (S)onic the (H)edgehog
SuperH competed with 68k and mips, perhaps? What do you think made the others more successful than SuperH? I suppose I dont know for sure that they were but I'll guess at least 68k lasted longer.
It was the first CPU I encountered that uses branch delay slots [1], meaning that the instruction immediately following a branch instruction is always executed, even when the branch is taken. That took a bit of getting used to, although I understand it's quite common on RISC architectures.
The SH-4 was most notably used in the Sega Dreamcast.
That’s been my impression of RISC-V in general. Whenever someone thinks another ISA does it better, there seems to be a very well thought-out reason for RISC-V’s decision, when you take into consideration that it’s built to scale from the smallest microcontroller to the biggest CPU.
The RISC-V standards is also evolving, if something isn’t there it can also be that it’s planned for future revisions. There’s a proposed revision coming up that would make RISC-V beat ARM in code density across the board, on real world embedded code.
And I think that RISC-V is right on repeating that design decision. ;)
It also looks like Alpha was introduced a bit earlier than SuperH and that makes me wonder why they still wanted branch delay slot. It is nothing but trouble across the board (tools, hardware design, etc) for couple of percents of execution speed. Which can easily be achieved just by using register bypass and that bypass logic will cost less and bring more, effectively reducing pipeline length by one stage.
> https://buildd.debian.org/status/architecture.php?a=sh4&suit...
Installer images are being built, too. But currently don’t boot due to an resolved bug in QEMU or the kernel:
> https://cdimage.debian.org/cdimage/ports/debian-installer/20...
If you’re interested in Linux on sh4, join the #debian-ports IRC channel on OFTC.
I’m the primary maintainer of the Debian sh4 port (and m68k, sparc64, x32, ia64, powerpc and ppc64).
Of several Busybox binaries for ARM, the v7m version is the smallest, and is (AFAIK) Thumb-only.
busybox-armv5l 2019-06-10 14:02 1.1M
busybox-armv7l 2019-06-10 14:02 1.1M
busybox-armv7m 2019-06-10 14:02 867K
busybox-armv7r 2019-06-10 14:02 1.1M
busybox-armv8l 2019-06-10 14:02 1.1M
busybox-sh2eb 2019-06-10 14:02 1.3M
busybox-sh4 2019-06-10 14:02 1.0M
https://busybox.net/downloads/binaries/1.31.0-defconfig-mult...Never heard of this and it makes no sense. The cache is pretty decoupled from the CPU pipeline. I saw a similar comment from someone discussion RISC-V mentioning that ARM/Thumb interworking is costly. I have no idea where these myths come from - really bizarre.
1) The compiler support for SuperH was beyond abysmal.
2) I loved that machine anyway.
I remember that you had to cross-compile from x86 to sh4 on windows and for that you had to first build the entire toolchain from scratch, my pentium 200MMX was not an ideal machine for that :D
I wrote SH-3/4 simulators that were significantly faster than the actual parts. (The cheat was that the simulator ran that fast on an early 200MHz Pentium while the SH-3 was something like 35Mhz.)
I also wrote a synthesizable[1] SH-5 hardware model that was cycle and signal accurate at every module boundary and ran >100k cycles/second on said Pentium. (SH-5 was a 64 bit successor to the SH-4 that also had a 32 bit mode that ran SH-4 code. I don't know whether it ever shipped.)
[1] The cache, TLB, and floating-point weren't synthesizable. Making them synthesizable would have killed the cycles/second.
[1] https://ed154c547559d2878d6a-e584b6b63c3a42919fe0cc5066a1430...
Aside from the Dreamcast, there's a few Japanese NAS devices that run the SH3 port too.
Debugging SH4 processors is tricky because they have a non-standard extension to JTAG standard: H-UDI.
Does anybody in this thread have details about the H-UDI proprietary SH4 JTAG extensions?
Context here:
It was a fun architecture: 16-bit instruction set; 32-bit bus, and 64-bit vector instruction set. I actually got Linux running on it, but ended up running the code without any OS. We switched to an off-the-shelf ARM-based SBC for the next version of the robot.
I was not the one who put Linux on it (it is not my name on the paper, but I was there around that time), but I had fun making a NetMeeting like demonstration application on it. Well, there was no sound but I could get 5fps video! :D
We have quite a collection of SuperH binary packages:
http://cdn.netbsd.org/pub/pkgsrc/packages/NetBSD/sh3el/9.0_2...
They currently seem to be in a "release tarbal" model of open source but know they ought to be in an develop on master branch in public repo model.