J2 open processor: an open source processor using the SuperH ISA
j-core.org
j-core.org
(This was me, in my bedroom at age 16, furiously preparing a robot capable of playing Connect-4 for an rtlToronto Lego robotics competition called Deep Yellow...)
>Numato: The cheapest usable FPGA development board ($50 US) the j2 build system currently targets is the Numato Mimas v2 (also available on amazon). It contains a Xlinux "Spartan 6" LX9 FPGA that can run a J2 at 50mhz, 64 megs of SDRAM, USB2 mini-B, and a micro-sd card slot.
PDS: Nice!
But, it would be an additional serious "would be nice"(!) -- if this could run on Lattice FPGA's / IceStorm Open Source Toolchain:
https://www.latticesemi.com/Products
https://news.ycombinator.com/item?id=12105913
https://news.ycombinator.com/item?id=20658584
Is the project still alive? There's no news for over 4 years.
I‘m Debian‘s sh4 maintainer and I have done quite a lot to keep the port alive and improve it.
Neat! What's your reason for doing this? Just a fan of retro architectures, or something else?
I‘m maintaining most unofficial architectures in Debian that are part of „Debian Ports“.
What does that entail?
What skillset does one need?
How do you connect with your users? The ones with the weird hardware?
I just received my Turtle Board with J2 and will give it a go how far I will get with that.
I wonder if this is the CPU core that's being used.
Since the J2 is quiet modular it might be possible to modify the timings but i somewhat doubtful that 2 vector units and two cores will fit on an affordable for hobbyist FPGA but i honestly don't know.
Edit: Changed DreamCast -> Saturn
Is there any chance you can expand on the issues there?
"We traded die space for speed so one instruction which takes one cycle on an SH takes 2 on J-Core."
If you are really interested i can see if i can come back to you.
https://www.copetti.org/writings/consoles/sega-saturn/
So, it's notoriously difficult to code for (if you want peak performance) and tough to emulate. I think each individual core is easy to emulate, but the whole salad of chips is another story.
Coding Secrets on YouTube has some videos on writing code for the Saturn, like this one. Not SH2 specific, but wow.
https://www.youtube.com/watch?v=czqwd43WQWM
https://www.youtube.com/watch?v=wU2UoxLtIm8 (mentions the incomplete documentation from Sega, as if the system wasn't complicated enough without iffy documentation as well)
* Two SH-2s. These CPUs occupied different places on the buses (plural) and were generally running different code to perform different tasks. It is incorrect to think of the two CPUs on the Saturn as being equivalent to x86 SMP; (or dual core, for that matter) the two CPUs were not interchangeable.
* An SH-1. This controlled the CD-ROM, and was only able to run code baked into ROM.
* The VDP1. (video display processor #1) It drew sprites and polygons.
* and VDP2, which drew backgrounds.
* A bus controller. The two SH-2s were hard pressed to deliver the computation needed for geometry transforms, so the bus controller contained a DSP which was often used to calculate geometry.
* A 68000. The Amiga, Macintosh, and Sega Genesis (and lots of others) used a 68000 as a CPU. This was generally given control of the...
* FM synthesizer. It was based on the YMF262, aka the OPL3, which drove the Sound Blaster 16.
I think that one of the SH-2s was tied more tightly to VDP1 and the bus controller DSP, and the other was tied more tightly to VDP2. I think I remember reading that a general common architecture was for game logic and VDP2 control to run on one SH-2, and the other SH-2 would feed geometry into the DSP which would feed into VDP1.
The VDP1 was a strange beast. Basically every hardware graphics accelerator you've ever heard of has used polygons as its basic primitive, but not the Saturn. The Saturn used quadrilaterals. It's not incorrect to think of the VDP1 as being a super duper sprite engine that was fast enough to build 3D stuff with sprites.
The VDP1 was not powerful enough to draw the whole screen. If a 3D Saturn game was to look good, significant parts of the screen would have to be drawn by the VDP2. It was mostly up to the task; it could display something like 7 layers of background, with two layers having full perspective control and 5 more with decreasing levels of scaling/translation. (SNES Mode 7 gave 2 layers, one with partial perspective control (I'm forgetting the technical term) and another with just translation) A good looking level/game required incorporating the VDP2's perspective background layer into the level. But since the VDP2 took different parameters than the VDP1 did, lots of games matched the VDP2's background very poorly with the VDP1's foreground.
That system was wild.
What's tragically hilarious about the Saturn is that despite the deep weirdness and absolute gobs of silicon, it's still not particularly good at anything.
...relative to its contemporary peers, I mean. The Saturn was an absolute beast at traditional 2D games, but even then, due to the paltry onboard RAM it required the 4MB RAM expansion cart to really shine and do things that the Playstation couldn't.
A good counterexample might be the original Amiga, which was "weird" and employed gobs of custom silicon and was indeed actually a generation or so ahead of its peers when it came to graphics and sound.
The Dreamcast uses an SH4, not SH2; which is a superset of the SH2 instruction set, let alone completely different timings.
Seems like it was in the pipeline, and still should be easier to add an MMU to a fully-fleshed system than create a whole new ISA. For that matter, there are plenty of RISC-V implementations without MMUs, aren't there? So they should be exactly the same.
The 64-bit thing I'll give you, although it isn't obvious to me how much that would matter for most places where people are using ARM or RISC-V so far; it limits you to 4G of memory, but in embedded uses (phones, routers, raspberry pi) who cares?
The reason why they build their own implementation of an CPU and not used some open Mips or Sparc is that they found that those used to much BUS bandwidth for their use case. SuperH was really well designed architecture with really good instruction set density.
Instruction set density is not a thing tagged on to a spec which many cores don't implement and if it is implemented you end up having a mess of mixed 32 and 16 bit instructions. SuperH is cleaner in that everything is a 16 bit instruction which helps with making decoder cheaper.
Additionally SH variants with full MMUs would have lost their patent protection after RISC-V was already released and gaining steam.
One example of such an head ache would be the branch delay slot. It works really well if you have the compiler perform an reorder of instructions for you so the chip can do more while waiting for memory.
In a modern pipeline CPU with register renaming, out of order execution this ISA level reordering would be at best complicating the implementation.
The branch delay slot makes sense where you aim to wait 1-3 cycles for memory loads but becomes a meaningless complication in the face of 10's or 100's of cycles waiting for memory seen on high clocked CPUs today.
You might say "J-Core failed" because of a lack of privilege and because of scope.
J-Core just wasn't pushed by an top university and had multiple big companies invest in it.
J-Core was never ment to be an industry standard covering use cases from embedded to server to super computer accelerators with an ISA designed by a comittee (i don't use designed by compute sneeringly here). J-Core was meant as a patent troll proof way for a small company to design an competitive chip for their use case on a "shoe string budget" (as far as CPU design budgets are concerned). The only reason why the wider world here about it, is because people in the company believe in open source and they didn't want anybody to have to redo this work if it fits someone's use case.
RISC-V was originally created because some greedy company wanted much money for their CPU architecture to be used in the course of an professor at an prominent university.
A CPU architecture is a set of trade offs. RISC-V traded compactness of the instruction set for the flexibility an 32 bit instruction set allowed and to gain back some density by having an optional compact instruction set at the cost of making the encoder slightly more complex. J-Core always has an Multiplier-Accumulator, while this optional for RISC-V. RISC-V traded suitableness for embedded use cases for agreeableness (You don't have to implement some stuff, if you don't need it), flexibility, mass appeal and i think those trade-offs worked out for what they had in mind.
If the team behind J-Core took that route they wouldn't have shipped a product and had a half finished obscure standards document instead.
And then there' the design for extensibility and only implement what you need that makes RISC-V appealing in the embedded space even apart from being free to use.
And the 32X used the SH2!
It seems to be a great candidate for IoT CPUs.
For example ARM has different ISAs depending on the application: A for more complex Application processors, R for Realtime and M for Microcontroller. The ISA for M is less complex, less complex to implement and likely to use less power than A (or than a full x86 implementation).
Not an expert but from what I've seen Super-H seems like a good candidate for a low power applications.
Intel did try... https://en.wikipedia.org/wiki/Intel_Quark
If you were to talk with a specific foundry and shell out the money to get access to proprietary simulators and other proprietary information you could get some simulated number. I am not aware of any such numbers.
I heard "decade on a button cell at chip inside free toy inside kids magazine cost" in a presentation once and that probably refers to a 180 nm process but this is hardly a concrete number.