Resurrecting the SuperH architecture
lwn.net
lwn.net
There have been some minor additions, he said: the J2 adds four new instructions. One for atomic operations, one to work around the barrel shifter, "which did not work the way the compiler wanted it to [...]
Is so intriguing! Does anyone know what was wrong with the original barrel shifter design? I tried reading up on it but failed to find much reference material. I followed the link to the J-core community site to read the code, but it wasn't immediately browsable, just available for download.
I assume there were compilers for SuperH back in the day, didn't they use the shifter? Why not fix the compiler to teach it the existing instruction, rather than adding an instruction just for this? How wrong can a shifter be, really? The questions just heap up.
On the original SH4 implementation, under certain conditions, there had to be one cycle in-between when a shift-amount was generated and when it was used, otherwise there would be a one-cycle CPU stall. A real right shift would avoid the need to schedule around this stall. This isn't necessarily something that needs an extra instruction to fix, the implementation could be designed to not need the stall, but it might difficult to work around. I don't to circuit design, but dynamic shift instructions typically look at as few bits in the shift amount to simplify and speed up the design of the shifter. The reason for the delay in the original SH4 is probably because it analysis and tags each register with information for the correct shift direction and amount, and certain units won't have this information ready for the shifter in time, hence the stall if the shift is too close the shift amount generation. (I've read this certain CPU implementations have done similar work in tagging if a register is zero or not, in order to help keep branch-on-zero/not-zero instructions quick.) If the instruction talked about is a dedicated right shift, it could be defined in a way that doesn't need a negation and extra tagging, would be much more compiler friendly, and faster.
Does that mean that the shifter is actually capable of doing rotates? Otherwise the negation part doesn't make any sense.
If you have 0xf0 and want to shift it three bits to the right to get 0x1e, no amount of negated-amount left-shifting is going to do that unless the instruction is a rotate.
If, on the other hand, you can do a 8-bit rotate left of 8-3 = 5 bits, that would produce the same result and need that "negation" (which is actually an inversion).
GCC has supported the SH series since at least the 2.7 era (though Hitachi's compiler seemed to produce better code in those days, but only ran under DOS).
Alternatively, I proposed the security enhancements for processors showing up in academia be applied to this or another processor with low market share as a differentiator. Some of these enhancements take almost no chip real-estate, esp simple tags & tag-checks. The chip designer could also make money for the semi-custom work. Time has passed, that didn't happen for low market chips, and did happen for AMD+Intel for non-security applications [that I know of]. Matter of fact, even though my scheme didn't happen, AMD is making so much money off the other half of my proposal that they could be cited when trying to convince chip-makers to do it for security enhancements with mass-market availability. So long as they don't bear the cost of failure (huge in ASIC's) they might go for it.
Finally, anyone wanting to deploy this or other things, remember the Structured ASIC's with FPGA conversions. eASIC [3] has a long track record in this with offers down to 28nm. Gigoptix [4] does S-ASIC's down to 28nm. Tekmos [5] offers a similar product at 350nm (good for budget masks). Just make sure you design in FPGA's with ASIC transition in mind from the start & follow published advice on that (available with Google or consultation). The result is you prove it in FPGA's, even use it on FGPA boards, and then move it to S-ASIC later for reduced costs/power + maybe speed increase. Authors are right that 180nm is a sweet spot although proving it at 350nm or higher first might be smarter given costs.
[1] https://www.schneier.com/blog/archives/2013/09/surreptitious...
[3] http://www.easic.com/products/28-nm-easic-nextreme-3/
[4] http://www.gigoptix.com/products/asics/asic-type/structured-...
[5] http://www.tekmos.com/products/asics/process-technologies
They might have had more actual work done with them back in the old days, so their code might be in better shape now.