How the Z80's registers are implemented
righto.com
righto.com
http://en.wikipedia.org/wiki/Barrel_shifter
The easiest way to shift bits isn't "adder like" with carry propagation across data lines (sorta), although that could be done. The easiest way is just to phase distort the bus (sorta) so what you called bit X is now bit X-4 for all bits on the bus, and use a 2x1 mux to select which, and cascade rollers in binary (you don't need a roller for all integer numbers if you can cascade them 16, 8, 2, 1 position only activating some rolls). If you can rotate, you can shift by adding one more stage at the end that does bit oriented ops to clear MSBs or LSBs. Also if your roller is capable of rolling all bits on the bus, you don't need two rollers one in each direction. Latency does add up of course, all those cascaded ops.
Also I saw in the notes the Rodney Zaks z80 book being referenced. I learned assembly from that book, back when it was new. I enjoyed that book. I also looked at the PDF scan of someone's beat up copy... remember when textbooks were only $10.95 each?
The z80 was quite the chip for its time and I enjoy reading these articles has they look back into the deeper inner workings of the chip.
One could master the z80, I'm not sure many modern processors could be mastered at that level by a single programmer. They seem to have gotten so complicated that they're beyond comprehension in a single person's mind. I could be wrong but that's the general feeling I get. Sure you can be really good at them and maybe an expert in many applications of a modern complex processor but I doubt many people have a singular understanding of everything each chip has to offer.
Or maybe I'm just getting old.. :-)
On the non-x86 side, modern ARM SoCs like the ones used in smartphones are not all that much simpler despite a less complex CPU core - they're still thousands of pages of documentation in total.
The nasty operations are the ones that are linear (or worse!).
Modern CPUs are NUMA. Don't treat memory as RAM any more, because that's not true.
Biggest problem with this is that not all CPUs have the same amount of cache. But you can get around this by treating the cache as the low area of RAM, with instructions to get the amount of cache available. Especially if cache is also paged.
Other issue with this is context switches, but this is conceptually no different than paging RAM to disk when required.
I haven't seen many people actually use these ops, because it's actually pretty hard to do better than the built-in cache allocation policies for most applications, especially if you take into account that your app is going to get swapped out consistently by the operating system task switches.
> I haven't seen many people actually use these ops, because it's actually pretty hard to do better than the built-in cache allocation policies for most applications, especially if you take into account that your app is going to get swapped out consistently by the operating system task switches.
And again, this is largely because the cache is implicit to the OS. There's no way to go "this is the stuff that was cached last time this process has control, when you can, reload it back in" to the processor, because you can't tell what in cache is "owned" by what in anything like an efficient manner - and even if you could, the moment you start executing a context switch you've overwritten random bits of cache.
It's like if the processor was set up to directly talk to the hard drive to do paging on demand, to the point that the OS wasn't even aware of it. In theory it's a good idea, but the more you look at it the more flaws emerge.
Conversions to C64 from Spectrum were routinely done by hand, routine by routine, as the C64 could "emulate" the Speccy this way. For instance games like The Great Escape ( http://www.crashonline.org.uk/35/greatescape.htm ) was converted this way, as were other games (it was a bit disappointing when this happened, as it basically meant not using the C64 graphic sprites and its tricks.
Chances are, someone is already working on an implementation.
I did it the same way when I designed a Z80-compatible CPU in a (graphical) logic simulator, although I used a dual-ported register file.