My, what strange NOPs you have!
blogs.msdn.com
blogs.msdn.com
I actually have a mostly-written LD_PRELOAD library lying around that exploits this for purposes more like the ones you describe though -- hot-patching code in memory at program load to dispatch system calls via little dynamically-generated trampolines so you can insert calls to arbitrary tracing functions. Perhaps I'll polish it up a bit and toss it on github...
Is it that the patch could be conditional and not incur as much of a performance penalty if the condition isn't met and the patch doesn't run? Would this have made a significant difference to performance?
I like it.
All the problems that Windows developers had to face for compatibility almost make me sympathize with them. Some context is given in Joel Spolsky's "How Microsoft Lost the API War" (http://www.joelonsoftware.com/articles/APIWar.html), where he talks about "The Raymond Chen Camp".
See http://www.piclist.com/techref/piclist/codegen/delay.htm
> Late in the product cycle (after Final Beta), upper management reversed their earlier decision and decide not to support the B1 chip after all.
After all that work, on compilers, processes, tools, etc, it was all scrapped anyway. I wonder if there is a lesson to be learned there. (The usual startup-mantras don't really seem to apply!)
I don't know what the correct approach should have been, but it seems that changes to support other architectures should have been made independently from other fixes and tagged such that they could have been backed out rather than contaminating your source and pushing it to gold. (Remember this is before the days of Windows Update!)
Where's the mistake? Wouldn't backing out the workarounds introduce new potential for error due to other changes made after the workarounds were put in? Why do the workarounds need to be backed out? They don't break anything. If it ain't broke...
My point is that it wasn't really a mistake; it was a change in circumstance.
Does anyone have any examples or explanations for how these CPU bugs are caused? I probably have enough knowledge of CPU architecture from school to follow an explanation, but not enough to make my own guess.
Some CPU verification code I wrote on an internship a couple years ago discovered a few bugs in a certain fairly widely-used processor, though I'm pretty sure they were all logic-level problems (i.e. RTL bugs, not circuit level ones)...
- The L1 D-cache tracked clean/dirty status at half-cache-line granularity, and if you did a store (with just the right timing) to one half of a cache line you had just explicitly cleaned with a cache-clean instruction, the dirty bit wouldn't get set on that half line, so as soon as the cache got flushed the data written by the store was lost.
- The prefetcher would shut down sometimes as a power saving technique, but if you laid out the right sequence of cache operations and branches in the last 32 bytes of a 4KB page, sometimes concurrent TLB misses would cause it to not get re-enabled, meaning the processor would lock up, stop fetching instructions and just sit there dead in the water until an interrupt came in (assuming interrupts were enabled).
There were a couple more, but I thought those were the more interesting ones. Granted, these weren't bugs that were likely to be encountered in normal usage for various reasons (in addition to being extremely difficult to reproduce sometimes -- i.e. on one in particular you could run the exact same sequence of instructions from system power-on and sometimes it happened, sometimes it didn't), but bugs nonetheless.
If the detection logic is wrong, you could easily end up forwarding 16 good bits + 16 bits of random garbage into one input of a 32-bit operation. That would explain Raymond’s "if all the stars line up exactly right" line, since the hole in the forwarding logic must have been really small (or it would have been caught in testing).
It's amazing how much hardware out there is buggy and fixed by operating systems and drivers.
However, since the original comment was about the Z80 which according to http://www.z80.info/decoding.htm treats invalid instructions as a NOP, the point is moot.
By the way, z80.info is the type of site that made me fall in love with the web -- it's full of information compiled by people doing it not for AdSense impressions, but just for the love of sharing knowledge. I miss the days when all of my Google searches would take me to these sites instead of the contentless content farms of today.
One of the commenters said that he would buy a book of this story and more and someone replied that there is in fact a book from him already:
"Old New Thing, The: Practical Development Throughout the Evolution of Windows" by Raymond Chen
Very strange for a frequency-encoded instruction set to let such short opcode sequences go to waste.
So by having a pattern to your instructions you can optimize the layout of the CPU. XCHG AX,AX does exactly that, with no visible effect, but the action is carried out.
That was then.
These days CPUs compile machine language into microcode, so you can pick any number for the opcode. The microcode is still switches based on the bits, but you don't see that.
Similarly, on x64 0x90 is a true no-op, while stuff like XCHG EBX, EBX clears the upper 32 bits of the register.