Breaking the x86 Instruction Set: how to find undocumented x86 instructions
youtube.com
youtube.com
0f0d/0-7 were all prefetch instructions, but probably behave like NOPs if not supported
0f18/0-7 are HINT_NOPs
0f{1a-1f} are also HINT_NOPs
0fae is a bunch of assorted instructions (FXSAVE, FXRSTOR, LDMXCSR, etc.)
dbe0 is FNENI
dbe1 is FNDISI
df{c0-c7} are x87 ops
f1 is ICEBP
c0,c1,d0,d1,d2,d3 groups have a few aliases (SAL/SHL)
f6/1 and f/1 are aliases of f6/0 and f7/0
0f0f are 3DNow instructions and it wouldn't surprise me if there were many aliases there
0fa7 was briefly used for the IBTS instruction on the very earliest 386s and then CMPXCHG for the very earliest 486s
(http://datasheets.chipdb.org/Intel/x86/486/Intel486.htm)
perhaps VIA continued to use it for a CMPXCHG alias
IMHO the 1-byte opcode map has basically been completely explored and documented, perhaps with the exception of some of the x87 stuff. It's the 2-byte (0F xx) ones where things start to get really interesting.Interesting. 0fa7xx is indeed where the VIA Padlock instructions live, but the last byte is probably being partially decoded.
I don't understand why more vendors don't do this (if anyone wants to comment to this, I would be interested to get another opinion.) While my experience is admittedly limited to obscure chips that require NDAs for access to the specs, I was always a little annoyed that almost every time there was not even a reference to the errata documentation that the vendor provided when a new version of the spec would come out.
Now in AMD's case, I would argue that they should be more clear that it's an errata, and mention that the updated spec differed from previous versions (which it alluded to by saying "This behavior is model-dependent".) Ultimately, the spec is THE document on how a user should expect the chip to behave. So sue me if I am blurring the lines between an errata and a mistake in the spec, but I just want my documentation to tell me what the chip does without having to refer to a dozen other secondary documents dang it!
https://www.symantec.com/connect/blogs/x86-fetch-decode-anom...
github: https://github.com/xoreaxeaxeax/sandsifter white paper: https://github.com/xoreaxeaxeax/sandsifter/blob/master/refer...
This kind of technique, and the exploitation of minor CPU errata, can be used to help differentiate processor models and steppings.
That in turn allows a currently widespread DRM system to download personalised portions of object code that rely on properties specific to the licensed hardware in order to execute properly, in an attempt to counter debugging, emulation and transfer - continuing a tradition practised in copy protection techniques since at least the 6502, maybe even earlier.
I don't know much about CPU internals, but would it actually be possible to 'patch' that through updated microcode?
It's such a clever program, will be intrigued to see what else it can find!
If this just searches the space looking for packets that shouldn't decode but end up getting executed, then it's unlikely to be anywhere as interesting as f00f.
In all likelihood we have already seen 2017's big silicon bug and it was AMD's Ryzen 7 1800X issue.
Nice job though
I mean if you worked for Intel and your manager said "make me a really secret instruction" would your best response be "lets just not document it and hopes noone notices"?
What I would give to read the full microcode of the latest Intel processor. I am guessing it is stored in a vault with the real nuclear codes, Alien cadavers and the Holy Grail.
You just need a set of magic register values, like how CPUID [0] instruction already works.
[0]: https://github.com/RUB-SysSec/Microcode/tree/master/updates
[1]: http://syssec.rub.de/research/publications/microcode-reversi...
https://hackaday.com/2017/04/25/an-analog-charge-pump-fabric...
I suspect that really keeping it out of view would also cost both silicon and propagation delays in what would probably be some of the most critical paths, but then I'm not a vlsi engineer, or whatever the correct title would be :)
Given this, I suspect wiring in a path all the way from the relevant versions of the relevant registers might be quite expensive. Plus part of the decode logic now needs to block on a register value - so a timing based attack might find these.
More qualified comments welcome...