Intel X86 Encoder Decoder
intelxed.github.io
intelxed.github.io
Aside: Instruction emulation is pretty finicky and bug prone. I'm not too familiar with Xen, but KVM has had at least 10 instruction emulation CVEs. There were talks at both KVM Forum and Xen Summit last summer mentioning the sketchiness of instruction emulation.
https://xenbits.xen.org/xsa/advisory-200.html
https://xenbits.xen.org/xsa/advisory-195.html
All XSA
Or, to put it another way: looks like a good library for adding "just a bit of JIT" to an interpreter, without going full LLVM.
Ex-Intel here. It takes years to gestate new instructions. First specs are in controlled documents available to Intel employees only. Eventually, when things are nailed down, preliminary specs are available under an NDA where a VP approves/signs the Intel side of the NDA. Tool vendors will get those specs. Finally, when the chip is announced the previously NDA documents become public. I was a CPU designer at multiple companies, and everyone follows a similar process.
Furthermore, if you are a big enough customer you can also negotiate for special isa extensions yourself (can't provide citation but heard from people at Intel), then of course you'd be the only one to have the documentation to take advantage of them.
The 286 had purposely undocumented instructions, of this sort: "Ooops this is b0rk3d. It will always be b0rk3d. Let's pretend it didn't happen." So for generations there were holes in the op code map that people tip-toed around. Especially since Intel (meaning the internal grey-beard collective) also forgot exactly what those opcodes were and what they were supposed to do. You care, why, exactly?
It's not like the NSA slips extra opcodes into executables that you compile with your own compiler in order to spy on you. They have much easier ways to spy on you.
Also, it's not like it is that hard to throw unused opcodes at the decoder and see which ones give you the illegal instruction exception, and which ones do something else. You now have a homework assignment. Have fun, let us know what you find.
http://web.archive.org/web/20000817193452/http://www.tbcnet....
http://web.archive.org/web/20000817084210/http://www.tbcnet....
Obviously, XSAVE/AVX was handled better.
I'd say some are eventually made public. I went to the IDF in San Francisco expecting this sort of information on Skylake. What a waste of a day.
Not that I'm complaining. Intel's information is excellent. It just arrives when it arrives and it is what it is at that point in time. There are many ways involving effort which glean more information: Agner Fog, articles, patents, .... BTW, Intel folks are helpful on their dev board.
@dbcurtis: Yes, that's what I meant. Thanks for clarifying.
Back in the 1990s, when I was doing assembly language on the Amiga, Motorola sent me the official 68000 manual for free when I called and asked them. It was a really cool book to have on the shelf and occasionally lead through.
I looked at abebooks.com, and there are Intel books, but I wouldn't know exactly which one would be the reference manual. Anyone got an ISBN?
http://styx.head-crash.de/stuff/intel_manuals.jpg
Now they still have hardcopy, although it's basically just paying for someone else to print and bind the PDFs for you:
Previously discussed here: https://news.ycombinator.com/item?id=10143295
Internet Archive collection (the Intel documents part): https://archive.org/details/bitsavers_intel
Library of Congress reference: https://www.loc.gov/item/lcwa00096459/
Can I assume that whatever is parsed by XED is going to be parsed in the same way by real CPU's?
That's one goal certainly. Not claiming it is perfect. I have been working on it a long time and many cycles have been run on it between Pin, Intel SDE and some internal simulators that boot OSes, etc.
Python would not be as suitable as the full implementation language, since C is much more easily accessed from other languages, which is a good property for a library like this.