Writing a self-modifying x86 factorial program
brianstadnicki.github.io
brianstadnicki.github.io
Does anybody know? I would be very interested in an example.
Your CPU has an instruction cache and a data cache, and on ARM (and x86, too, but I'm not sure) these caches are not coherent. So if you modify your instruction stream with a write, you have to clear the instruction cache to ensure that your modified instructions are actually executed by the processor. If you do this a lot, this will make things S-L-O-W, because it forces you to go all the way to main memory to find the next instruction to execute.
This means that if you do want to generate code at runtime, you want to batch the modifications into large groups, so that you have to invalidate the i-cache less frequently. This actually is useful -- it's what JIT compilation is! The reason that JIT can be helpful (even in statically typed languages like Java or Haskell) is that programs often get passed functions as arguments (eg, qsort in C). A static compiler can't optimise these functions much, because you have to know what the function argument will be to do much. But at runtime, you do know what the function is, and by inlining it your code can be made much faster.
The jit can help if the value of the pointer varies dynamically and in an unpredictable way (otherwise PGO would also help).
I was curious about your comment regarding arm and risc-v not having coherent instruction and data caches. Is this a toggle on these chips hen for turning it on and off? I think I remember reading about some SoC that have this configurable.
These queues where "visible" in their effect on self-modifying code. After modifying one of the instructions that could be already in the queue, you had to do a jump to flush it.
If you know about this, it is obvious how certain code only works when single-stepped through, or perhaps when run on an 8088 with its shorter queue. But few people did, even among experienced programmers.
IIRC the 486 and everything newer can detect when a cache line containing code is changed, so this is no longer necessary (but bad for performance as other commenters said).
ARM quite famously requires the explicit cache flush, and will usually fail to work without it. However, some emulators, e.g. QEMU, don’t require the cache flush, which can lead to confusion if you usually test on emulators.
There might be benefits to "modify-once" code in some circumstances because of how branch prediction works. If you determine that a branch is being consistently mispredicted you can speed it up by changing the code so that it is correctly predicted - invert the test, swap the branch targets, apply a prefix instruction, etc.
It is not unusual to have CPU-specific bits of code gated behind a CPUID check; you could maybe gain a bit by doing the check and then hardwiring a branch or pasting the code into the middle of a loop.
A related trick is the "micro-interpreter". You build a hyper-specialised language and run short programs in that - which can then use self-modification as well, because they're counted as "data" and not code. I think there was a game that used this (Flashback maybe?)
That depends; how broad is your definition of "self"? Within a process, this is a fairly common approach already (JIT).
It's for me, because you can definitely do one less loop by modifying:
cmp ebx, 0
to
cmp ebx, 1
Init: EBX = 2; EAX = 2;
Loop: EBX = 1;//gets decremented
CMP EBX, 1;//it's 1 so we exit
Final result is 2. Sounds the edge case is not that of an edge after all