Movfuscator: A single-instruction C compiler
github.com
github.com
The loop jump address sometimes changed for some effectful operations like printing or for optimizations like executing addition.
It took me about four hours to find the correct password. In the course of there three hours I wrote 1) an executor that used i386 debug registers to look for current MOV addresses, 2) a tracer that produced a trace and 3) a compactor which identified common instruction sequences and presented them as some macrocommand. It turned out the original source code has used macros in the opposite way. The final challenge was to write brute force password finder, which is not that hard at all (for 32-bit checksum).
All in x86 assembler. I guess it was about 1995-96, somewhere there.
Now I'd use the same technique, but on higher level. Instead of peephole compacting I'd use graph analysis, but that's about it. You can get pretty much everything from the program trace, I think this way you can get even more information than from disassembly.
So in my opinion, it is one hell of a cool experiment. But try not to use it as a real obfuscation device.
That one taught me it was far better to start by working backward from the result, although self-modifying code tends to be more difficult that way.
Do you have some illustrative example?
Let me hypothetically apply movfuscator to some not too complex program and look at the assembly. I believe I'll get nothing useful from it, which is fair. But if I trace the program, I can get some useful results - like filling some array with data (increase of the addresses accessed), looping over indices (accesses within some range), see through generated code (table lookups here and there are identical), etc.
Getting the trace is the tricky bit, I had to write an msp430 emulator.
(Actually seeing this example requires completing the rest of the levels, but you should do that anyway, especially if this is the sort of thing you're interested in)
I've always felt these were more of a trick in being single instruction set because you are using some of the addressing bits to encode an opcode.
That's somewhat different from the move-based code discussed here where the MOVs are actually performing the computation.