x86 API Hooking Demystified (2012)
jbremer.org
jbremer.org
Essentially given an idempotent data-transformation process, you can inject a dll into it and hook its filesystem calls, maintaining a list of all files it touches. Upon process termination, serialize out the filesystem state (size, timestamp, MD5) of each of those files, saving the data in a dependency file whose filename is the hash of the command line.
When you run the same command line again, the dll can first look up the dependency file from the previous run, and if all files on the filesystem are in the same state as previously, you can short-circuit the process execution entirely.
This is orders of magnitude faster in many cases for highly-parallel data build jobs where only some small percentage of the source data changes each time you run the build job, and has the advantage that you don't need manually maintain a list of dependencies for each process type (no new dependency can be added without changing one of the existing dependencies).
You can implement and mount such virtual FS entirely in userspace using FUSE.
There's also:
push rax
mov rax, 0xDEADBEEFDEADBEEF
xchg qword ptr ss:[rsp], rax
ret
but i dont like it as much since it touches stack (technically more detectable since you overwrite stuff at rsp - 8). RSP & RAX are original values after that gadget though.
Also fun fact. My library supports JIT-ing, so you can create the stub that the hook jmps to at runtime and it will JIT translation logic for the calling convention and pack the args + ret value into a structure that can be modified. So you can hook unknown functions at runtime.
I am guessing the length of the jump itself is probably important for some reason. But if you could afford to overwrite ~16 bytes, maybe you can store the address inside the imm64 of another instruction. Length issues aside, it shouldn’t break nested hooking at least.
works too
I guess also though, at ~16 bytes its probably deep enough into the function that it may no longer be position independent, or hell, maybe the function isn’t even that long to begin with.
The application would ask the operating system for the default audio device. By intercepting this request, we were able to re-route it to our own, virtual audio device. Our program would then fetch the audio data from the virtual device, and replay it to the "real" audio device. At the same time, the audio gets saved to ram, and finally to disk.
The benefit of this method was that we were able to actually isolate the audio from all other sources on the computer. So you could, in theory, mute the playback, while still being able to let the recording run. Ultimately, we abandoned this method, as it proved quite unreliable. But it was fun to come up with, and finally implement.
API hooking sometimes even works in hostile environments, like on software that tries to guard against patches and modification, simply because it can be challenging to detect, and you can do it early in (like inside a DLL entrypoint). So if you can do all of your work at API boundaries, you can get away with a whole lot, even if an app is packed with a strong VM packer.
Worth noting that for many less difficult use cases on Linux, you can use LD_PRELOAD to somewhat similar effect.
(I wonder how this interacts with glibc symbol versioning, now that I think about it.)
https://github.com/baldurk/renderdoc
I used graphics API hooking in a Source Engine game where I didn't have the source for the engine but needed to display the 2D Flash based GUI (Iggy) at the correct time. It was fun to get it working.
Valve's Steam also does something similar since it has the ability to superimpose it's GUI over a running game.
Similar on Linux, it doesn’t have IAT but it has PLT (Procedure Linkage Table) which is basically the same thing as IAT.
Point 3 might include logging, passing the request through to the original ISR (possibly with changed parameter values), changing the return values, anything really. Easy & fun. =)