Hotpatching a C Function on x86
nullprogram.com
nullprogram.com
[1] https://gcc.gnu.org/onlinedocs/gcc-4.7.2/gcc/Function-Attrib... [2] https://gcc.gnu.org/onlinedocs/gcc/Extended-Asm.html#GotoLab...
In our Deviare Hooking Engine/Deviare In Process [gihub-1] and RemoteBridge [github-2], we have a disassembler in place to hook Win32/COM/C++ vtables, so the hooking process is smarter if there are changes in the prologue. You can obviously take a look at the source code and learn a lot from it since it is state of art and perfectly competing with Microsoft Detours [3]. For an old comparison with Microsoft Detours you can check [4].
For anyone else looking for an extremely easy to use and higher level API, Deviare Hooking Engine makes extremely easy to hook and handle functions parameters and return value. Simple like this:
[snippet]
notepadPID = LaunchNotepadAndGetPid();
//in first place, hook DllGetClassObject of the target dll/ocx
hookDllGetClassObj = spyMgr.CreateHook("shell32.dll!DllGetClassObject", (int)eNktHookFlags.flgOnlyPostCall);
hookDllGetClassObj.Attach(notepadPID, true);
hookDllGetClassObj.Hook(true);
hookDllGetClassObj.OnFunctionCalled += OnDllGetClassObjectCalled;
[/snippet]Docs are available here: http://www.nektra.com/products/deviare-api-hook-windows/doc-...
[github-1] https://github.com/nektra/Deviare2
[github-1 bis] https://github.com/nektra/Deviare-InProc
[github-2] https://github.com/nektra/RemoteBridge
[3] http://research.microsoft.com/en-us/projects/detours/
[4] https://www.reddit.com/r/programming/comments/22crn0/gpl_alt...
If I want to use it with the GPL license, that only means my own code using your library needs to have a GPL compatible license, right? I can use your library with my GPL compatible code to modify other closed-source or GPL-incompatible code, right? Just checking because some may view it as using a GPL plugin with GPL-incompatible program, which violates the GPL.
I believe it's where the "ms" in "ms_hook_prologue" came from.
According to the present article GCC emits an eight byte prologue (LEA RSP,[RSP+0x0]) at the start of the function, but Raymond Chen says that Microsoft's compiler emits five NOPS before the function start address and an overwritable two byte prologue (MOV EDI, EDI) at the start of the function itself. To me, Microsoft's approach seems more efficient - but I've never written any serious x86 assembly. Anybody knowledgeable want to comment on this?
[0] the only reference I could find at the moment, but I've seen it documented elsewhere too: http://homepage.ntlworld.com/jonathan.deboynepollard/FGA/fun...
To change the configuration, the program would write new values to the struct, and then patch the executable. Since I knew how to find the offset of the struct instance in the executable based on the runtime address of the instance, this was easy.
The big advantage to this was speed - being on floppy disk systems, having to do a file lookup/read was very slow.
Sadly, what killed this technique was when antivirus software appeared and decided that an executable patching itself was malware.
Since linking on a Mac Plus took upwards of 11 minutes for our project, I'd would typically replace any particular line with a JMP $PlayMem to patch up whatever registers or memory wasn't right, and then jump back into the middle of the routine.
Oh, and then there was Steve Jasik's "The Debugger"[2] which would let you crash and then switch to a (capabilities limited) MPW environment to let you edit and re-link individual routines which it would then hot-patch back in place. I think for that one, he actually shipped own linker, which was a hand patched copy of Apple's.
Life in 24 bits of flat address space was more fun. Shorter, but more fun.
[1] http://www.mactech.com/articles/mactech/Vol.01/01.10/TMONDeb...
It's neat, but I'm wondering why simpler things like updating a function pointer wouldn't be sufficient.
What I'd like to see is hotpatching the function for another process, but I guess that's very hard to do with ASLR. Probably doable with some tricks, though.
Edit: Come to think of it, gdb is able to attach to a running process w/o debug symbols and find function addresses. So in other words, I just need to dig into and grok the gdb source.
But I doubt you could disable ASLR of a running process, for somewhat obvious reasons...
ASLR shouldn't be a big problem though. Only the base address of code in each file (executable or library) is changed, and you can easily find it. Functions within one file are not shuffled around. ASLR only exists to stop you from hard-coding function addresses, to make exploits harder.
Let's say you're Red Hat, and want to release a super critical security fix for some package, but you absolutely cannot restart the affected process. This technique allows you to write some code that redirects functions to fixed versions, patching the process while it runs.
If you're Red Hat and you do have the source, but the service cannot be stopped, and yet it must be upgraded, yet you didn't plan this into the software (and yet you didn't write in it in Lisp which could save your ass here) then you're plain stupid. If you're stupid, you're not going to pull off this highly technical, intricate hack, that is looming with pitfalls.
An example of that is garbage collection safepoints, although in that case you also need to be able to patch a function in the middle (any potentially unbounded control flow loop)
It's by far the simplest hooking framework on the planet.
on x86_64, a pointer-wide value is 8 bytes, which is enough space to put some instructions that detour to another function. so it's a question of whether you CAS the global function pointer, or the first 8 bytes of the function.
the "atomic CAS" part means you don't benefit from caching either, since you need to make sure the write is globally visible or threads on different cores will do different things. this kind of trickery combines two terrible things: reasoning about virtual method invocation (essentially) and lock-free programming techniques. have fun!
My guess is the OP's method can probably be extended to allow for more advanced hooking, like returning to the original function definition after running the new code or something. And, it means the caller doesn't have to know that the function is hot patchable, which I'm guessing is the most likely reason.
Ages ago, I tried to figure out how to dynamically probe OpenGL extensions. My intent was that the first time a method call was attempted, the method proxy would
look for a concrete implementation
on success, swap its pointer to the concrete implementation
on fail (not found), swap its pointer for a generic error method
then call itself again
By doing it dynamically, an app would only be probing extensions it actually used, versus every extension from every vendor. Back then, all the stubs were code generated. To avoid probing everything, we'd manually modify the headers, which I didn't like maintaining.I didn't get very far. Looking at your implementation reminded me of that effort.
I haven't done OpenGL in probably 15 years. I don't even know if extension probing is still a thing (useful).
Thanks for sharing.
Another question, as I understand it, one must go to c11 before C gains any "standard" thread awareness -- does that mean that the use of a naked int for x in this article, probably should've been a _Atomic(int) x;? [ed: if the article conformed to c11 as opposed to c99, that is]
I suppose it depends where x ends up being stored, if multiple threads across multiple cores will always see/modify the same version of x?
But that the write is atomic does not mean that it is synchronized with other cores. For example, thread A in core 1 could write the value of x, but later thread B in core 2 could read an outdated from its L1 cache. You should use memory barriers that force cache refreshing, so all the threads always see the same version of the variable and do not rely on possibly outdated cachés.
About the unlocked stdio, I don't think there would be a clear advantage. For best performance, I/O is done in blocks as big as possible, so locking mechanisms should not matter that much (you spend much more time in the actual I/O than in the locking). It should only affect significantly when doing lots of I/O calls, but even in that case, removing the locks would not improve performance as much as grouping and batching those calls.
https://lwn.net/Articles/620640/
Other unix kernels do the same for much the same reasons.