You can't copy code with memcpy
devblogs.microsoft.com
devblogs.microsoft.com
Very fast & easy.
Of course, it broke when COMDATs were introduced.
I did a similar thing with my text editor. The colors were configurable. The usual way was to have a configuration file, which the editor would read upon startup. But floppy disk systems were unbearably slow. So what I did was take the address of the configuration data in the data segment. I'd work backwards to where those bytes were in the EXE file, and patch the EXE file. This worked great!
Until the advent of virus scanners, which broke that. Virus scanners hated self-modifying EXE files.
You can name it: 'Bright Moments'
I'd be happy to place an advanced order if it leans you in that direction!
For a modern take on similar idea see also this discussion from a while ago:
"Improving startup time in Atom with V8 snapshots"
I was amazed that I never ran across another DOS program that used the technique.
Do you unfork then unexec or unexec then unfork?
Turns out there was an undocumented Windows API function along the lines of "AliasCsToDsRegister" or something like that - I've tried to find a reference to it but I can't find it. It allowed me write into the code segment (the CS was global and read only - as it was shared among all processes) and replace the first few bytes of the connect function call with a jump to my code which would then put it back, make the call to the socks server, do some other magic, put my jump hook back in and the return to the caller. Good times!
Kind of surprised I remember this and more so that it actually worked.
Some other functions had similarly great names:
>>Anyway, I am not sure what type of person writes a function called "BozosLiveHere" and puts it into USER.EXE
>This started out life with a non-bozo name as an undocumented function in Windows 3.0.
>Windows 3.1 removed the undocumented function, but we found that some programs were using the undocumented function and started crashing.
>So we reluctantly put the function back, but changed its name to "BozosLiveHere" so that nobody else would use it in the future.
>A similar story exists for "TabTheTextOutForWimps".
KnowledgeBase Archive
An Archive of Early Microsoft KnowledgeBase Articles
Q89560: Creating Dynamic Code Segments Using PrestoChangoSelector
An example function I memcpy and run:
https://github.com/skimmilk/liblibinject/blob/master/src/lib...
https://web.archive.org/web/20091104065428/http://www.di.uni...
I am stealing this line.
I wonder how it looks internally. Like, can anyone write there. Apart from a university I worked at management would never let me write any "blogs" about work related stuff.
I deliberately obfuscated the script by replacing characters with various Unicode confusables, both the letters and the symbols.
I still felt bad and ended up deleting the post.
Hopefully nobody tried to type it in manually instead of cut & pasting just to see what would happen...
Delete locks and the like will stop this… unless you bulk delete them first.
The admins will get warning emails that they will finish reading in abject terror some time after their entire cloud tenant has gone to heaven.
The modern equivalent of “Operating system not found” is “Click the guides below to get started with your new cloud account.”
# this would wipe all EC2 instances, assuming you are authenticated and there is nothing else need to get a list of them
Get-EC2Instance | Remove-EC2Instance -Force
If there is a way to enumerate objects without explicitly specifying some properties then PS makes it very easy to pipe that output to remove or disable cmdlets.
I never worked with AWS but if I needed to write such killer script I would kill instances (as in above example), remove all S3 objects (quite similar, Get-S3Object | Remove-S3Object) and finally wipe out all IAM roles and users (similar Get/Remove-IAMGroup, ..IAMUser etc). And thesubscription/account things, as @jiggawatts says, I'm just not familiar with the AWS lingo. And secrets/creds, of course.
By the time a human starts to investigate why the monitoring gone mad (assuming it wasn't hosted with all other infra on same account, lol) there would be nothing to do, too many thing are gone already, too many things are on the way to be purged and even if they could be restored there would be too many things to tie it again to something resembling a functioning system.
As an aside, whenever I set up a Windows PC for me or a family member, the first thing I do is uninstall any third-party antivirus that may have come with the computer. I have found that anti-virus software likely makes my computer more insecure by having a big attack surface, not to mention slowing it down.
It's almost like the extensibility necessary to make third-party security products work requires creating entirely new attack surface for those products to work.
Just wipe and reinstall Windows.
For those who remember, its early versions (before the MS acquisition) required the Visual Basic runtime.
I'm a pentester and red teamer, and yes, bypassing Defender isn't all that hard, but neither are most 3rd party AVs, and Defender does not bring in all the instability and additional attack surface. Also I think it can take advantage of newer kernel APIs that non-MS AV can't.
> The customer is a major anti-virus software vendor! The customer has important functionality in their product that that they have built based on this technique of remote code injection, and they cannot afford to give it up at this point.
Oh now it makes sense.
1) VirtualAllocEx some memory to hold the path to a DLL you want to inject.
2) WriteProcessMemory to write the path to that memory.
3) CreateRemoteThread passing in the address of `LoadLibrary` and the pointer to the previously allocated memory.
It's no coincidence that `LoadLibrary` has a signature which can directly be passed to `CreateRemoteThread`!
However it should be possible to use GetModuleHandleEx to find the DLL base addresss in a remote process, then ReadProcessMemory to implement your own remote GetProcAddress.
Yes, the loader will create file-backed memory mappings and not redundantly store read-only parts. However, it is free to load it at a different address in each process. This can happen via ASLR, or if the mapping is already claimed by the time the module loads.
They may get the same base address repeatedly in multiple processes and work most of the time, but it's not guaranteed.
> That's not true. You're not considering different virtual addresses backed by the same pages.
technically I suppose, but PEs don't tend to be relocatable, so if it mapped it in at different virtual addresses that would be extremely unlikely to be backed by the same pages as much of the just-mapped-in code would need relocs
- injected DLL can change some process-related settings and break main program
- injected DLL can call non-reenterable function from a system library simultaneously with main program
(And as far as the terminology, these aren't "bugs" since they aren't defects in the software.)
Linux can do this too with ptrace. And mac has something similar I'm sure, you gotta be able to debug.
Now, the target process can implement countermeasures against you, that's what anti-debug is, but it's impossible in the general case to defend yourself against a debugger with the same privilege level as you. (it is an arms race though, so sometimes it's easiest to resort to kernel-mode anti-anti-debug techniques even against pure user-mode anti-debug. If you do that you need to disable KPP, and you can't disable KPP in the supported way because the supported way is to attach a kernel debugger to the kernel, which will make it obvious that someone is debugging!)
GDB will use this if you tell it something like `p foo("bar")`, as it needs to allocate memory for that string somewhere.
unsigned short main[] = {0104525, 0xb8e5, 32, 0, 0xC35D};
and I made a tool to generate it automatically from a function [0].
It uses exactly this technique to run a thread in cmd's process to actually change the directory. It's kept working from XP on up to Windows 11 now. I am always amazed it works, I fully expect it to go boom some day, probably with an error along the lines of "Don't do that, please".
[1] https://github.com/seligman/ccd/blob/master/RemoteThread.cpp...
Forgive me if I got it wrong, I am definitely not a Linux kernel committer.
This harkens back to the days when you could "download" a math coprocessor for your SX system, which was a TSR which likely did the same catching and handling of illegal instructions.
[0] https://github.com/rkeene/sse3-emu/blob/master/libsse3.c
For example:
- On most RISCy Arm CPUs with Harvard style split instruction and data caches special architecture specific actions would need to be taken to ensure that after the memcpy any code still lingering in the data cache was cleaned/pushed out to the intended destination memory immediately (instead of at the next cache line eviction).
- Any stale code that happened to be cached from the destination (either by design or coincidence) needs to be invalidated in the instruction cache.
- Depending on the CPU micro architecture, programmer unknown speculative prefetching into caches as a result of the previous two actions may also need attention.
If using paging you may need to invalidate the TLB entry which contains execute permission for the page.
On x86 if using segments, after changing segment attributes you need to reload the segment selectors.
The execution pipeline may need to be flushed, using a serialising instruction.
When modifying code in place that may be being executed by another thread on another core at the same time, some modifications may trigger CPU errata.
On particular CPUs there may be other kinds of caches or state invalidation required, but hopefully the OS provides a "flush I-cache" function that covers all of them.
"The customer reported that this code “worked just fine on 32-bit x86 and 64-bit x86”, but it doesn’t work on Itanium."
"Actually, I’m surprised that it worked even on x86!"
And that is probably how Spectre and Meltdown were found!
Differences in the cpu's.
When writing code like this, as Chen says, you are bound to the architectural rules regarding how to appropriately locate code and safely invalidate code caches etc.
So if you follow the rules it doesn't work then typically I figure you'd take it up with the CPU vendor.
While you could potentially do it and have a good reason to in userspace code, it should be heavily scrutinized because it's so unconventional.
I figured it was likely that code that aggressively scans and modifies other running executables would be written as a kludge, an unorthodox way of abusing the compiler-loader-runtime chain.
Also, the client should have embedded their code in the executable file name so they just have to jump to the appropriate offset in argv[0]. This way, future updates just require renaming the file!
Dynamically loading code is indistinguishable from self-modifying code, and each architecture has special steps you must take in order for it to work.
No problem, I just fixed the compiler to compile it!
Then you also get namespaces, modules, compile time execution, type safe macros (aka templates).
It is a bit more than syntax sugar.
For example the following code you know what the assembly is going to be.
strcmp(char* a, char* b);
strcmp(str1,str2);
If you do the above as a template you can run into some weird issues that you may not be expecting. So while tedious, you would need to make your own wscmp. You also have to be very careful so that you don't pull in ANY libraries. Since your code needs to be 100 % independent and do the loading itself.
C++ exceptions are implemented at the OS level in windows. C++ exceptions using SEH, while there's also VEH and unhandled exceptions. You can easily use SEH for your shell code, it's just not documented well. But sadly you have to manually set this up by having something like
SetExceptionHandler(curAddr,Handler) // Where curaddr can be found by doing something like call $+5 so you remain position independent.
When I was first shown this I was like 'What non virus use case does this have!?!?'
I used it to copy out a (forgotten) password from a password inputfield in another program, which you cannot read remotely (for security reasons). Worked fine for that one use-case, and I haven't used this trick it anywhere else ever again :)
Suppose you are able to inject and execute remote code into the process containing your password in a text field
How did you read the password out of it? Did you know that the password variable was stored at some particular memory location?
Basically you use VirtualAllocEx() to allocate some memory in the remote thread. The returned pointers are in the context of the target process.
You can access that remote memory with ReadProcessMemory() and WriteProcessMemory(), which uses those "remote" pointers to copy data to/from your process.
You can then use these memory areas to pass global handles and other stuff around.
For accessing the actual password field data, you use standard Window-Messages with SendMessage() etc.
Edit: Looks like NVDA still does. https://github.com/nvaccess/nvda/blob/master/source/NVDAObje...
To this day (Win10/Win11) you can hide your program from Task Manager using this technique and any malware that respect itself does it.
https://en.wikipedia.org/wiki/Shatter_attack
You could do that to call WM_GETTEXT, against an input-control, for example.
:)
This is incorrect for C/C++ though. Modern compilers definitely treat null dereferences as UB, with real consequences (e.g. eliminating redundant null pointer checks). The compiler is part of the architecture.
Even Linus got to enjoy that fun.
1. To substract the two function pointers after casting to BYTE*.
2. To read the bytes through those pointers.
3. To cast the copied bytes back to function and invoke it.
The article only focuses on 3, and only about how it's undefined because the compiler is not required to generate position independent code. But an optimizing compiler in theory can just optimize the whole function away from looking at the very first undefined line.
The message is that yes, you need to step out of the standard to do stuff like this, but you may have to consult a lot more about your implementation defined behavior than you originally signed up for.
The failure happens because the programmer assumed that the code is "self contained" and position-independent, both of which are concepts outside of the C abstract machine.
[1]: Function pointers on Itanium are actually fat, but IIRC most compilers hide this by making the "function pointer" point to some kind of thunk instead.
There's literally decades worth of people trying to use undefined behavior in clever ways and ultimately failing and yet here we are...
My only point was that, for better or worse, UB is not the culprit in this code. C could have well-defined abstract semantics for copying functions or aliasing function pointers through datatype pointers, and this code would still be platform dependent and would still break on different hosts.
Saying that C++ (the code in the article is not C and the two languages differ about the treatment of pointers to functions) could have well defined semantics to handle this is entirely moot... it's about as relevant as saying that a C++ program could run on top of the JVM with a garbage collector [1] and hence eliminate all types of memory errors. Even if a C++ program ran on a platform that had guaranteed garbage collection the fact would still remain that C++ as a language does not have well defined semantics for what happens to a dangling reference and as such a C++ compiler is free to exploit that to make very strong assumptions about the runtime behavior of the program for the sake of generating efficient code.
The fact that this is undefined behavior gives a compiler the freedom to perform optimizations under the assumption that runtime behavior will never engender said behavior. Raymond makes use of this property when he discusses COMDAT folding which is a common optimization to elide multiple copies of the same function or to even produce multiple versions of a single function optimized for different scenarios. From the article:
"Even without Profile-Guided Optimization, compile-time optimization may inline some or all of a function, so a single function might have multiple copies in memory, each of which has been optimized for its specific call site."
This property has nothing to do with x86, or Itanium or PDP-11, it's a purely logical optimization permissible only because of the various forms of undefined behavior in C++ with respect to the treatment of function pointers.
C/C++ compilers would then not be portable to ISAs with Harvard architecture.
Fortunately, I followed some of the techniques from “Programming Applications for Microsoft Windows” book and Detours project to intercept and execute custom code mostly based on loading custom DLL in target remote process and using DllMain() to execute.
So yes, the mechanism you posit exists, and its the virtual memory manager.
(VirtualProtectEx looks like it would do that. Never used winapi, not sure.)
What boggles my mind is how they went on to ask MS for help fixing their obviously wrong vulnerability-and-crash-introducing software.
Also, in this particular case the case is old enough that it's not necessarily any AV that's still around anyway.
So actually the opposite is true, if the code wasn't position-independent and was statically located, the assembly code offsets wouldn't need to be updated and you might be able to call a memcpy'd function. Position-independent code could only possibly work if you updated the reference to the symbol metadata structure in the ASM after the memcpy, but at that point you're re-implementing libdl and no longer just using memcpy.