Can I exec a new process without an executable file? (2015)
unix.stackexchange.com
unix.stackexchange.com
[0]: https://github.com/impl/systemd-user-sleep/blob/666cf29871b1...
[1]: https://github.com/impl/systemd-user-sleep/blob/666cf29871b1...
int main(int argc, char *argv[]) {
#define TINY_ELF_PROGRAM "\
\177\105\114\106\002\001\001\000\000\000\000\000\000\000\000\000\
\002\000\076\000\001\000\000\000\170\000\100\000\000\000\000\000\
\100\000\000\000\000\000\000\000\000\000\000\000\000\000\000\000\
\000\000\000\000\100\000\070\000\001\000\000\000\000\000\000\000\
\001\000\000\000\005\000\000\000\000\000\000\000\000\000\000\000\
\000\000\100\000\000\000\000\000\000\000\100\000\000\000\000\000\
\200\000\000\000\000\000\000\000\200\000\000\000\000\000\000\000\
\000\020\000\000\000\000\000\000\152\052\137\152\074\130\017\005"
int fd = memfd_create("foo", MFD_CLOEXEC);
write(fd, TINY_ELF_PROGRAM, sizeof(TINY_ELF_PROGRAM)-1);
fexecve(fd, argv, environ);
}
Who here is brave enough to run my C string? #define _GNU_SOURCE
#include <sys/mman.h>
#include <unistd.h>
I've tested in Fabrice Bellard JSLinux with tcc (x86 arch) and on https://replit.com/languages/c (x64). I failed to see any side effect at all. gdb "catch syscall" doesn't show anything interesting too. Looks like TINY_ELF_PROGRAM is not doing anything. push 0x2a
pop edi
push 0x3c
pop eax
db 0x0f, 0x05 ; invalid? ; set the first syscall argument to 42
push 0x2a
pop edi
; select syscall 60 (sys_exit)
push 0x3c
pop eax
; sys_exit(42)
syscallEdit: Wouldn't mov be shorter than push/pop? (I am not very familar with x86)
Nowdays with a combination of ebpf, apparmor, cgroups, kvm, nx-stack and a strict firewall it's possible to almost entirely prevent external code from being run (after performing in-depth profiling of its intended behaviour). Sadly nobody does that, and if anything Linux on the desktop is missing most, if not all, of the process isolation features Android and iOS have.
You mean fortunately? Android and iOS are siloed walled gardens, not general-purpose OSs.
Or, better put, this is not like making Linux a "walled garden" but being able to put walls around each app you run - which is different.
As for the side-question of switching between 64 and 32 bit mode in the same process, this is classically known on Windows as "heaven's gate" and a similar technique on Linux seems possible too: https://gist.github.com/rqou/1a1834b784283add7955af430097311...
See something I published just a month ago: https://github.com/anvilsecure/ulexecve/
See for example https://www.anvilsecure.com/blog/userland-execution-of-binar.... The implementation is on GitHub and rather clean if I may say so myself (am the author).
Edit: I was wrong about the names given to memfd objects, I thought they showed up under /dev somewhere but they’re purely for debugging purposes.
It's truly great for situations where APIs refuse to take anything other than files and you don't worry about cleanup. Ex: loading certs from memory into a python openssl context.
(I understand this is not necessarily a solution to the question and the existing solution is probably a better fit, but I'm curious)
switch_to_64bit:
pop edx ;EDX=return address
xor ebx,ebx ;EBX=selector
.next_sel:
add bx,8 ;try next
jc .exit ;none found, -> segfault
lar eax,bx ;load access rights
jnz .next_sel ;failed?
and eax,0x60F400 ;mask bits
cmp eax,0x20F000 ;64bit code selector?
jne .next_sel ;no
.exit:
or bl,3 ;set RPL to ring3
push ebx ;selector
push edx ;offset
retfd ;go thereThen there’s miscellaneous stuff like cloexec. Not privileged, but atomic.
The binary you exec may load code at the same address you’re using for code, unless it’s PIE. Not insurmountable, but tricky.
See https://github.com/anvilsecure/ulexecve/blob/main/ulexecve.p... for details. Especially the `CodeGenerator` classes with implementations in x86, x86-64 and aarch64.
Unfortunately we don't use this all the time because some Kubernetes unit tests started failing when we first added this protection (the size of the binary is added to the memory usage of each container which caused some Kubernetes unit tests to use more memory than they did before). Ironically this exact protection would've protected us from Dirty COW and other such bugs but it's disabled by default (instead we make a temporary read-only bind-mount that we then exec which is slightly less safe but doesn't add ~10MB to every containers' memory usage).
But the actual answer to your question is that this was not originally intended behaviour (when we mentioned we were doing this to the mm and fs folks they weren't happy) and there have been patches posted recently to make this feature something you have to explicitly opt-in to.