Executing an array as a function
kahlonel.io
kahlonel.io
short main[] = {
277, 04735, -4129, 25, 0, 477, 1019, 0xbef, 0, 12800,
-113, 21119, 0x52d7, -1006, -7151, 0, 0x4bc, 020004,
14880, 10541, 2056, 04010, 4548, 3044, -6716, 0x9,
4407, 6, 5568, 1, -30460, 0, 0x9, 5570, 512, -30419,
0x7e82, 0760, 6, 0, 4, 02400, 15, 0, 4, 1280, 4, 0,
4, 0, 0, 0, 0x8, 0, 4, 0, ',', 0, 12, 0, 4, 0, '#',
0, 020, 0, 4, 0, 30, 0, 026, 0, 0x6176, 120, 25712,
'p', 072163, 'r', 29303, 29801, 'e'
};
(example is for Vax-11 or pdp-11, and courtesy of http://www.ioccc.org/1984/mullender/mullender.c)http://jroweboy.github.io/c/asm/2015/01/26/when-is-main-not-...
Sure, dlopen does tons of more stuff behind the scenes, but ultimately it is about loading bunch of bytes into memory and executing those as functions.
That's pretty much mprotect, though. dlopen mostly does the other things, and it does a lot.
Now if only someone would write a compiletime C compiler in C++ templates...
I have some home computer books from the late 70s/early 80s that use this trick extensively.
On the era where home computers had only slow BASIC interpreters and no assemblers (or compilers), the usual way for speed up was to type in a long sequence of numbers (or characters) that were actually a machine language program.
So you have a line like:
1000 DATA 100,32,65,12,44,32,52,11,255,12,55,22
and on and on, which hold a sequence of bytes (the machine language program)
and later, READ statements would read each byte (of the machine language code) and POKE them into memory, that is, write it into a specific address of the RAM...
... later you CALL to that specific address, which basically instructs the BASIC interpreter to "jump" to the machine language code at that location.
The important thing to know is that 8051s have separate CODE and RAM address spaces. (Actually the RAM is divided up into multiple flavors too: direct, indirect, external and bit addressable.)
It turned out it wasn't worth the overhead. The interpreter and ancillary bits took up too much space and slowed things to a crawl. It was generally easier to rewrite sections of the code in a way that made the compiler and optimizer happy. By various techniques I managed to reduce the footprint of the system code by at least fourfold after a number of refactorings. The application kept needing new features, so it always barely fit in 16 code address space (some of which was already consumed by the bootloader).
Enough said.
Remember in our scenario, we need the file to compile with a gcc on a 64 bit system, without any special modifcations to the compiler flags, so that means there is no special compile flags, nor can we include any custom linking steps and we want to use GCC inline AT&T syntax.
Edit: whoops, that was from a link to a similar thing in the comments. I apparently missed the real article, which I'm off to read now.
This is a bit different from constructing the function "from scratch," though. I tried to inspect the guts of the function body itself, but was met with lots of segfaults.
x86 will allow this if the OS does. There have been many CPUs which don't allow it, such as PowerPC, which had separate instruction and data caches. After loading code, the loader had to make the pages executable and cause a cache flush before the code could run.
[ (Hello World) /print cvx ] cvx exec
Of course, procedures are just shorthand for executable arrays, so { (Hello World) print } exec
does the same. Oh did someone say "C"/machine code? Never mind ;-)