What's the most portable way to include binary blobs in an executable?
tratt.net
tratt.net
We do this for Python applications, by combining a ZIP containing the "link tree" of sources/packages/modules, with a shell bootstrap script that automatically sets up the environment, import path, etc, and Python itself has built in support for importing pure-python modules from a ZIP file. All that's needed for native modules is a simple import hook that extracts the native objects into temp space and then loads them appropriately.
https://blog.frankel.ch/creating-self-contained-executable-j...
https://coderwall.com/p/ssuaxa/how-to-make-a-jar-file-linux-...
Hey, I wrote objcopy and objdump as debugging tools for developing bfd targets when I was designing bfd. Neither was (originally) intended as a production tool. OK I still use them myself, and I haven’t looked at the bfd source code in almost 30 years.
I forget who came up with the tool chain target names, which I think was your real complaint, though I remember when. Perhaps Ian Taylor (later author of gold, which was notable for, among other things, not using bfd).
Objdump should have a feature to generate assembly or c code for an arbitrary blob (with correct byte swap, if needed, of course).
> the tool chain target names, which I think was your real complaint
Yes, this was solely what I was referring to (slightly thoughtlessly) as "ugly", and even then only in the sense of "where did those magic names come from?" I certainly wasn't referring to objcopy itself!
We have recently open-sourced a small tool called Postject (https://github.com/postmanlabs/postject), which is able to inject arbitrary data as proper ELF/Mach-O/PE sections for all major operating systems (with AIX support coming). The tool also provides C/C++ cross-platform headers for easily traversing the final binary and introspect whether the segment is present or not.
The tool is based on the LIEF (https://github.com/lief-project/LIEF) project.
At Postman, we are making use of this on our custom Node.js single-executable applications and soon on our custom Electron.js builds too.
https://news.ycombinator.com/item?id=32201951
C++ added "std::embed" https://open-std.org/JTC1/SC22/WG21/docs/papers/2020/p1040r6...
My solution is [1]. It generates a C file with a specific array name passed in through the command-line. It also has a few other niceties that I need.
It works on Windows, Mac OSX, Linux, and the BSD's, no matter the compiler or linker.
I use it to generate the arrays for the help texts ([2] and [3]), as well as two math libraries ([4] and [5]).
People are welcome to adopt and adapt it. Just follow the license, as per usual. I've even adapted to my other software. [6]
[1]: https://git.yzena.com/gavin/bc/src/branch/master/gen/strgen....
[2]: https://git.yzena.com/gavin/bc/src/branch/master/gen/bc_help...
[3]: https://git.yzena.com/gavin/bc/src/branch/master/gen/dc_help...
[4]: https://git.yzena.com/gavin/bc/src/branch/master/gen/lib.bc
[5]: https://git.yzena.com/gavin/bc/src/branch/master/gen/lib2.bc
[6]: https://git.yzena.com/Yzena/Yc/src/branch/master/tests/strge...
It did 13875457.34 bytes per second, which means my program would have done the 82 MiB file he had in 6.20 seconds, faster than hexdump, GCC, and Clang.
It used a max of 3223848 Kbytes, which is about the size of the file it was processing. (The file was 3300000020 bytes exactly.)
I also tried with a file as close to 82 MiB as I could get. It used 85168 Kbytes max, and it took 6.77 seconds.
My code could probably be optimized too. It tries to skip a header comment. It also reads all of the input file in at once, when it could probably stream it on demand. It is also checking for stuff to exclude, which takes time.
If I take out the if statement that begins with:
if (!strncmp(in + i, bc_gen_ex_start, strlen(bc_gen_ex_start)))
That should remove most of everything else that does not matter.If I run that on the 82 MiB file I made, it uses 85124 Kbytes and 5.40 seconds.
This is still far slower than objcopy and incbin, but maybe it's good enough, right?
I think I could optimize it more, but I still think that's a pretty good showing against the competition, especially for portability.
Edit: I forgot to mention that I did these tests while running a fuzzer (AFL++) on 15 of my 16 cores and while watching YouTube. I didn't want to stop the fuzzer just for this (it's been running for more than 24 hours).
Mentioned it to an old Unix hand, she replied "hmm, what about `xxd -i < file`"?
Live and learn.
But yes, xxd is an easy solution.
Granted I can replace it with xxd and two lines of awk, but that would have taken longer, and what's done's done, it's not a task with shifting requirements.
It's not "portable" per-se, but all modern platforms [1] have a way to interrogate the binary contents of the currently-running executable.
[1] _NSGetExecutablePath, GetModuleFileName(), getexecname() etc
EDIT: Apparently https://github.com/gpakosz/whereami will manage a lot of this complexity for you
Back to this particular case, the binary will fail strict code signing validation on macOS. It may still run because the kernel does not access the binary past the coverage of the code signature (and all the bits there are still intact), similar to how multiarch binaries work, but you will at least severely be hampered to distribute your binary, since Gatekeeper won't be happy either.
I've written loaders for all of the executable formats you mentioned, and maybe a dozen more. I know of none where this would violate the strict interpretation of the word of the spec.
That being said, valid file != happy OS
[0] https://www.sqlite.org/src/file/ext/misc/appendvfs.c
Edit - the very first comment explains that is uses a trailing string
https://github.com/ebu/libear/commit/40a4000296190c3f91eba79...
This is a cmake function which generates C++ files using no external tools. It's probably not very fast, but if you don't need to handle big files and are already using cmake this is easy to integrate, adds no dependencies and works on all platforms.
What I found is that many compilers don't like to compile very large source files; so if the binaries you'd like to integrate are big, it might be better to integrate their constituent objects one by one (if applicable).
https://www.nongnu.org/txr/txr-manpage.html#N-0389D15E
There is a 128-byte area prefixed by the character sequence @(txr):. It normally contains all zeros (empty null-terminated string). If you put a non-empty UTF-8 string there, it gets executed.
Of course, the problem of including a binary blob is trivial if it can just be declared as an array; the interesting problem is doing it to the executable, without doing any compiling or linking.
Just template generate and store the data as a bit array on the language of your choice.
For example, if you are using C/C++ you can zip everything then use a small python script to generate a C/C++ header where this data is available as a uint8_t array.
Keep in mind that all this data will be loaded to memory, so I don’t recommend this approach for anything north of 10mb.
.incbin "string_blob.txt"
...
printf("%s\n", string_blob);
Text files don't have to have a NUL termination. The proper way to embed data with the .incbin directive is to add a label after the file and use that directly for pointer arithmetic or compute the size with another assembly directive.The thing I was patching up was a large rust program (30+mb release stripped), and so it was undesirable to always have blob changes require a relink of the program.
Embed as it is available in a few common languages now/soon is very convenient, but it is quite painful once the program link stage is expensive.
Probably doesn't work on windows, no measurable compile time cost (unlike massive arrays of comma separated chars).
xxd --include name < binaryblob