Show HN: A 166 KB file for cross compiling glibc for any version, any target
github.com
github.com
In master branch of Zig right now (before https://github.com/ziglang/zig/pull/10330 is merged), status quo:
$ cat ../lib/libc/glibc/*.txt | wc -c
205219
$ cat ../lib/libc/glibc/*.txt | xz | wc -c
22976
Zig supports targeting every version of glibc for any target architecture. The information required to do this takes up 200 KB installation size / 22 KB tarball size, and it has the following problems (which are causing several bugs):* fails to represent symbols that migrate from one library to another between glibc versions
* fails to represent a distinction between functions and objects
* fails to represent object sizes
With this new glibc-abi-tool:
$ cat abilists | wc -c
169421
$ cat abilists | xz | wc -c
24880
165 KB installation size / 24 KB tarball size, and all the above issues are solved. It also should be slightly faster to load/parse for the compiler since it is fewer KB to load from disk.The code that uses this file is here: https://github.com/ziglang/zig/blob/0d7331b6f6af85bd45829514... It creates c.s, pthread.s, dl.s, etc., which are assembled into libc.so, libpthread.so, libdl.so, etc., which are placed on the linker line when cross compiling. Thanks to this improved dataset, they can now accurately model the glibc that will be on any CPU architecture, any glibc version.
If I just naively shipped every version of the .abilist files:
$ cat (find glibc/ -name "*.abilist") | wc -c
37041906
$ cat (find glibc/ -name "*.abilist") | xz | wc -c
205036
...it would be 35 MiB installation size, 200 KB tarball size. So I have effectively achieved a compression ratio of 219:1 by implementing a bespoke encoding of this information.I think over time the C/C++ way is going to go the way of the dodo; IIRC, rust and go both use the prefix-[]. It creates a clear, unambiguous way of defining the type (no "spiral typing" jokes for the modern languages).
fn main() {
let mut a: [[i32; 2]; 5] = [[0,0],[0,0],[0,0],[0,0],[0,0]];
println!("{:?}", a[0]);
}
Note that the Rust array declarations are also homologous to other types in the ecosystem, such as nested Option types or nested collections (eg: Option<Option<T>> or List<List<T>>).Contrast w/the C version, where it's obvious to a skilled C programmer but you need to remember to read the bounds from right to left to determine the memory layout (ie: an array of 5 arrays of two elements each):
int main() {
int i[5][2] = {{1,2},{3,4}, /*...*/};
printf("%ld\n", sizeof(i[0])/sizeof(int));
return 0;
}
In this case you are required to mentally combine the type on the left of the variable with the bounds on the right of the variable to parse a type declaration in your head that eventually looks like the Rust version: (int[2])[5]
I've been doing a lot of professional C work lately so it's currently "swapped in", but I find that any time I pick C up after spending time away I find it ambiguous and I have to reason through the ordering of the declarations.Obviously this is all IMHO and YMMV, but this developer does find it significantly easier to unambiguously parse Rust types.
To me, [[i32; 1]; 10] is homologous to Items<Items<T, ZeroOrOne>, Unlimited>, not Vec<Option<i32>>. To me, it make more sense to put the size before the inner type (either [1; i32] or [1]i32), since you need to index the array before getting an instance of the type.
Say if I am using a language/toolchain which pulls in a C/C++ compiler, are you able to substitute "zig cc" there and explicitly cross-compile to a different GLIBC version?
Can I use that outside of Zig, for compiling C++ code?
Versioned symbols get added when a function or object breaks binary compatibility. You keep around the old version, and compile against the new version. This might be something simple like a struct layout change, or a change in the size of an array. Both versions have the same name. When you compile, you'll get an unversioned symbol reference. When you link, that will get resolved to whatever the latest version is.
Windows handles this by giving a new name to new symbols. For example, MoveFile(), MoveFileA(), MoveFileW(), MoveFileExA(), MoveFileExW(). You get the correct symbol when you compile, either by calling the correct function directly, or by referring to it by macro (MoveFileEx is a macro for MoveFileExA or MoveFileExW). Newer versions of Windows also make you embed the GUID of the versions of Windows you support in your application manifest, and Windows will run your application in a compatibility environment (at runtime) matching the latest available version you specify.
On macOS, similarly, it uses the preprocessor, but it's a bit different and macOS also supports something called "weak linking" (not the same thing as weak linking in GNU Binutils / ELF). Weak linking allows you to link against an old version of a library, but use a symbol from a newer version... if the symbol is not available, it is NULL. This is done with the preprocessor and the linker in tandem... there's a preprocessor macro which specifies which minimum&maximum version of macOS (or iOS, etc) you target, and that affects which symbols are declared as being weakly linked.
The long and the short of it... you can use the latest macOS SDK to make binaries compatible with old versions of macOS, and same for Windows, but glibc maintainers have not made this possible for glibc.
The zig language doc is pretty good and short, and the lib/std folder is full of pretty easy to understand examples.
This RH blog post is good intro on the topic: https://developers.redhat.com/blog/2019/08/01/how-the-gnu-c-...