A clean solution would be to have stub executables that link to a central libClang.dll/libLLVM.dll (-DCLANG_LINK_CLANG_DYLIB=ON), but this is not supported under Windows because a DLL can export at most 2^16 symbols. Some work would need to be invested to make process launching on Windows work differently than on other platforms, but then disk space is not that much of a problem.
(Also note that in this case you don't need symlinks per se, just identical binaries with hardlinks. Each executable could inspect argv[0] and figure out what it's invoked as, and behave accordingly.)
What I don't understand is why can't they just bundle everything into a single DLL, then make whatever stub executables they want just call the exported main() in that DLL. I think that might be what MSYS2's copy already does, though I haven't checked. That DLL wouldn't need to export anything else, just the main() function that each .exe stub would forward to. And it can handle everything else internally if it so desires.
Symlinks have been supported since Vista, though by default creating a symlink does indeed require Admin rights. Hardlinks can be created by anybody and are extensively used by Windows itself.
I'm not aware of even Windows 10 or Windows 11 being able to create Symlinks on FAT32 and exFAT, but I could be wrong? But it doesn't matter, the code as written is cross platform, it works on Apple Journaled File Systems, APFS, exFAT, etc. We then compile it for Macintosh, compile it for Windows, and compile it for Linux.
We can then spend extra time and carefully detect each filesystem and each platform and then make the optimization if we can. And this is a valid criticism that we have not done this yet. But no matter what we need this general code that will always work FIRST, what the links are is a space optimization to save valuable SSD space when it is possible.
Or are you saying your heuristic for determining the underlying engineering quality changes its output based on 1 bit flip? Identical executables is awful engineering, but a few bytes different (out of 90MB) is great engineering?
Really? You genuinely don't see why I would think a case with >99.99% identical executables might be relevant to a case with 100% identical executables?
With 100% identical executables, there does not appear to be a good reason to have multiple copies under any circumstance.
With anything < 100% identical, well, maybe there’s a good reason to have multiples. Who knows? Id probably give someone the benefit of the doubt and figure there was some engineering challenge that made it faster/easier to do it that way.
So yes, 100% identical is completely different than almost 100% identical.
They're completely (!) different? And you're saying this despite the fact that the comment I replied to was discussing cases where one could "just execute one binary n times (with different argv[0] if they'd like)"... which is something you can do with different executables just as well as with identical ones? It's not just a little different but completely different? So different that not only you don't see any similarity, but you also cannot fathom why I might think there's some similarity?!
If they're 99% the same, it's generously easy to assume that there's a material difference. That assumption is completely nonsensical if they're 100% identical.
So no, for the sake of every bit of context in this conversation, it does not make sense that you'd bring up an unrelated scenario of "similar" binaries.
Context is important. For example, chimpanzee DNA is 99% identical to human DNA. Without context you could make the argument, “They are almost identical!” and be correct. But in the context of discussing the ability to fly to the moon and return safely, the two genomes are completely different.
But with computer code, you very much can (for example) easily replace all the executables with a combined one that gets hardlinked with different names, such that they behave differently depending on argv[0]. They don't even need to be 100% identical for us to be able to do this; it works just as well with 1% identical code—you just end up with bigger combined binaries. This is the Busybox approach (notice it combines executables that have hardly anything in common), and it's in fact exactly what Clang already sometimes does (like on MSYS2); one would think they could take the same approach here. This is also precisely the context of the comment I replied to was saying, right? This is the context I was replying to. It's so confusing to me that you claim I'm missing context when I was in fact addressing the context I saw directly—and it was the other comments that were not.
You can instead configure LLVM’s build system to build a single dynamic library and have the tools link to it, and this eliminates all of the duplication. However, it apparently comes with a “substantial performance penalty” [1] due to the nature of dynamic linking. (This actually surprises me, and I wonder whether it’s only referring to the inability to do LTO, or whether even LTO-less static linking is faster. Aside from the startup time issue.)
A theoretical alternative would be to build all of the tools into a single executable, à la Busybox, where the combined executable would inspect argv[0] to figure out which tool’s code should be run. That way you could statically link the LLVM libraries without duplicating them in multiple executables. LLVM’s build system does not support this. I think it would be nice if it did, but it would be nontrivial to implement.
I think the latter is also the case (though to a much lesser extent) on x64. One of the unfortunate features of x64 is it lacks direct 64-bit jumps, so every jump to an external library ends up being an indirect call. (In fact, with a potential memory load on top of that.) This was kind of surprising for me when I learned it too; it doesn't apply to x86.