Mold – A really fast linker
github.com
github.com
Mold linker: targeting macOS/iOS now requires a commercial license - https://news.ycombinator.com/item?id=34141912 - Dec 2022 (74 comments)
Mold linker may switch to a source-available license - https://news.ycombinator.com/item?id=33584651 - Nov 2022 (206 comments)
Mold linker creator considers changing the license - https://news.ycombinator.com/item?id=33495528 - Nov 2022 (19 comments)
Mold/macOS is 11 times faster than the Apple's default linker to link Chrome - https://news.ycombinator.com/item?id=31769699 - June 2022 (116 comments)
Using the mold linker for fun and 3x-8x link time speedups - https://news.ycombinator.com/item?id=31604772 - June 2022 (41 comments)
Using the mold linker for fun and 3x-8x link time speedups - https://news.ycombinator.com/item?id=31592678 - June 2022 (1 comment)
Mold 1.0: the first stable and production-ready release of the high-speed linker - https://news.ycombinator.com/item?id=29568454 - Dec 2021 (65 comments)
Mold: A Modern Linker - https://news.ycombinator.com/item?id=26233244 - Feb 2021 (122 comments)
Mold: A Modern Linker - https://news.ycombinator.com/item?id=25410312 - Dec 2020 (1 comment)
Which would also be a good test of how good of a "drop-in replacement" mold is given that the Blender build process isn't trivial.
Given the final program needs to be storable in RAM + Virtual Memory it surprises me that we still need the intermediate step of pushing to the file system only to then immediately reopen and merge those files.
Does someone have more info on this? Or is the reason just legacy? Or is the reason just some ideological “single responsibility” thing?
In some places, where builds are particularly optimized, there are special distributed filesystems just for object files. In this case, it's not necessarily true that the object files are backed by disk even.
Being backed by disk locally mainly helps for incrementally building so that you can change one file and only recompile intermediate files for translation units that depend on this file. Disk/FS caching presumably helps a lot with redundant I/O, and I think most of the benefit of building with tmpfs winds up being putting the source itself in RAM.
Edit: also it's worth noting that many compilers can output object files which can be linked by other linkers, allowing you to mix output from different compilers in some circumstances.
> Being backed by disk locally mainly helps for incrementally building so that you can change one file and only recompile intermediate files for translation units that depend on this file.
For the local incremental usecase I’d love to see a more state-full compiler instead. One that could change the bytes of a binary instead. e.g. give all functions some additional “empty padding”. Then any modifications could directly fit into the binary as needed until some defragmentation process which creates the final output binary.
That said, it accomplishes this still using an on-disk datastore.
The first thing to point out is that reading a recently-written file isn't all that expensive: it's stored in the filesystem cache anyways, so you bypass the disk for the reads.
Moreover, the filesystem is actually a decent database for multiprocess communication. If every translation unit is compiled into an independent file, which is then combined into a single final output, there is no need to build any complex locking mechanisms or the like and you still get to take advantage of the embarrassingly parallel nature of compiling.
Incremental compilation is an incredibly important tool. If you make a small change to one file, it's frequently not necessary to rebuild most of the code. Making the output of individual file compilations work in a way that allows incremental compilation to happen requires basically building .o files--and there's very little savings to be had by not emitting them to disk.
Finally, I'll note that very frequently, debug builds of large applications cannot fit in RAM. Debugging symbols bloat builds tremendously, especially in intermediate object form (since many symbols end up needing to be duplicated in every single .o file). A debug build of a large application may take up 80GB in disk space, to build a binary that (without debug symbols) would be perhaps 100MB in size.
Keeping everything in RAM just isn't feasible at scale.
You don't. I recommend you learn about unity builds, and how major build systems support toggling them at the project and subproject level.
The main reason why most people haven't heard about the concept and those who did the majority doesn't bother with it is that a) you have little to nothing to gain by them, b) you throw incremental builds out if the window, c) you ruin internal linkage and thus can introduce hard to track errors.
Also, it makes no sense at all to argue how the released software needs to run in memory to justify aspects related to how the software is built. At most you have arguments over code bloat and premature optimization, but at what cost?
Specifically, to produce a .o file the compiler has already read through and created indexes of module::function/struct names, their layouts, dependencies, etc.
My understanding is that the .o files need to be re-parsed by the linker to create these indexes, layouts, and dependencies. Especially with LTOs I'd imagine there would be additional inlining work (stuff the compiler is already good at).
This is all just wasted time, including the IO bottlenecks - even if those are marginal.
There were subtle differences to gold and ldd that I didn’t have time to chase down, but it seems like the future.
Also the speedup is only for Debug builds, since Release ones use LTO anyway, and no amount of linker magic makes that fast.
I think you need to use a log scale for this. It's a step or two toward perfection, but perfection is infinite steps away.
> Open-source license: mold stays in AGPL, but _we claim AGPL propagates to the linker's output_. That is, we claim that the output from the linker is a derivative work of the linker. That's a bold claim but not entirely nonsense since the linker copies some code from itself to the an output. Therefore, there's room to claim that the linker's output is a derivative work of the linker, and since the linker is AGPL, the license propagates. I don't know if this claim will hold in court, but just buying a sold license would be much easier than using mold in an AGPL-incompatible way and challenging the claim in court.
Regardless of their current stance, this type of policy changes on a whim led me to remove mold from any of my systems, since I don't want all of my code in the future to automatically become AGPL, even by accident.
[0] https://bluewhalesystems.blogspot.com/2022/11/mold-linker-ma...
Ehh...
"I want to share another idea in this post to keep it open-source [..] Let me know what you guys think" is not a "policy change on a whim". It's an idea. It was not "walked back" on, because it was ... just an idea.
Your comment is a horrible misrepresentation of what's actually in the post.
> mold stays in AGPL, but _we claim AGPL propagates to the linker's output_ [...] I don't know if this claim will hold in court, but just buying a sold license would be much easier than using mold in an AGPL-incompatible way and challenging the claim in court.
Sounds like protection money to me
One practical problem, even for developers who are fine with licensing their code to whatever open source license is easiest, is that not all open source licenses are AGPL compatible. Take for example the Mozilla Public License, which is inherently incompatible with AGPL because of the terms imposed; this means that any project using MPL licensed libraries could no longer be linked with Mold.
License incompatibilities can be a huge pain (see: ZFS + Linux). If you develop software for yourself this isn't a problem, but if you intend to distribute your software this becomes more of an issue.
This is probably also the main reason why normal linkers/compilers don't impose licenses on the produced work.
The issue was that they wanted to claim that AGPL was contagious - That by using mold, your outputs would also be required to become AGPL.
The unusual thing here is that the creators of a linker are apparently trying to have the copyleft licence propagate to code that is input to the linker. Others have pointed out that GCC has exceptions for this kind of thing, despite that it is released under a strong copyleft licence (GPLv3+).
Also, your account of copyleft is still incorrect. It's true of the GPLv2 and GPLv3 licences but not true of all copyleft licences. The AGPLv3 licence, which is the one relevant here, doesn't apply only on distribution.
[0] https://www.gnu.org/licenses/copyleft.en.html
edit I think I was mistaken in putting propagate to code that is input to the linker, though. As lokar's comment points out, it's instead about the output of the linker.
Not sure we can say that a linker is the same as a compiler in this sense, but if so, maybe it is indeed worrisome.
[0] https://sourceware.org/git/?p=binutils-gdb.git;a=blob;f=READ...
The standard compiler license exception (this applies to LLVM to, e.g.) says that any such code that gets combined in with your application code doesn't count. Note that it's still a potential license violation to use that code elsewhere (say, using those routines in another compiler).
This isn't a concern for linkers because linkers don't really provide anything in the way of code, everything being provided by the compiler as a compiler or language support library. The largest code it might add to your program is probably the PLT stub code, at best a couple of instructions long.