We use lld currently but had to disable threading as sometimes in CI several of our test binaries would get linked at the same time. Whenever this happened the lld instances appear to have spawned enough threads to overload our Jenkins slave to the point that the master wasn't able to reach it and failed the build.
Unfortunately GNU Make offers no such mechanism. And for Ninja build generation I don't think Meson does it either.
If not, any plans to support this in the future? :)
But even without multi-threading, mold is still faster than other linkers. I can think of various reasons why, but I don't know which attributes how much. I believe the biggest contributor is its efficient data structure -- it is hard to make program faster by writing fast code, but it can naturally be achieved by designing efficient data structures. That said, it is hard to compare two or more programs to find out why one program is faster than the others unless their designs are similar.
My assumption is that future machines will have more cores than we have today on average, so I'm optimizing mold for such computers.
From the manpage of mold-1.3.0, I get the impression mold is designed to scale up to 32 cores, but not more.
man mold | rg -C2 32
--threads
--no-threads
Use multiple threads. By default, mold uses as many threads as the number of cores or 32, whichever is the smallest. The reason why it is capped to 32 is because mold doesn't scale well beyond that point. To use only one thread, pass --no-threads or --thread-count=1.
Is this correct or does the manpage need to be updated?Thanks for the great write up as well as mold itself!
Thanks for your good work! I still remember the O(n^2) complexity of ld.bfd when linking C++ code.
I know this is a type, but it is interesting to see knowledge as a score or commodity you can "earn".
( ͡° ͜ʖ ͡°)