Firefox switching to clang-cl for Windows builds
groups.google.com
groups.google.com
"At the moment, performance is a mixed bag. Some tests are up and some are down. In particular I believe Speedometer is down a few percent."
"Note however that clang-cl is punching above its weight. These builds currently have neither LTO nor PGO, while our MSVC builds use both of those. Any regressions that we're seeing ought to be short-lived. Once we enable LTO[1] and PGO[2], I expect clang to be a clear performance win."
[1] "LTO (Link Time Optimization) is a method for achieving better runtime performance through whole-program analysis and cross-module optimization." http://blog.llvm.org/2016/06/thinlto-scalable-and-incrementa...
[2] PGO = Profile Guided Optimization
1. https://bugs.chromium.org/p/chromium/issues/detail?id=580389...
Performance is only one aspect of it. It also reduces code bloat, reducing the program's size footprint. Most tests (yes, I know, not all, but most) should not make it into the final binary users are running. I also don't see what's "hacky" about making a foo.test.cc file when I want an alternate implementation for foo.cc. It seems to be quite a positive and clear way to document the fact that an implementation is only needed for testing, and vice-versa. And not only that, but it reduces compile (& link) times, since you only need to compile one of the two implementations for each use case.
Unfortunately a lot of C++ developers don't really understand how the linker works, and because of that the build system is black magic.
I'll reconstruct my first take from memory if you have comments on it:
Current CPUs are pretty good at predicting indirect branches, and it's hard to tell beforehand which virtual calls still turn out to be perf problems - and it's wasted effort and a fallible process to attempt their compllete elimination beforehand.
Depends. Virtual calls can be done purely in code, with different implementation files you need to go round via the build system. The former just always works independent of platform whereas often for the latter it's more configuration work. Also, ignoring YAGNI, you could say the virtual one makes it easy to quickly test things out etc, is easily mockable, ...
I don't fully understand the proposal here. Do you have an example of a code in the wild that uses the proposed scheme?
// widget.h
class Widget { virtual void throb(); };
// widget.cc
class WidgetImpl : public Widget { void throb() { ... } };
class WidgetTest : public Widget { void throb() { ... } };
then, instead of that, just do // widget.h
class Widget { void throb(); };
// widget.cc
void Widget::throb() { ... }
// widget.test.cc
void Widget::throb() { ... }
where you only compile widget.cc for the production build, and only compile widget.test.cc for the test build.Feel free to post a more specific example and I can take a stab at seeing if I can make this transformation to it (or why it might not be possible, if you think it's not).
Essentially, it was easier to write a compiler / linker optimization, than change the source code. On the bright side, everyone wins!
I mean there's no reason you have to change everything by tomorrow. You could introduce it as a policy moving forward, the same way I'm sure you make any other changes to other patterns that need to be followed. I'm sure you have a process for this?
> Essentially, it was easier to write a compiler / linker optimization, than change the source code.
I agree, but see above and below.
> On the bright side, everyone wins!
Well, you waste space, time, and energy making the compiler and linker do work that doesn't inherently need to be done. You also lose parallelizability of the compilation of the two implementations if they currently reside in the same file. Finally the compiler (er, linker) might not be able to actually apply the devirtualizations you expect, whereas in this case it's just a matter of inlining since the target method is already known.
I don't know what you're referring to as the "suboptimal code" here. If it's the old pattern using virtual everywhere, then the ship's already sailed... that code is already in production. If you mean the two-file approach I proposed then that makes no sense; if it's worse than the last approach then why would you even use it.
So in this example you have to weigh the benefits of the two file tests against having tests that work uniformly.
It's repetitive, error-prone work, better to have the compiler and linker do it than rely on the programmer getting their use of the preprocessor right.
And even if the programmer gets it right, doing it via macros means every tool that you want to apply to your codebase - IDEs, profilers, coverage, instrumentation - needs to understand your macro. Are you sure they'll all get it right?
Better to write plain old standard code that every tool will work correctly with, and the worst thing that can ever happen is a slight performance penalty.
Like the other commenter said, I'd rather do simple, virtual calls and then fix it using LTO.
You're suggesting that instead of a compiler wasting time space and energy, that we get actual people to waste their time space and energy by manually applying the transformation that the compiler is able to do anyway?
"We do it that way here" is not something that has to be always applied.
void DoSomething(widget* param, OtherClassWithvirtual* param2) {...}
My DoSomething function needs to link to widget, but now I have two different libraries that widget could be in. The linker encodes in my binary which library to load to get widget. All attempts to use the other library instead are undefined behavior.The only way I know of to work around that is build DoSomething twice, linking the real widget once and the test widget the other. Note that DoSomething takes two parameters, and those are classes that may themselves also take other classes with virtual functions. All the possible cases quickly gets into a combinatoric explosion and my build times goes from an hour (already way too long) to days.
If I'm wrong please tell me how, I'd love to not have so many interfaces just because someone doesn't want their unit test to become a large scale integration test.
Because that’s what we need, more magic in the build system.
// widget.h
class Widget { void throb1(); virtual void throb2(); };
// widget.cc
void Widget::throb1() { ... }
void Widget::throb2() { ... }
// widget.test.cc
void Widget::throb1() { ... }
void Widget::throb2() { ... }
class TestWidget : public Widget { int x; void throb2(); };
void TestWidget::throb2() { ... Widget::throb2(); ... }
If not, could you present code to illustrate the problem? I won't try to address more issues without actual code.Very cool. The potential to inline Rust and C++ (and C, of course) code together at the LLVM level is one of those things that's been theoretically possible for a long time now but has yet to really mature in practice (at least, I'm not aware of anyone aside from Mozilla attempting it at scale). Really curious to see how this affects Firefox performance once they get it all wired up.
I'm more interested in seeing whether the maintainability of Firefox will increase over time as the proportion of Rust code increases. That is, will Mozilla be able to accelerate Firefox's development because the source code is in a safer language?
https://support.mozilla.org/en-US/kb/firefox-uses-too-many-c...
If it is still happening after you try those, use https://perf-html.io/ and share a profile.
Using Firefox over Safari can cost hours of battery life, so it’s very hard to recommend it. Is this just a Mac issue?
https://bugzilla.mozilla.org/show_bug.cgi?format=default&id=...
Also nightly contains all sorts of debug stuff
I'll see how it goes with the normal release
> We've discussed this on several occasions but apparently nobody ever filed a bug.
> There are a number of reasons using clang-cl for our official builds would be good:
> * The ability to patch the toolchain as we do on other platforms
> * The ability to diagnose and upstream fixes for compiler bugs we encounter
> * It would provide motivation to switch to clang on all platforms, which would make it so we only have to worry about a single compiler for things like performance optimization (I don't assume we would drop support for MSVC or gcc, we just wouldn't have to worry as much about compiler quirks)
Compiler Version Compiler flags
z:/build/build/src/clang/bin/clang-cl.exe -Xclang -std=gnu99 -fms-compatibility-version=19.13.26128 19.13.26128 -Qunused-arguments -nologo -wd4091 -D_HAS_EXCEPTIONS=0 -W3 -Gy -Zc:inline -Gw -wd4244 -wd4267 -Wno-unknown-pragmas -Wno-ignored-pragmas -Wno-deprecated-declarations -Wno-invalid-noreturn -we4553
z:/build/build/src/clang/bin/clang-cl.exe -fms-compatibility-version=19.13.26128 19.13.26128 -Qunused-arguments -Qunused-arguments -TP -nologo -w15038 -wd5026 -wd5027 -Zc:sizedDealloc- -wd4091 -wd4577 -D_HAS_EXCEPTIONS=0 -W3 -Gy -Zc:inline -Gw -wd4251 -wd4244 -wd4267 -wd4800 -wd4595 -wd4065 -Wno-inline-new-delete -Wno-invalid-offsetof -Wno-microsoft-enum-value -Wno-microsoft-include -Wno-unknown-pragmas -Wno-ignored-pragmas -Wno-deprecated-declarations -Wno-invalid-noreturn -Wno-inconsistent-missing-override -Wno-implicit-exception-spec-mismatch -Wno-unused-local-typedef -Wno-ignored-attributes -Wno-used-but-marked-unused -we4553 -D_SILENCE_TR1_NAMESPACE_DEPRECATION_WARNING -GR- -Zi -O2 -Oy-Since the headline doesn't mention MSVC, this could be interpreted as relevant to the old GCC vs. Clang issue, but it really isn't in this case.
To use Google Groups Discussions, please enable JavaScript in your browser settings, and then refresh this page.
NoScript users will be automatically redirected to this version (which doesn't require it):
https://groups.google.com/forum/m/?_escaped_fragment_=topic/...
Sadly it probably doesn't make sense for them to allocate resources to that. On the other hand in my experience at least GCC dominates in performance department so that would help with not worrying about compiler switch causing performance regression.