OpenMandriva, the first Clang-built Linux distribution
openmandriva.org
openmandriva.org
For the end user, it comes down to - do you care about copyleft and what FSF/Stallman are advocating for, or do you just want a free (as in beer) code, (LLVM).
I'd argue as a user, the GPL cares more about your freedom, so unless you have a specific reason not to use GCC, go with that over LLVM, (yes, am aware of the exception re binaries compiled with GCC not having to themselves be GPL).
Am I crazy in worrying about this? I know initiatives like reproducible builds are supposed to help solve that kind of threat, but it's still not clear to me how it all fits together.
I'd agree with this, since GNU projects tend to have at least some devs who are in it for the ideology.
That may have been true at one point, but I don't think it is any more.
Neither. They're equally unsafe with relatively low risk of this specific attack outside maybe distribution. Although clever idea, AsyncAwait's ideology doesn't work since spies will pose as them. Good news is the Karger compiler-compiler attack has only happened two or three times that I know of.
What's burned projects many more times and most worth worrying about are security-related compiler errors. They transform your code in a way that removes safety/security checks or adds a security problem (eg timing channel). So, the real problem is compiler correctness more than anything. That requires so-called certifying compilers that put lots of effort into ensuring each step or pass is correct. CompCert and CakeML are probably the champions there with formal verification. You could also do rigorous testing, SQLite-style, of each aspect of the compiler on top of using a memory-safe language. If you restrict features, then the bootstrapping can be done in an interpreter written and tested the same way before being ported by hand to designed-for-readability assembly.
It didn't stop there, though. Recent work in verification camp is doing compilation designed to be secure despite using multiple, abstraction levels such as source or assembly or mixed languages. Here's a nice survey on that:
http://theory.stanford.edu/~mp/mp/Publications_files/a125-pa...
The group putting it all together the most is DeepSpec. They have both papers and nice diagrams here:
I'm not sure if the removal of safety or security checks were caused by compiler correctness issues rather than misunderstanding of language semantics / memory model of a programming language. If we take removal of a memset (to erase sensitive data) or of erroneous integer overflow checks (because of relying on undefined behavior) as an example, they are based on language / programmer error than compiler errors. These issues should be fixed at the language level so that first, a programmer can express his intentions more easily and second, that it is hard to write code which doesn't align with the programmers intention.
https://pdos.csail.mit.edu/papers/stack:sosp13.pdf
One of the main ones that can do it without undefined behavior, or at least what I was told was undefined behavior, is optimizations getting rid of "dead" code. It doesn't have to be using memset: just an assignment. That assignment would sometimes get removed because the compiler thought nothing would be done with the assigned data. I never read if that was in C specification since it seemed to be a common problem in optimizations. Here's a recent solution just in case you find it interesting:
https://www.usenix.org/system/files/conference/usenixsecurit...
"These issues should be fixed at the language level so that first, a programmer can express his intentions more easily and second, that it is hard to write code which doesn't align with the programmers intention."
Being a fan of Ada, SPARK, and Rust, I can't agree with you more. The problem is legacy code, esp useful FOSS, that isn't getting ported any time soon. The OS's, web browsers, and media players come to mind. We need ways to analyze, test, and compile them that mitigate risks. Hence, all these projects targeting things like C.
https://twitter.com/dakami/status/668888298677407744?lang=en
That being said, in the grand scheme of things any performance improvement would probably only show up on benchmarks.
That assertion isn't very realistic. Performance benchmarks comparing gcc and clang are mixed, with performance differences being marginal at best.
https://www.phoronix.com/scan.php?page=article&item=gcc9-cla...
FreeBSD since 10.x on i386/amd64, not sure about the status on other platforms.
OpenBSD since 6.1 for arm64, 6.2 for i386/amd64. This is both default for the base system, meaning kernel and userland. And also the ports tree, for compiling 3rd party packages, very few ports still depend on gcc.
And also, while not the default system compiler yet, LLVM/clang is compiled and installed on macppc/sparc64 and mips64 systems.
Android and ChromeOS are also built with Clang. I'm curious about the distinction of "first." Does anyone know the timelines here?
"Mac OS X Internals: A Systems Approach"
https://www.amazon.com/Mac-OS-Internals-Systems-Approach-ebo...
Since then, has macOS become even further away from BSD, with its own Network stack, XPC, SIP and many other architecture changes.
following the lineage and where active development happened at the time (80s/early 90s), one can easily make the case that BSD is unix.
This is cool. I wish ubuntu/debian would move towards this.
[1] https://www.debian.org/doc/packaging-manuals/python-policy/p...
That by itself would be enough to make me choose clang over a compiler that didn't have that option. Gcc deliberately doesn't support that, right?
[edit] ipad auto correct
It also adds a data-dependency (zeroing out a stack buffer depends on the length of the buffer) which is insecure.
In fact, I'd rather have a flag that clobbers values with random data, to make sure that uses of uninitialized values are caught as soon as possible.
This is absolutely the best option. That said,
> I'd rather have a flag that clobbers values with random data.
That's roughly what happens in practice as is, doesn't it? Barring the first option I'd rather have an option that fails predictably and reproducibly. An arbitrary but deterministic garbage number maybe? Like --set-uninitialized=0xdeadbeef. That might be getting too elaborate, haha.
Initializing values just because seems wasteful. That's why global and static variables are already initialized in this way.
For me it would yet make sense that Gcc allows it even if it's not a very good practice it may have its uses.
Great work btw. I've found a few actual bugs with ASan and MSan already. Not once a false positive.
* Video: https://www.youtube.com/watch?v=QinoajSKQ1k
* Slide deck (PDF): https://llvm.org/devmtg/2019-04/slides/TechTalk-Rosenkranzer...
> Python has been updated to 3.7.3, and we have successfully removed dependencies on Python 2.x from the main install image (for now, Python 2 continues to be available in the repositories for people who need legacy applications);
Besides the Python2->3 thing, are there actually any practical advantages to using a lower-level language for managing packages?
Those legacy tools never were updated for Python 3 because they had no maintainers or developers. When the distribution switched to DNF, they were able to adopt actively maintained software that replaced those that were ported to Python 3.
Would be fun to see some benchmarks.
A smaller binary and no external dependencies is more audit-able, has fewer dependencies for bootstrapping itself on a server or a container image.
When you multiply this by the total number of images, this makes a big difference.
Does that help?
arguable.
BSD and commercial Unices used PCC and PCC derivitives for much of their history; by this token, GCC is the 'other woman', and this is itself ignoring the clear differences in philosophy between MIT/BSD and GPL licening
This is probably not true (because clang is in C++ it can't be less complicated than anything), and gcc is complicated in some parts because it uses much better algorithms (the LLVM register allocator is not as good as LRA). LLVM also has some very ugly DSLs like the .td files.
But GCC's codebase does have lots of added complexity from the extremely weird GNU coding style where they want you to pretend you're writing Lisp and all commits have to update a changelog file. Plus terrible GNU software like autotools and recursive make.
https://penguindreams.org/blog/the-philosophy-of-open-source...
And most likely, given the BSD state back then, it would mean we would just keep using either commercial UNIXes, or Windows would have won the UNIX wars.
But lets get rid of Stallmann's GNU concepts.
you mean being persecuted by over-bearing commercial unices? (e.g. SVr4 and ATT)?
lets not mis-confuse the issues to our personal ends - the argument is just as valid that without BSD UNIX, stallman would also not have a system to base a clone on..
GNU attempts to redefine the existing cultural status quo of open-source software dating from the dawn of computing to its own personal ends
GCC only got manpower when Sun changed the way UNIX SDK used to be given to customers.
And who knows, maybe companies would have been more willing to dedicate manpower to HURD.
Though NetBSD is a bit behind on switching to the LLVM toolchain, here's a recent update https://blog.netbsd.org/tnf/entry/final_report_on_clang_lld
Meanwhile FreeBSD (on amd64, i386, armv6/7, aarch64) has been buildable with clang since some point in 9.x (2012-13), comes with clang only since 10.0 (01.2014), and since 12.0 the bootstrap linker on amd64/i386/armv7 is LLD (which was the case on aarch64 from the beginning iirc)
Edit: last I checked (more than two years ago), GCC was faster on more micro-benchmarks.
People promote Clang because it has a permissive license and opens the door for Google and Apple to inject proprietary crap into Linux.
Looking around just now, there's what seems like an initial experiment from someone, but it doesn't seem to have gone anywhere. :(
https://sourceware.org/ml/binutils/2017-03/msg00044.html
Asking because with the LLVM 8.0.0 release (a few months ago) it's one of the standard supported architectures.
This doesn't make sense. You can compile proprietary code with gcc without problems already. Clang doesn't enable anything new here.
e.g. the clang memcpy can be 100x faster than the gcc memcpy, when the size and alignment is known.
And gcc-9 added serious regressions on some platforms, that you need to blacklist it. gcc-10 probably not being better.
Also with FDO I get better results with GCC over Clang/LLVM, my main test subjects are rendering (Blender), archivers, encoders and emulation.
However with straight up -O2/-O3 I very often get better performance with Clang/LLVM. I haven't benchmarked on ARM though, my results may be very different there.
The existence of a distribution using Clang for building itself makes the ecosystem way stronger.
Linux distros, package repositories, etc. are basically giant compilation farms, compiling packages making sure they work well together so that you don't have to.
So switching to clang might impact their resource usage.
---
For you, the user, the performance of binaries compiled with GCC or clang is pretty much on par. Some binaries are a bit faster with clang, others are a bit faster with GCC, often in negligible ways.
If you are doing something that's very resource intensive, recompiling that software yourself tuning it to your use cases is probably going to have a much larger impact on resource usage than whether the shipped package was compiled with gcc or clang.
> LLVM/clang 8.0.1
Guess they have a time machine, as LLVM 8.0.1 isn't released.
There's an rc2 available (3 days ago), but that's not really a release:
https://github.com/llvm/llvm-project/releases
Jumping the gun a bit maybe? ;)
https://github.com/freebsd/freebsd/commit/48cf3d0825d200d26e...
Who cares about "really™ final releases"? :D
but I'm not really sure. I guess if you want a recent kernel, KDE Plasma based distro and you work with LLVM/clang.
I really don't know where it fits with Mageia/PCLinuxOS and other Mandrake descendants.
History IIRC -- Mandrake was RH Linux with KDE; Mandriva was a continuation of that which split; OpenMandriva were devs from that split that took ROSA Linux (still doing KDE4 I think) and then continued their project from that base.
1. What arch are you targeting.
2. What version of Linux are you trying to build? Mainline, stable, next?
3. What configs are you trying.
4. What version of clang are you using.
For example, pixel 2 kernel is arm64, 4.4 stable kernel, limited configs, and clang-4.
Things for the most part are pretty green with released versions of clang. There are some long tail configs or combos of the above, but it's pretty minimal and we have a good handle on them.
https://clangbuiltlinux.github.io/
X86_64 required the implementation of ASM goto, which we just shipped. You'll need to build clang from source, but the feature will be in clang-9.0 release. Other arches should build with clang-8 (technically x86_64 will build pre 4.19 kernels) but we shipped pixel 2 kernels w/ clang-4.0 so older Clang's may work depending on your target arch/tree/configs.
"On the AMD side, the Clang vs. GCC performance has reached the stage that in many instances they now deliver similar performance... But in select instances, GCC still was faster: GCC was about 2% faster on the FX-8370E system and just a hair faster on the Threadripper 2990WX but with Clang 8.0 and GCC 9.0 coming just shy of their stable predecessors. These new compiler releases didn't offer any breakthrough performance changes overall for the AMD Bulldozer to Zen processors benchmarked.
On the Intel side, the Core i5 2500K interestingly had slightly better performance on Clang over GCC. With Haswell and Ivy Bridge era systems the GCC vs. Clang performance was the same. With the newer Intel CPUs like the Xeon Silver 4108, Core i7 8700K, and Core i9 7980XE, these newer Intel CPUs were siding with the GCC 8/9 compilers over Clang for a few percent better performance."