GCC 7 Release Series – Changes, New Features, and Fixes
gcc.gnu.org
gcc.gnu.org
What does that even mean? I'm dabbling with Ada as of recently.
Also, woo!
Support for the RISC-V instruction set has been added.
On topic, I'm using GCC5 and 6 as well as experimenting with 7, and I also have Clang - all available to me via switch/simple shell functions. I always tend to have my junk work on all of those, so am in a kind of a simplistic position to compare each. From a daily user perspective, there's not much difference at all (clang and newer gcc). Clang's pretty warnings are almost in GCC now as well, speed is still on GCC's side (for me at least) and all is good in GCC and LLVM land. I'm leaning more towards GCC since I've been using it since it became united in late 90s. I believed in it, even though people thought of it, at the time, to be slow and shit. Look at it now! I'm using it on MacOS (yes), Windows and Linux. I'm sure more heavyweight users will find a lot of things GCC can improve upon though.
I know it might not be strictly in the scope of GCC's domain, but I would like to see some static analysis and linters (like MISRA, etc.) included out of the box with it. That would be swell.
I know nothing about Ada but there is a feature called called 'executable bit' in pretty much all desktop processors since 1-2 decades. What this bit does is it marks certain regions of memory as not executable, i.e. you're not allowed to put the instruction pointer onto that region and let the CPU run the instructions. This is a hardware-level feature and can be disabled in BIOS/UEFI. Usually the stack is marked as not executable as a security feature, so buffer overruns can't just write code on the stack and then it is executed.
Why Ada needed this disabled is beyond me. Does it execute things on the stack?
I believe this is due to nested function support. There is some explanation here: https://stackoverflow.com/questions/34982151/executable-ada-...
If the error message of one is not very illuminating I often try the other compiler on the same piece of code. It is a good practice anyway and I should be doing more of that.
One claim that everyone will stand by is that the competition between the two has been a huge help.
It actually is a good idea to regularly build and test C/C++ codebases with both gcc and clang because not only error diagnostics are different but also warnings and optimizations, including crazy optimizations exploiting undefined behavior.
I discovered several bugs simply by compiling some code with different compilers and running the test suite each time.
foo(bar(), bar());
(where bar() was stateful)Discovering that g++ and clang++ compiled to one that did what I wanted and one that didn't was an interesting experience.
And before somebody wonders about currying and eager evaluation, the problem is not functions but type constructors, which aren't curried.
void test(a, b, c)
stack[0] == a
stack[1] == b
stack[2] == c
But to get the items in the stack in that order you have to... puch c
push b
push a
If you wanted to do that, and still allow for left to right evaluation you could... temp_a = a
temp_b = b
temp_c = c
push c
puch b
push a
That would be 2x instructions for every (already expensive at that time) function call. Defining unordered evaluation as UB was to avoid having to do temp_a, temp_b, temp_c. Now compiles can do that and trivially optimize that sort of thing out.Instead, the stackframe is created in one go with a single change to the stack pointer and then the values are filled in in whatever order the compiler / language desires.
Here is a typical example using GCC:
movq %rsp, %rbp
subq $16, %rsp
movl $12, -4(%rbp)
addl $1, -4(%rbp)
movl -4(%rbp), %eax
movl $14, %edx
movl %eax, %esi
movl $12, %edi
call f
Which moves the parameters using registers (subject to availability) as a further optimization. The 'subq' reserves space for the stackframe with addresses relative to rbp being the parameters.This is the accepted proposal:
http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2016/p014...
Although I have a long gcc bias I do use clang first as it is still faster to compile. Gcc still generates faster code for our codebase (and more importantly: faster where it matters to us) but they are pretty close and if I really had to decide between one and the other it would hardly matter.
G++-7-not-quite-released appears to have better C++17 support. We may be the only people who care yet :-)
In terms of exploit mitigation clang now has control flow integrity and safestack, as far as I'm aware nothing like this is available in gcc.
msan may not be such a big deal, because you could use clang for testing and gcc for production. But I hope more people adopt stronger exploit mitigations like CFI, and if gcc doesn't deliver them clang will win - at least for security sensitive areas.
asan, ubsan, tsan, and msan are important debugging features, but for my every day use in core analysis, local variables in gdb are very important.
Clang: https://godbolt.org/g/BVfy81
GCC: https://godbolt.org/g/uIN1gG
See the famous and now 12 years old GCC bug: https://gcc.gnu.org/bugzilla/show_bug.cgi?id=18501
Many warnings are available in both compilers and effort has been made to use the same names. Both compilers have some warnings exclusively. Clang prefers to raise warnings in the frontend only. Gcc in contrast will sometimes raise more warnings at higher optimization levels. If you can afford the overhead build with clang dbg, gcc O3 and MSVC.
(I'd say that this warning could be explicitly disabled for the scope is the macro... but alas, GCC broke _Pragma inside macros.)
It attempts to move evaluation of expressions executed on all paths to the function *exit* as early as possible, which helps primarily for code size, but can be useful for speed of generated code as well. [Emphasis added.]
Typo? PRE hoists upwards, so it would make sense to move closer to entry, not exit.This is identical to what LLVM's load-store motion does for loads, except for all expressions and not just loads.
This is identical to what LLVM's load-store motion does for
loads, except for all expressions and not just loads.
It's actually not.
It's what GVNHoist does, but not MLSM.
MLSM only handles diamonds.In any case, because the GCC implementation is written on top of a sane PRE infrastructure, it is like 50 lines of code to do this :)
It's actually not. It's what GVNHoist does, but not MLSM. MLSM only handles diamonds.
Fine, it's what MLSM aspires to: http://llvm-cs.pcc.me.uk/lib/Transforms/Scalar/MergedLoadSto... (:Maybe it happens, but it ain't gonna happen with anything like this code :)
But not because of "That is, if a function terminates early, you're not evaluating the expressions unnecessarily, no?"
If the function terminates early, and you've moved the expression computation before or after an early termination point, you've by definition changed what paths it is computed on (unless it was already computed there). That is not legal in all cases (it's speculative PRE/PDE)
PRE and VBE (which is what this is) guarantee that the expression is still computed at on the same paths. They just make it so it's computed once.
What is happening here is really a size optimization. If it can prove that it is always executed, it has one copy of the computation, instead of multiple ones.
IE
if (a)
foo = a + b
else b
foo = a + b
->
foo = a+b
if (a)
else (b) ->
if (a)
else (b)
foo = a+b if (int_value) { ... }
Will now be warned/failed. -Wint-in-bool-context
Warn for suspicious use of integer values where boolean values are expected, such as conditional expressions (?:) using non-boolean integer constants in boolean context, like if (a <= b ? 2 : 3)
Or left shifting of signed integers in boolean context, like for (a = 0; 1 << a; a++);
Likewise for all kinds of multiplications regardless of the data type. This warning is enabled by `-Wall`.[1]: https://gcc.gnu.org/onlinedocs/gcc/Warning-Options.html
if (int_value != 0) { ... }
To be clear about what I mean.Per the examples in the documentation, it does trigger when you do something silly like "if (i < 3 ? 7 : 0)".
In fact if you read the documentation and the warning messages more closely, it's looking for suspicious use of integer constants, not integer variables.
> Escape analysis is available for experimental use via the -fgo-optimize-allocs option.
I've been seeing this comment for the past 7-6 years, and GCC still delivers better performance on the majority of real-world programs I benchmark (mostly compression).
What I find very strange though is that some people seemingly want one of these projects to die.
As is evident from many of the replies in this thread, having two great compiler toolchains with which to test your code is a great advantage.
Secondly, looking at how GCC development picked up greatly when Clang/LLVM came on stage, it shows that GCC was stagnating with the lack of direct competition, should one of them disappear now, the same thing is likely to happen to the surviving project.
On the contrary, I would prefer having even more competition in this field.
I don't think people want, I think people are worried that this will happen. It's pretty clear that commercial backing largely favours LLVM for obvious reasons. A compiler monoculture nobody really wants back.
I'm not sure which side you were addressing here, so I'll cover both
Stallman wants the LLVM project to die for political reasons (he described it as "a terrible setback for our community" [1]). His argument is basically that LLVM can be used by non-free software, so it's mere existence is negative for the world because it enables non-free software. Also obviously it takes away resources that could have gone to improving GCC (although in my opinion a lot of them wouldn't for the reason below). It's an extreme argument, but it's the kind of extreme position Stallman has consistently taken so it's not surprising.
On the other side, one of the problems people have with GCC is that it's run by people who actively want to make worse software for political reasons. i.e. they'd rather software not support something at all if supporting it might benefit non-free software. That's fair enough, but it shouldn't be a surprise when users of the software prefer to use and support a project that isn't deliberately designed to make doing certain tasks very difficult. Academics and other people with an interest in hacking on compilers were obviously going to prefer a project that wasn't architected to try and prevent the very kinds of things they were doing.
The opposition to refactoring tools for emacs based on GCC is the most recent (2015) example of this[2], but the problem is a long standing one. Fundamentally people want to use compilers to do more sophisticated things with their code than just compile it, and that is seen as being incompatible with the political goals of the GCC project.
[1] https://gcc.gnu.org/ml/gcc/2014-01/msg00247.html
[2] https://lists.gnu.org/archive/html/emacs-devel/2015-02/msg00... and the rest of that thread
Well, for an end user those 'political' reasons are often practical benefits.
Having features only available as proprietary add-ons or through proprietary forks, or being locked out of running the code of your choice on hardware you've bought, are things I find very unappealing.
Of course there are downsides with copyleft as well, because there is no perfect solution to this problem.
Personally I'm favoring permissive licensing for projects where there is little or no incentive for commercial proprietary forks, and copyleft for projects where there is (typically end user targeted).
I wish they stop playing the political stand here since LLVM is being adopted more and more.
The free software community is more than ever before at risk of being replaced by a monopoly culture controlled by large corporations.
I wish people, particularly devs, thought more about licenses and their long term impact on freedom of the user & dev. I get tired of hearing about how "but copyleft is less free because it restricts me", but to me that's like saying "individual liberty under the rule of law is less free because it prevents me from punching that dude in the face". It's some strange form of anarchism argumentation that fails to respect the rights of others.
With all the security issues cropping up lately, I think it should be obvious to big picture thinkers that, while not the solution in itself, any real forward thinking solutions for cyber-security must focus on keeping black boxes out of the picture. BSD style licenses are dangerous to me because they allow hard working peoples code to be abused and used for abuse of others.
To be fair to this particular argument though, LLVM does fall under the LSCA license which is gpl compatible, it simple isn't copyleft, so my above rant is more a general comment than on the topic of clang.
If everything in computer software is copylefted, the status quo in the rest of the non-software economy persists.
Further, examples of AGPL 'free software' for the web backed by the dominant cloud service provider essentially give them an unlimited monopoly on that particular service, since, as copyright holder, they will be the only party able to create a proprietary fork which is better than the competition..
More philosophically:
The spirit of the law is always greater than the law itself..
If every free software project was copylefted, it would be effectively impossible for any company that is smaller than Apple to maintain everything they need to make software proprietary. Want font rendering? Reimplement libharfbuzz. Want to write any C program? Reimplement glibc. And so on.
> Further, examples of AGPL 'free software' for the web backed by the dominant cloud service provider essentially give them an unlimited monopoly on that particular service, since, as copyright holder, they will be the only party able to create a proprietary fork which is better than the competition..
Only if they have a CLA or don't accept any contributions. I can't think of any examples of such AGPL projects, but even with the GPL such projects are quite rare (and CLAs like the FSF actually don't allow them to create proprietary forks). You're just spreading FUD, please stop.
Why does it need to be proprietary in the first place though. I feel like the arguments I am hearing are making assumptions about the desired outcome to support their reasoning.
The purpose of the GPL was to encourage companies to publish their software under the GPL (want to use readline, got to use GPL). The long term goal was to create an environment where GPL software was so much better than the alternative that closed-source software would just whither.
This made sense in the 1980s. Remember that gnu was started when RMS found that he couldn't modify a program he needed - at the dawn of proprietary software.
This was a time when software writing was a small-scale operation (emacs was written by one? individual, Unix by two or three, etc.) and college students/professors could easily outnumber commercial software houses.
It actually worked for a while - gnu actually won a objective-C compiler purely due to the GPL.
Now, on the other hand, nowadays software companies are huge and have huge teams. Even without llvm's academic base, Apple would have enough cash to build it on their own.
Why make it easier for them? Very few companies can be like Apple (and even Apple didn't write everything from scratch). Compromising allows more companies to wrong their users, it's not helpful to treat companies like people -- without a profit motive most companies won't liberate their software.
Additionally, a lot of GPL software, like MySQL and BerkeleyDB, create a more closed community because that company has a lot more ability to create their own black box projects through dual licensing as closed source works. Postgres, in comparison, is BSD licensed, making it much harder for a company like Oracle to buy out pieces of the community and run away with the source and make deals others can't. The GPL has a lot less community power to counter those sorts of situations.
[1] http://llvm.org/devmtg/2013-11/slides/Robinson-PS4Toolchain....
In contrast, Linux does not and GCC's FSF copyright assignment includes a promise that the software will remain free.
With BSD-style licenses, everyone is also on an equal footing, except that now users have no guarantees that the software they use will be maintained as free software. At any point, a treacherous developer or company could scoop up the talent from the community and make a maintained fork of the project proprietary. The original project dies because of lack of talent and now free software has helped expand the reach of proprietary software.
Note that even with copyright assignments, there are some good ones. The FSFs (optional for projects) copyright assignment has specific wording that guarantees they will always keep the code free (and even go further to state that it will always be copyleft and in keeping with their well-documented philosophy). If I had a single foundation I had to pick to assign my copyright to, it would be the FSF.
Going GPL also has drawbacks. Look at how Apache 2 and GPL 2 weren't compatible because of the Apache license's patent clause. That lock-in effect of the GPL means that your software package can't be used by the larger free software community. Which if you're satisfied with it, fine, but some people don't want to encumber their software as such. There's no such thing as a universal benefit.
I don't understand what this sentence means? Do you mean they take the developers and stop them from working on the original GPL version? In this context I'm referring to the most common case which is a GPL project that has more than one copyright holder.
In _that_ context is is not possible for a developer to take the existing work of the developers, make a proprietary fork, and convince the developers (in a moment of weakness) to switch and start working on the proprietary code. They can create a new project, but that's always true and not possible to restrict (nor would anyone want to).
> That lock-in effect of the GPL means that your software package can't be used by the larger free software community.
And this is the whole point of the "or any later version" clause. Every complaint you're bringing up has already been resolved by how the GPL is used and has worked for >20 years. Of course the GPL does have its downsides (the whole MPLv1/CDDL thing is a real shame) but "lock-in" is not one of them (unless you explicitly decide to lock yourself in, which is your own fault).
I'm referring to the GPL's "no other restrictions" clauses that keeps things from interoperating with the Eclipse public license for IBM wanting to maintain choice of venue, the MPL, CDDL, 4-clause BSD, etc. Each of which the FSF says, "well, if they just used our license." The GPL is absolutely a barrier to cooperation.
Where can I download the source code for the specific version of clang used in OSX?
(AFAIK, it is unavailable, which has been a bummer in chasing bugs which show up only in it)
https://opensource.apple.com/release/developer-tools-81.html for example.
Sometimes takes a while to show the latest version though.
As a dumb example, the Open Watcom license is an "open source" license but is not a free software license because it is too restrictive. In the Open Watcom instance, the OSD does not protect the freedom for users to have private copies of software, the four freedoms do (or at least the modern interpretation does).
There's also the whole TiVo thing, where the modern interpretation of freedom 0 is restricted by DRM while "open source" software does not have any problems with DRM restricting users.
[1]: https://www.gnu.org/philosophy/open-source-misses-the-point....
In addition to GCC/clang and Firefox/Chrome already mentioned, think Microsoft/Apple, Intel/AMD, nvidia/AMD, Boeing/Airbus, Playstation/XBOX...