My favorite C compiler flags during development
nullprogram.com
nullprogram.com
-Wvla
Warn about any use of stack-allocated variable length arrays. There are a couple of reasons why these are bad: without care they can cause exploits where a caller specifies a very large array which crashes the program. And they are a problem where stack space is limited (kernel code, threads on 32 bit machines). -Wframe-larger-than=5000 -Wstack-usage=10000
Prevent large stack frames. Same reasons as above. -Wshadow
Prevent shadowed variables, eg. a local variable with the same name as a global. -Werror
Turn warnings into errors. Only use this for developer builds though, because newer compilers introduce new warnings and that will cause problems for downstream packagers. -fno-omit-frame-pointer
https://rwmj.wordpress.com/2023/02/14/frame-pointers-vs-dwar... -fvisibility=hidden
When building a library, make sure all symbols between compilation units but within the library are not exported. Basically reduces the load-time overhead and prevents random ELF symbol interposition. If you use this, then to export a symbol you have to use '__attribute__((__visibility__("default")))'. -fno-strict-overflow -Wno-strict-overflow
(Controversially) unbreaks the C compiler by allowing signed overflows to do the obvious thing.agreed on 64bit. Also this advise is old it hasn't popped up out of nowhere. I still recall it being crucial back in the early 2000's to improve performance. And today it remains very valid because of 32 bit IoT systems.
var-tracking-assignments in on by default when you use any variant of -g (which you should). You can also configure the targeted dwarf version there (which by default is 5 these days) and the level of debugging info, which you should set to the max. This makes your dwarf2-related options redundant.
-funwind-tables is also there by default in C++, not sure for C.
-rdynamic exports global symbols from the main binary so that they're available to satisfy dependencies of dynamically linked libraries and modules. For example, a Perl or Lua interpreter, or an Apache daemon, will use -rdynamic so that its symbols can satisfy the dependencies of binary modules, rather than requiring the engine implementation to be linked via a shared library itself, which is more brittle. It's sort of midway between static and dynamic linking in terms of ensuring everybody is using the same implementation of a shared framework.
Looking at it from an attacker’s point of view, a VLA is a primitive for adding an arbitrary offset to the stack pointer: extremely powerful and dangerous.
You can’t use VLAs for large allocations because the stack has a limited amount of space and you don’t know what that limit is. So you must check the size is reasonable before declaring a VLA. So you might as well declare a fixed-size array that is large enough for all reasonable uses.
Yes, one should not let an attacker control the size of the VLA. But in any case, one should use -fstack-clash-protection .
With stack clash protection if an attacker can control the size of VLA and it becomes to large, you can a trap with is likely DoS. With a fixed size array which you overflow instead, it is more likely a RCE. If you check the size, it does not matter.
That you can't use the stack for large allocations has nothing to with VLAs and also depends on what you call large and how large your stack is.
The people at WG14, Google and Microsoft (VLAs are one of few C99 features not ever coming to MSVC) think otherwise.
So you're basically saying they are wrong.
I mean, I'm yet to see a definitive example where automatic* VLA was actually right tool for the job.
I know they were supposedly introduced for the numerical analysis, but I fail to see what problem it actually did solve. Yes, the syntax way neater than piecemeal allocation, but it could have been solved with functions and macros anyway. Did VLA just hit some sweet spot in performance between fixed size arrays and heap-allocated ones, which made number crunching better optimized?
* as opposed to VM types used in function parameters and for allocating multi-dimensional arrays on heap
In terms of security, there are some issues but it is largely overblown and misunderstood. A dynamic buffer is always a sensitive piece of code when dealing with data from the network. VLA were the language tool of choice and then got a bad reputation because they were involved in many CVEs. But correlation is not causation. The main real issue is that - if the attacker can control the size of the VLA -, it is possible to overflow the stack into the heap. This can be avoided by using stack clash protection which compilers support only since a couple of years. With stack protection I believe VLAs can be safer than fixed size arrays due to improved bounds checking.
I think the main gripe most programmers have with VLA is lack of control over them. Fixed size array is, nomen omen, fixed and predictable - can even be tested beforehand. malloc() at least returns NULL on fail (although memory overcommitment muddies the situation), so program can take some action. But what happens when VLA fails? Segfault? Stack crash protection if compiled with it or worse if without? None of those options is graceful from the perspective of end user.
> This can be avoided by using stack clash protection which compilers support only since a couple of years.
As you yourself say, it's been only few years since this protection made its way into compilers. But that's not the issue. The issue is that `-fstack-clash-protection` isn't part of C language, it's part of compiler. What's the incentive for the developer to use less certain feature when there are easier alternatives?
Your second point I do not understand: How is it the fault of the C language, if it is poorly implemented? And yes, if you use a poor implementation, you may need to avoid VLA. I am not blaming programmers who do this. My point is that it not inherently a bad language feature and with a good implementation it may also be better than the alternatives depending on the situation.
VLA are poorly defined feature for which some implementations offer additional protection.
And I agree, automatic VLA aren't inherently a bad feature; it's just a poor feature of C99, C11, C17 and sadly still C23*.
I have faith we will be able to finally tackle it in C2y, but until then, I'm with Eskil in opinion ISO C would be better off without them all those years.
* Thank you for your work on N3121, hopefully we will vote it in during next meeting :)
Similar to UB and many other related issues, the main issues are primarily about what compilers do or not do. I have countless examples were C compiler could easily warn about issues or be more safe, but don't. Very similar to the arguments against VLAs there are people who want to throw the complete C language away. The story is always the same: Compilers do a poor job, but the language is blamed. Then programmers blame WG14 instead of complaining to their vendor. I do not see how VLAs are any different in this regard compared to other parts of C.
And the answer can never be to go back on fundamentally good language feature (a bounded buffer - damit!) but always push towards better implementations.
* automatic VLA (i.e. stack allocated ones) will stay optional
* VM types (e.g. VLA in function parameters) will be mandatory feature again
Actually, they allow for some limited size checking in function parameters [0], but GCC added this feature only very recently
[0]: [link redacted]
https://godbolt.org/z/q9qsax7qY
(also compiler support is still improving)
It basically shows that if you use automatic VLA you have can do some checks on VLA.
But if I prohibit automatic VLA altogether and use fixed size array, I don't need to worry about that at all and the example falls back to what I presented with UBSan doing just its regular thing.
You may want to rethink this statement
Unfortunately, it warns about all VLA, not only stack allocated ones.
What you might want is rather:
-Wvla-larger-than=0
This flag will indeed report only automatic VLA (but for some reason only if optimizations are turned on [1]).GCC now has this elaborate (kind of beautiful) way of highlighting a path through your code which leads to a potential oopsie like use after free.
`-fno-omit-frame-pointer` has been useful for profiling using perf.
if (4 < x) {
if (x < y) {
if (y == 0) {
// Unreachable since the above three conditions cannot
// all be true at the same time (without multithreading).
}
}
}
For a more detailed explanation, see [2]. (Also the inspiration for the above example.)[1] https://en.m.wikipedia.org/wiki/Transitive_relation
[2] https://github.com/gcc-mirror/gcc/commit/50ddbd0282e06614b29...
If you're on Mac with Xcode, the static analyzer and the sanitizers ASAN, UBSAN and TSAN are simple UI checkboxes.
scan-build makeThat being said, codechecker is not that great in terms of it's interface. Over all, it really needs to be implemented in a GUI application as there are many many fine tuning flags you can set with the clang analyzer (as well as potential for parallelization) that either cant be taken advantage of, or is too cumbersome to use codechecker to do.
https://codechecker.readthedocs.io/en/latest/analyzer/user_g...
I noticed that codechecker can also ingest results from many different tools other than Clang:
https://codechecker.readthedocs.io/en/latest/supported_code_...
The article mentions the address sanitizer keeps track of allocations and their sizes. I wrote my own allocator, how do I integrate it with the sanitizer? I added malloc attributes to all allocation functions but they don't seem to do anything.
In those functions, you need to implement the sanitizer yourself, and have it panic if it detects an invalid access (address sanitizers are usually implemented with shadow memory). You'll have to give your sanitizer enough shadow memory during initialization, and also add your own alloc/free sanitizer callbacks to your allocator (so the sanitizer knows when things are allocated/freed).
I have an example [1] that implements a basic ASAN and UBSAN in a kernel (it's written in D, but could be easily adapted to C or C++). Hopefully it's helpful!
[1]: https://github.com/zyedidia/multiplix/blob/master/kernel/san...
https://mcuoneclipse.com/2021/05/31/finding-memory-bugs-with...
https://interrupt.memfault.com/blog/ubsan-trap
I've tested the UBSan, it worked with some limitations but helped me catch some old hidden bugs.
clang --analyze $(SRC_WITHOUT_SQLITE) $(shell cat compile_flags.txt | tr '\n' ' ') -I$(shell pg_config --includedir) -Xanalyzer -analyzer-output=text -Xanalyzer -analyzer-checker=core,deadcode,nullability,optin,osx,security,unix,valist -Xanalyzer -analyzer-disable-checker -Xanalyzer security.insecureAPI.DeprecatedOrUnsafeBufferHandlingSetting the compiler to treat floating-point literals as floats instead of doubles helps too, but I still add the trailing ‘f’ just to be on the safe side.
After testing and fuzzing with -fsanitize=integer for a while, I came to the conclusion that implicit casts are a feature and C even has an advantage there compared to other languages which require explicit casts. To test for unwanted truncation, you really need two different casting operations. One which silently truncates in the rare cases where you really need it and another one which warns about truncation in test builds.
So my recommendation is: Remove all explicit casts and warnings like -Wconversion. Test with -fsanitize=integer. If UBSan detects truncation, you can often adjust some types so that no truncation happens. Only use explicit casts if you must convert between signed and unsigned chars for example.
I find I'm increasingly annoyed with the advice to 'just add a cast' to fix conversion issues.
And safely casting between signed and unsigned is something that should have been added to C about 40 years ago.
This right there.
https://substack.com/profile/135747695-nobody-has-time-for-p...
Also the problem goes both way. I often see people trying to learn React with little grasp of JS, or Django like it's a language, and not something based on Python.
I work on networking. People always want to learn "packet capture analysis" because Wireshark has a cool name and looks impressive.
imo there is no such skill as packet capture analysis.
Once one learns the OSI model and a bunch of Ethernet and TCP/IP implementation behaviour, then what Wireshark presents becomes obvious and supported by that contextual knowledge.
Wrong and outdated!
Apparently, in India many schools still didn't update the curriculum and teach Turbo C!
Here’s the GitHub actions for those platforms and a link to the Makefile:
GitHub Action:
https://github.com/williamcotton/express-c/blob/master/.gith...
Makefile:
https://github.com/williamcotton/express-c/blob/f2e1dde2f5a7...
I guess would be acceptible in CI pipeline too, but never tried.
I get that sometimes there are circumstances where the warning is a false positive, as the article mentions, so these escape hatches are useful. It's not the case with my former co-worker who seemed to reflexively use these escape hatches whenever he didn't understand why the compiler didn't like his code. It didn't help that he was dependent on his IDE to hold his hand. I suspect today he'd just let one of the LLMs write his code for him.
https://www.gnu.org/software/gnulib/manual/html_node/manywar...
That's used in dev mode for various projects like GNU coreutils etc.
When distros etc. are building though these warnings are not used. It's worth stating as I've seen this issue many times, that -Werror at least should not be used by default for third parties building your code
https://gcc.gnu.org/onlinedocs/gcc/Standards.html
> The default, if no C language dialect options are given, is `-std=gnu17`.
I strongly prefer clang‘s
-Weverything -Wno-c++98compat -Wno-c++98-compat-pedantic -Wno-padded
to gcc‘s -Wall -Wextra -Wyes-really-every-single-one -Walso-this -Wjust-turn-it-all-on-please -Walso-this-obscure -Wetc -Ware-you-getting-itFor instance, a lot of projects I've worked on prefer to be ready for the day if/when the code has to be compiled with a different compiler, so -Wpedantic is great.
Updated for gcc-13.1.0: https://gist.github.com/zvr/3814cc3e5b2b54e90320c09e9d074e64
How do you usually define different sets of flags (for debugging, for production...) in your Makefile?
Half jokingly, choosing build system for a project is a very important step and one of things that kills desire to do C or C++ projects on the side. The article is a very friendly description of at least a decade of experience.
Debugging build systems that do a lot of things ( unity files, caching, multiplatform and distributed builds ) has become an art in itself. Personally I prefer to capture commandlines that are passed to compiler as result to make sure that options are correct.
In continuation of my AAA project example. For a standalone utility that targets only one Linux platform and will never need to consume libraries from AAA project a makefile is still acceptable.
Its quite simple. Games is what I do for living for quite some time. They are a decent specific example of a software project that employs lots of people who need to build software every day on a job. This way I can say what industry does instead of what I think it should do.
We do, not every project is of scope of AAA game.
Heck, I don't even know if some of our projects could be converted from hand written Makefiles to something else.
Usually what can change between, say, dev and release are optimization flags and stripping symbols.
You can also do
DEBUG ?= 1
ifeq ($(DEBUG), 1)
CFLAGS =<whatever you want>
else
CFLAGS=<whatever you want>
endif
and then build with make DEBUG=0
for the release build.In BSD make, you can use something like
.ifmake target
CFLAGS := <whatever you want>
.else
CFLAGS := <whatever you want>
.endif #define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
int main(void)
{
char *x = strdup("hello world");
printf("%s\n", x);
free(x);
return 0;
}That might sound drastic but that in fact enforces discipline and the removal of warnings and is IMHO very good practice especially if you start doing that from day one.
For legacy projects this can be a goal after an comprehensive phase of warnings removal.
I've worked on projects with and without "all warnings as errors" and ultimately it is well worth it.
it depends on the promises the maintainer makes to themselves and to their users. I'd rather break the build and get immediately alerted by (hopefully moderately) upset downstream packagers than having this warning become visible in my users builds.
It isn't just about getting rid of a warning, but the new compiler has changed some important logic so this warning might highlight potentially undefined behavior. Am I, the maintainer OK with that - that is the main question I believe.
Might be useful to do regression with any new compiler being released especially when a product is used in safety or security critical systems.
-std=c17 -lm
And let downstream add their own flags (-Owhatever, -fsomething,...) on top of that.Do you know of a way to get rpmbuild to keep the -ffat-lto-objects in the ar libs?
The only way to ensure discipline and rid a large code base of warnings is to use -Werror
It may sound drastic until you're the poor intern that needs to refactor an active codebase with thousands of warning. Other than the technical debt it creates you're setting yourself up for having serious issue slipping through in all the noise.
"But some warnings might be false positives" ?
If one really must override the behavior it should be done explicitly and with a strong justification. This is possible either with a pragma[1]
#pragma GCC diagnostic ignored "-Wfoo"
or by globally[2] removing the entire warning category. -Wno-foo
[1] http://gcc.gnu.org/onlinedocs/gcc/Diagnostic-Pragmas.html[2] https://gcc.gnu.org/onlinedocs/gcc-4.1.2/gcc/Warning-Options...
And this works as long as everybody has the same setup as you. This is unfortunately never the case. Kernel changes, libc changes, the compiler changes and sooner or later you will break the build.
Many replies seem to imply open source by default. Of course projects may have specific requirements, and open source is indeed a specific use case, but the majority of software is commercial, closed source.
Regarding compiler versions, many projects I worked on actually check that during the build because the code may only have been designed for or tested against a specific compiler/compiler version. In general upgrading the version of the compiler is an activity in itself because, indeed, this may require solving new warnings/compilation issues and a full test campaign.
That may be true for 'end user applications', but definitely not for libraries (thankfully). I use exactly zero closed source (precompiled) dependencies in my hobby and professional work because they are such pain to work with. Thankfully these days there's an open source alternative for pretty much everything, 15 years ago the situation was much more dire.
Edit: the applications you use on your desktop represent a very small portion of the software running in the world, including software that crosses your path on a daily basis and so 'software devlopment' is not synonymous with open source.
But when I look at the tools and applications I use daily, there's only two important applications that are closed source: VSCode and Chrome (and even this is debatable because both of those tools are just thin wrappers around Chromium: https://chromium.googlesource.com/chromium/src.git).
That doesn't prevent the compiler from expanding an existing warning, but that seems to break code less often than a blanket `-Werror`.
Hard disagree. No warnings means CI should fail (or otherwise scream loudly), not every build of your in-progress working tree. You can add a commit hook if you want but if it needs to be enforced then it isn't really discipline.
Even for CI you probably don't want -Werror but instead run the build completely so that you can track all warnings.
And for source released to others -Werror is downright rude as you have no idea what warnings their possible future compiler will have.
Many of the suggested warnings will reject valid-but-marginal code that you might not have time to rewrite (or might not be able to, eg. it's in an external library that you don't understand well).
Sanitisers and static analysers very often complain about quite valid code that doesn't need to be fixed at all.
Some hardening flags cause significant performance regressions.
Who is to say that -g3 is the best option? It causes huge binaries to be generated which might not be useful if you have other means to debug the code or can't/don't use gdb.
It's just that some people never bother to read the docs.
do you maybe want to consider taking a step back and re-evaluating your opinions?
https://gcc.gnu.org/onlinedocs/gcc/Warning-Options.html#inde...
https://gcc.gnu.org/onlinedocs/gcc/Warning-Options.html#inde...