GCC should warn about 2^16 and 2^32 and 2^64
gcc.gnu.org
gcc.gnu.org
/edit: had that mixed up in my head. Interestingly, clang does warn about the overflow:
main.c:6:17: warning: implicit conversion from 'int'
to 'int16_t' (aka 'short') changes value from
65536 to 0 [-Wconstant-conversion]
int16_t s16 = 65536; -2^16 = (-2) ^ 16 = 0b111...1110 ^ 0b10000 = 0b111...11011110 = -18
if I'm not mistaken.But since philosophically a signed int has '1's in every bit from the MSB to infinity towards the left, it should thus be -0 not 0.
I code in C/C++ too, although not exclusively. I facepalmed a few seconds later, but it goes to show the value of having this be a compiler warning.
Does depend on what kind of lint you want to catch but something like this 'just' needs some lexing and parsing to give a usable AST that's reasonably clean so it's not too hard to write lint rules against it. (Not trivial for C & C++ hence 'just')
Building the full compiler is much more work, especially if you want a decent optimizing compiler.
In fact most of the essential complexity is in the front end; you make do with a simple back end, but the front end is irreducibly complex if you want it to actually compile the language to spec.
And you need the front end to lint, especially if you want to avoid false positives, the bane of linting tools.
[1] http://blog.reverberate.org/2013/08/parsing-c-is-literally-u...
A test passing just means that the code and the test agree, not necessarily that both are correct.
I'm pretty sure Dennis Ritchie at some point himself said that not including a linter was a mistake, but I cannot find a quote of this with a quick search.
In defense: configuring a package that is compiled by lots of people on lots of machines with lots of versions of various compilers requires a lot of attention to warning flags and such. If you aren't starting out a package from a blank buffer it's very hard to get this right.
But I am frequently shocked by the number of compiler warnings I get from code downloaded from public repos and compiled in precisely the environments it's been documented to have been tested on.
"To encourage people to pay more attention to the official language rules, to detect legal but suspicious constructions, and to help find interface mismatches undetectable with simple mechanisms for separate compilation, Steve Johnson adapted his pcc compiler to produce lint [Johnson 79b], which scanned a set of files and remarked on dubious constructions."
Some may even have released code that still has these time bombs waiting to occur.
Every external tool is a lost opportunity for quality enforcement.
https://docs.microsoft.com/en-us/visualstudio/code-quality/u...
Just making the point that they have been largely ignored.
Ever heard of PC Lint from Gimpel? It was my first one.
Or do you think I would be critic of how many use C, without adopting best practices myself?
One learns a lot about code quality, or lack thereof, when doing enterprise consulting.
Pure truth, there. One has seen neither really awful code nor really great code until they have seen the wide array of code quality produced by an enterprise with a few decades of software development history. Any enterprise will suffice, it doesn't have to be a huge corporate entity.
There are a very few developers who are really good and who really care, and there are very few who really do not care and who are really not good, and most are somewhere in between.
Check "Which of the following tools do you or your team use for guideline enforcement or other code quality/analysis?" results.
A rather unfortunate operator choice on the part of C: open to confusion and precluded adding ^ as an exponentiation operator to the language.
Upon first reading, I'd assume they are seeing the maximum unsigned values, but their variables aren't declared as unsigned -- and C/C++ default to signed values. Secondly, representing the nth power of two actually takes n+1 binary digits, so e.g. 28 is a one, followed by 8 zeros -- overflowing any 8-bit variable, signed or unsigned! This is undefined behavior as-declared (signed values), but would be entered as a zero for explicitly declared unsigned variables.
Personally, if I were to be implementing this, I'd use bitwise operations to generate the correct values, and explicitly acknowledge in a comment something like "the following code knowingly breaks the abstraction of numeric values. It is likely not portable to other hardware, nor should you assume easy extension to other variable types. Das Quellcode ist nicht fur der gefingerpoken!"
XOR may be rarely used, but 2^8 is not only one of the most common accidental XORs, it's also probably one of the more common legitimate uses of it.
Nope, the warning would be for literal numbers in the source. IMO anyone who doesn't define constants for the flag bits deserves the warning.
The preprocessor does macro expansion, the compiler compiles the result.
The compiler does not see FLAG_BIT1 ^ FLAG_BIT3, it only sees the result, e.g. 2^32.
Therefore to catch only explicit 2^32 the warning should come from the preprocessor... Nice mess created right there.
Interestingly, msvc can also do this, by virtue of not even having a distinct preprocessor phase at all.
j = 0
while condition():
++j
In Python, there is no "increment" operator, so `++j` is parsed as two unary plus operators: it's a no-op. x = 12 //10
What they where apparently trying to do was to change the value of x while commenting out the old value for reference. Except in python // is integer division and not comments so x was set to 1 instead of 12. a = b, a = ["long", "list" "of", "strings]
which become a = ["long", "listof", "strings"]
due to automatic string concatenationWe could roll -Werror for most packages (I think we have a few packages, e.g. 3rd party libs, which emit warnings), but only use error=format and error=format-extra-args.
To be fair, we disable these warnings (-Wno-...): unknown-pragmas, long-long, register, unused-command-line-argument.
That may not look relevant in short periods of time, but there are large odds that somebody will be biten by your choice of using -Werror 5 or 10 years down the road.
OTOH 5 to 10 years is a good way to accumulate plenty of warnings (assuming a lack of development discipline), which at some point makes it easy to miss that one critical warning that should really be fixed. Maybe I should mention the worst case for an error in the code I write is "people die" (static analysis of safety critical stuff), so that's a good motivation to have no warnings ;-)
And forget new versions of compilers, how many things are only tested with GCC and fail to build with Werror with clang because of different warnings in Wall?
see https://www.cpan.org/authors/id/T/TY/TYPO/typo-2.46.pl
// START: Mon Jun 17 14:49:53 2019
C:\src\test\test.c (1): using ^ on a number 37: short#maxshort=2^16;
C:\src\test\test.c (2): using ^ on a number 37: int#maxint=2^32;
C:\src\test\test.c (3): using ^ on a number 37: long#maxlong=2^64;
Link to the 1999 Perl Conference paper for the script:https://web.archive.org/web/20090729104734/http://www.geocit...
It’s uncommon to use base 10 when bit-twiddling with the XOR operator - so this makes sense.
Of course, it would be better if C-family languages didn’t use ^ for XOR in the first place.
I think at the time they were created there simply wasn't an ambiguity in anybody's mind. POW is a relatively complex operation, I don't think anybody expected it as a builtin operator (much like SQRT and friends). Arguably so is DIV but I suppose that's too convenient not to have.
XOR on the other hand is something you really want to be a builtin operator in a low level, bit-twiddling language like C, so getting a dedicated operator makes sense. I'm not sure if an other sigil would've made more sense, after all ASCII lacks ⊕.
And why would anyone in their right mind do 23 and not 1<<3? At least among the crowd who knows pow is expensive and what xor is.
bitwise x = 0x1234; x |= 0x1000;
(0x indicates a bitwise value.)
By making 0x literals of bitwise type, they can only be used with bitwise operators. Decimal literals are regular integers so can only use arithmetic operators.
If you want to use a hexadecimal value with arithmetic operators or decimal values with bitwise operators, use the operator that converts one type to another.
Warning about constant subexpressions where the LHS and RHS of ^ are constant literals makes a lot of sense to me, but if the LHS is a constant expression but not a constant literal, then not so much. Also, warning about this will cause some false positives where macro expansions are involved -- probably not a good thing.
I find it amusing that people really do that though.
On the other hand, piping and redirection are relatively uncommon outside of the shell. There is certainly a clearer mental distinction for me between "I am in shell" and "I am not in shell" than "I should double star".
YMMV or course, but distinction between a low-level language like C, and a higher-level one that can support exponentiation, matrix multiplication, thread spawning etc. as a first-class language construct is as clear to me as the difference between shell and non-shell.
X ** Y
would be parsed as X * (*Y)GCC should also change your diaper, burp you, and issue random diagnostics like
foo.c:32: warning: did you test this yet?If for whatever reason I have to include GCC in an embedded this sort of "misleading indentation" and "suggest parentheses" nonsense just wastes precious flash space.
Better to leave this to code analysis tools.
Edit:
2^32 uses the ^ operator exactly as intended, and has perfectly legitimate uses, especially when dealing with registers of HW peripherals.
It's not the same as warning on "if (a=b)" as someone replied.
From a language perspective, 2^32 is exactly the same as 2+32.
Many build environments are set to treat all warnings as errors.
The point is that 2^32 is a perfectly compliant C expression that is neither misleading nor ambiguous, and that also won't create any variable overflow. It uses ^ exactly as intended. Why should the compiler complain? Why should I get a warning/error when following the spec to the letter?
A new type of misleading and error-prone statement has been found, so it seems entirely reasonable that this is added to that list.
This seems to be the entire intent of compiler warnings (these days anyway)
Mostly written as y ^ (1 << x), which could easily resolve to these expressions. Mostly at runtime, sure, but there are exceptions. Especially if you like descriptive constants. I would expect it to be quite difficult to separate the "correct cases" without people starting to just suppress compiler warnings or trick it with writing the same stuff in different words (which might be better).
On the other hand, setting max_short to something like 18 is probably a really god prank in larger programming environments. The compiler would just ruin all the fun here.
We’re talking about a parse-time check; things that resolve to those values after identifier binding won’t throw the warning. It’s not a warning about what you wrote (semantically), it’s a warning about how you wrote it (syntactically).
Things that resolve to this don't trigger the warning, and it is a compile time warning.
And then specifically looking at 2^X and 10^X occurring in source code.
You're suggesting that using the C language as specified and intended should produce a warning... That makes no sense.
As for limiting the base, that is a more fluid matter. I guess you want to avoid warnings for very large numbers. Because those are more likely legitimate.
There are many technically valid expressions and statements that will still generate warnings in most compilers.
if(a=b) {} comes to mind first and foremost.
"if (a=b)" uses the "=" operator exactly as intended with perfectly legitimate uses. It's warned about because it's an error-prone construct, not because it's incorrect.
Seems to me 2^32 is also error-prone, if you wanted to combine flags you'd use 2|32, using xor here is weird.
On the other hand ^ expects integers so 2^32 is exactly what is expected.
Most the replies I saw here try to second-guess or claim that ^ should be used in a specific way. Not so, ^ is just doing XOR of two integers.
Apparently, I am having an incorrect opinion, though, so I will self-censor and remain silent.
The problem is, what proportion of 2^32 in C are correct? I'm will to bet it is as close to 0 as doesn't matter. The gcc devs aren't stupid, before enabling a warning like this, they will run a test compiling a substasal chunk of debian, and see if there are any false positives.
if (c)
printf("Hello");
exit(0);
exit(1);
Produces a warning in gcc: test.c:7:5: warning: this ‘if’ clause does not guard... [-Wmisleading-indentation]
if (c)
^~
test.c:9:9: note: ...this statement, but the latter is misleadingly indented as if it were guarded by the ‘if’
exit(0);
^~~~
"2^32" is of course just an xor, but it makes no sense (the "exclusive" part of it is not being exercised). If you really mean 2 bits being set in a constant then "2 | 32" is much clearer. If you're dead set on xor, "32^2" won't trigger the warning.Stating that 32^2 should not trigger a warning while 2^32 should shows that this proposal has not been thought through.
It's not desirable or sensible to raise a warning on the premise that the expression might mean something else in another language, which is what this would do.
Just because something is specced doesn't mean that it's reasonable. My "if statement" example is perfectly valid C, but it's not reasonable. 2^32 is perfectly valid, but also unreasonable. Frankly, any bitwise operations with >10 decimal constants are either intentionally obfuscating, or done by someone who doesn't know what they are doing. Consider "2^32" vs "0x02^0x20". The latter is much better.
It is entirely sensible to warn on the premise that it might mean something in another language: The construct seems to be an actual point of confusion (the authors sited numerous real code examples where people are accidentally doing this in the wild), and the construct makes no sense at face value. I think you'd be hard pressed to find a real example of 2^32 in code where it isn't a bug. That's enough to make it a reasonable warning.
error[E0308]: mismatched types
--> src/main.rs:4:8
|
4 | if x = y { println!("eq") }
| ^^^^^
| |
| expected bool, found ()
| help: try comparing for equality: `x == y`
|
= note: expected type `bool`
found type `()`
error: `<` is interpreted as a start of generic arguments for `u32`, not a comparison
--> src/main.rs:4:17
|
4 | if i as u32 < 2 {
| -------- ^ --- interpreted as generic arguments
| | |
| | not interpreted as comparison
| help: try comparing the cast value: `(i as u32)`We need to kill the myth that good programmers don't need help from their tools.
Also: it's incredibly naive to think that "bad programmers" are the only ones that would make a mistake like this. I'm very familiar with all C/C++ operators and know perfectly well that the caret represents xor, but I could easily imagine myself slipping and writing 2^16 instead of 1 << 16, because that's how the rest of the world writes exponentials.
Let he/she who has never written a typo and had the compiler save them cast the first stone.
In all honesty while I don't agree with this particular instance, the reasoning isn't ridiculous. If we assume it's true that it'd only affect bad programmers, then you probably wouldn't want to add it, since it'd increase the false positive risk and/or otherwise potentially get in the way of everyone else. (Kind of like how you wouldn't make cars start honking when they're turned on as a means of protecting against bad drivers, even thought that might well save lives.) Some tools just need to assume some base level of expertise to be effective.
That's not the reasoning given. The reasoning given is that since only bad programmers would make this mistake, it shouldn't be a warning at all, since those kinds of programmers "will create a lot of other problems regardless". That is nonsensical. You could make the same argument against ANY warning. The whole point of a warning is to alert the developers that, while their code technically follows the rules of the language, they've probably made a mistake.
Bad programmers are the ones who need the most help from our tools. Saying there’s no need for a warning because only bad programmers will benefit is like saying there’s no need for crossing guards because only stupid children get killed crossing the street.
I get that you are a full-time C programmer, but unless you're volunteering to take on all my C troubleshooting work, this seems almost punitive. Like building a balcony with no railing, because nobody would just walk off a ledge -- and yet, sadly, professionals still die from falls.
To say that all such people are "clueless", regardless of skill or training or experience, is to fall back on the old test pilot mentality: the good survive, therefore, if you didn't survive, you (retroactively) must not have had The Right Stuff after all.
Joel Spolsky explained the reasoning more eloquently than I could:
"Now, even without going through with this experiment, I can state with some confidence that some of the users will simply fail to complete the task, or will take an extraordinary amount of time doing it. I don’t mean to say that these users are stupid. Quite the contrary, they are probably highly intelligent, or maybe they are accomplished athletes, but vis-à-vis your program, they are just not applying all of their motor skills and brain cells to the usage of your program. You’re only getting about 30% of their attention, so you have to make do with a user who, from inside the computer, does not appear to be playing with a full deck."
I trust myself and the people I work with, but I still wear a hardhat when they're working overhead, and I tie all my tools to my belt when I'm working at height. It's the clueless people who don't learn from the mistakes of the past.
I can imagine a casual C coder could easily end up using ^ as "pow" for very much the same reason if they're used to it working that way in other languages. On top of that the proposed warning has extremely low chances of triggering on a false positive on actually legit code. I think it's a good proposal.
>Without to mention the kind of bugs created by such incorrect usage would be simple to spot most of the cases.
That's not obvious to me at all. Bogus bitfield values could be rather tricky to track down, especially since they might work correctly "by chance" in simple cases (for instance even if the value of the constant is completely wrong, as long as it's != 0 you can set it with |= and test it with & and it'll appear to work mostly correctly).