Integers in C quiz
acepace.net
acepace.net
Rust not having a default "int" type and forcing you to explicitly cast everything is such an improvement. Truly a poster child for "less is more". Yeah it makes the code more verbose, but at least I don't have to worry about "lmao you have a signed int overflow in there you absolute nincompoop, this is going to break with -O3 when GCC 16 releases 8 years from now, but only on ARM32 when compiled in Thumb mode!"
D follows the C integral promotion rules, with a couple crucial modifications:
1. No implicit conversions are done that throw away information - those will require an explicit cast. For example:
int i = 1999;
char c = i; // not allowed
char d = cast(char)i; // explicit cast
2. The compiler keeps track of the range of values an expression could have, and allows narrowing conversions when they can be proven to not lose information. For example: int i = 1999;
char c = i & 0xFF; // allowed
The idea is to safely avoid needing casts, in order to avoid the bugs that silently creep in with refactoring.Continuing with the notion that casts should be avoided where practical is the cast expression has its own keyword. This makes casting greppable, so the code review can find them. C casts require a C parser with lookahead to find.
One other difference: D's integer types have fixed sizes. A char is 8 bits, a short is 16, an int is 32, and a long is 64. This is based on my experience that a vast amount of C programming time is spent trying to account for the implementation-defined sizes of the integer types. As a result, D code out of the box tends to be far more portable than C.
D also defines integer math as 2's complement arithmetic. All that 1's complement stuff belongs in the dustbin of history.
Rust's current handling of this issue is by no mean perfect, although it's been steadily improving and I definitely don't use `as` as much as I used to.
>D also defines integer math as 2's complement arithmetic.
I think modern C standards does so as well, but I think than signed overflow is still UB, so it's mostly about defining signed-to-unsigned conversions. There are flags on many compilers to tell them to assume wrapping arithmetic but obviously that's not standard...
char a,b,c;
a = b + c; // error
a = cast(char)(b + c); // ok
produces a truncation error, and requires a cast. But char types have a very small range, and overflow may be unexpected. So D makes the right choice here to promote to int, and require a cast.So for instance the following code generates an error (I add to add the function call indirection otherwise rustc would even refuse to compile the straightforward `200u8 + 100u8`):
fn main() {
println!("{}", add(200, 100));
}
fn add(a: u8, b: u8) -> u8 {
a + b
}
// thread 'main' panicked at 'attempt to add with overflow', test.rs:6:5
// note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace
IMO that's a more robust and generic solution than hoping that `int` promotion will prevent the overflow. The drawback is that it incurs a performance hit and is therefore only enabled in debug builds by default.Not sure I like that D promotes implicitly and then complains about the result of the addition not being assignable to a `char`.
But in theory, every integer could just have the maximum range of values encoded in its type, and for every arithmetic operation the compiler could check if the maximum range of the integer type is exceeded.
You could set the range explicitly:
fn mult_col(color:u8<0..31>, factor:u8<0..8>): u8<0..248> { color*factor }
Or make the function generic over any range of inputs where the possible range of outputs provably fits the output type, which would be the default.There are crates that add ranged integers as types, but I haven’t found them satisfactory to work with.
D has core library functions to check for integer overflow. In my not-so-humble opinion, a targeted solution like this is better than a global switch to turn it on and off. Adding is deeply embedded into computation, and very very few are at any risk of overflowing.
That depends on how the casting is provided. For the C-style casting, or Rust's `as` casting, yes that is a problem. However, another way casting could be provided is through conversion functions that are only infallible if information isn't lost. For example, let's say we have the functions `to_u16` and `to_i16`. For an `i8` the first function could return `Option<u16>`, the second `i16`, while for a `u8` they would return `u16` and `i16`. That way, any change to the types that could cause it to now silently truncate would instead cause a compiler error because of the type mismatch.
Rust almost gets there with its `Into` and `TryInto` traits which do provide that functionality, but trying to use them in an expression causes type inference to fail, which just makes them a pain in the ass to use.
This means that if first_thing used to have a non-overlapping value space so that converting it to other_thing might fail, so you wrote
other_thing = first_thing.try_into().blahblahblah;
... if you later refactor and now first_thing is a subset of other_thing so that the conversion can never fail, the previous code still works fine, the try_into() call just never fails. In fact, the compiler even knows it can't fail, because its error type is now Infallible, a sum type with nothing in it, so the compiler can see this never happens, and optimise accordingly.Sort of like putting a nut on a grade 8 bolt rather than a hardware store bolt.
People don't always agree on what's better:
https://digitalmars.com/d/archives/digitalmars/D/Movement_ag...
and sometimes the rationale for things isn't obvious at all, and even counter-intuitive.
And hence unsigned shorts are promoted to signed ints.
C should also make char unsigned (D does). Optionally signed chars are an abomination.
The peculiar vagueness of the early C standards with respect to integers was probably due to the variety of hardware we used in the 1970s. For example:
The CDC 6000/7000s had 60-bit words, ones compliment arithmetic, 6 bit chars (but machine addresses were only of the 60 bit words).
The peripheral processors of the CDC machines (responsible for I/O) had 12-bit words as I recall.
IBMs very popular 360 mainframes supported several native hardware number formats: twos complement, unsigned, zoned decimal (one decimal digit per 8-bit byte), and packed decimal (two digits per byte).
I believe C was first developed on the DEC PDP-11. This was a very popular computer with 16 bit words and 8 bit bytes. It was twos complement, but mixed endianess (big endian for the 16 bit words making up a 32-bit long and little endian for the bytes in the 16-bit words).
Other DEC computers of the 1960s used 12 bit program counters and accumulators; I did some assembly language programming on the PDP-6 (it cost $300,000 back then and 24 were sold).
I seem to recall using a system with 18 bit registers and 9 bit bytes, but I can’t remember the details.
One system that I did assembly language programming on for a year was the TI 960. This machine was notable because it had two separate 64KB address spaces, one for data and one for instruction, known as a Harvard architecture instead of Von Neumann architecture. The advantage of the Harvard architecture was that a program couldn’t inadvertently modify the code it was running, a frequent and difficult to untangle occurrence when programming assembly. The Von Neumann architecture became dominant because it didn’t chop the available memory in two and partly because self modifying code was believed to be sometimes useful (yikes).
Unisys OS 2200 uses one's complement[1] and Unisys MCP uses signed magnitude[2]. Both are still around.
[1] - https://public.support.unisys.com/2200/docs/CP19.0/78310422-... (page 108) - "UCS C represents an integer in 36-bit ones complement form (or 72-bit ones complement form, if the long long type attribute is specified)."
[2] - https://public.support.unisys.com/aseries/docs/ClearPath-MCP... (page 304) - "ClearPath MCP C uses a signed-magnitude representation for integers instead of two’s-complement representation. Furthermore, ClearPath MCP C integers use only 40 of the 48 bits in the word: a separate sign bit and the low order 39 bits for the absolute value."
Rust gets points in my book for being more explicit and avoiding arbitrary promotion rules. Plus the standard lib has specific functions for things like add with overflow which reflect what the CPU flags are doing.
True. I've done a lot of assembler programming. Getting the right jxx instruction after a compare was a prolific source of bugs. I've had a lot fewer bugs with signed types :-)
> Plus the standard lib has specific functions for things like add with overflow which reflect what the CPU flags are doing.
D does, too:
https://dlang.org/phobos/core_checkedint.html
along with integer types that automatically check for overflow:
Rust's 'as' will silently throw away data to achieve what you asked. Sometimes you wanted that, sometimes you didn't realise, and requiring into() and try_into() instead helps fix that.
For example suppose I have a variable named beans, I forgot what type it is, but it's got the count of beans in it, there should definitely be a non-negative value because we were counting beans, and it shouldn't be more than a few thousand, so we're going to put that in a u16 variable named enough. let enough = beans as u16;
This works just fine, even though beans is i64 (ie a signed 64-bit integer). Until one day beans is -1 due to an arithmetic error elsewhere, and now enough is 65535, which is pretty surprising. It's not undefined but it is surprising.
If we instead write let enough = beans.into(); we get a compiler error, the compiler can't see any way to turn i64 into u16 without risk of loss, so this won't work. We can write beans.try_into().unwrap() and get a panic if actually beans was out of range, or we can actually write the try_into() handler properly if we realise, seeing it can fail, what we're actually dealing with here.
And of course there are times I do want to bitcast negative values to large positive unsigned values or vice versa, without error. So while I do understand why maybe you shouldn't use as, at the end of the day, it just ends up being easier to use it than not use it.
I always carry with me a tiu() alias method for exactly that reason (in a TryIntoUnwrap trait implemented for every pair of types which implements TryInto).
You might use try_into(), but then you risk panicking on a 32-bit system, instead of getting a compile-time error.
Never mix unsigned and signed integer comparison unless you know exactly what you are doing.
And never do arithmetic on boundaries definitions like INT_MAX, they are boundaries , why you need to compute a derived value on boundaries?
If you do not need arithmetic behavior, do not use signed integer, use unsigned. Because computer does not understand sign, so if you do not need a sign, do not use it.
You do not need to be a language committee member to write C , you just need to understand the reason behind its design.
The unsignedness of the integer lets readers of the code know that it's a constant of some kind, and that comparisons should be performed with care. Normally, when I've worked with C (fairly limited), unsigned numbers normally get used as flags for bitwise comparisons. Sometimes your flags may have a partial ordering defined on them (so a sense of flag A being 'smaller' than flag B), so you might need some arithmetic to do those checks you need to be explicitly careful with the boundary checks though.
I'd argue something stronger: if you care about boundaries like INT_MAX, you should never be comparing them using your regular comparison tools. I.e., even though there are correct ways to compute whether x + 1 will overflow, don't bother trying to do that and instead always use __builtin_add_overflow since you can't fuck it up. Trying to do these sorts of edge checks are incredibly hard and it has lead to numerous security vulnerabilities due to checks being optimized out. The builtins do exactly what they say and you don't have to worry about UB blowing your foot off.
That's interesting. In my career (embedded development) I've learned to do the opposite. Always use signed unless you have a reason not to. Even if a value can't naturally be negative, use signed. Use unsigned only if you need the extra bit, or if you're doing bitwise operations.
> Because computer does not understand sign
Computers understand signs just fine, we're long past the days of the 6502 N-flag being a glorified bit 7 check. All CPUs have signed instructions.
In this case, since you make sure that the extra size gained by unsigned is not important to you all, then you can also go with signed by default. Basically it is the tradeoff between 1. robust in majority of the use cases 2. capability to do signed arithmetic 3. additional positive integer range. As long as you make a consistent selection and be mindful when you are in the danger zone, it can be handled.
Can you elaborate on what benefits this approach has? I would feel that, especially when a number cannot be negative, unsigned integers seem like a proper representation of the data?
If you accidentally subtract and go into 'negative' numbers, using an unsigned in C will just happily give you the wrong answer.
Now in something like Haskell, it's much more robust advice to make invalid states unrepresentable. But C's compiler and runtime don't typically help you, when you are trying to break your representation.
I think the biggest source of these kinds of bugs is because size_t is unsigned but ptrdiff_t is signed. So the moment you mix pointer subtraction with sizeof, footguns abound.
I don't have 16 years of experience writing C, but
> If you do not need arithmetic behavior, do not use signed integer, use unsigned
does seem in a roundabout way to match the advice that I've gotten from other folks -- usually stick to signed (because you never know if somebody is going to want to use an integer in, say, a downwards loop). Your comment just seems to highlight the less usual case, where you can be sure that nobody will ever need that arithmetic behavior... maybe it depends on the type of applications, though.
(unsigned short)1 > -1
Correct answer is: implementation-defined.The left operand of > has type unsigned short, before promotion. On C implementations for today's popular, mainstream machines, short is narrower than int; therefore, whether it is signed or unsigned it goes to int.
In that common case, we are doing a 1 > -1 comparison in the int type.
However, unsigned short may be exactly as wide as int, in which case it cannot promote to int, because its values do not fit into that type. It promotes to unsigned int in that case. Both sides will go to unsigned int{*}, and so we are comparing 1 > UINT_MAX which is 0.
Maybe the author should name this "GCC integer quiz for 32 bit x86", and drop the harder choices like "implementation-defined".
---
{*} It's more nuanced here. If we have a signed and unsigned integer operand of the same rank, it could go either way. If the unsigned type has a limited range so that all its vallues are representable in the signed type, then the unsigned type goes to signed. My remark represents only the predominant situation whereby signed and unsigned types of the same rank have overlapping ranges that don't fit into each other: the unsigned version of a type not only lacks negatives, but has extra positives.
> All other things being equal, assume GCC/LLVM x86/x64 implementation-defined behaviors.
* windows - https://godbolt.org/z/WWWzarEsK
* linux - https://godbolt.org/z/57WK45o6e
same for gcc on linux/windows (mingw)
(Edit, since some respondents seem to miss this, explanations about efficiency or ISAs might justify promoting to u32 (though even that's debatable), but not i32. A design that auto-promotes an unsigned type, where every operation is nicely defined and total, into a signed type, where you run into all kinds of undefined behavior on overflow, is simply crazy.)
The unsigned preserving rules greatly increase the number of situations where unsigned int confronts signed int to yield a questionably signed result, whereas the value preserving rules minimize such confrontations. Thus, the value preserving rules were considered to be safer for the novice, or unwary, programmer. After much discussion, the Committee decided in favor of value preserving rules, despite the fact that the UNIX C compilers had evolved in the direction of unsigned preserving. uint8_t x = 4;
extern volatile uint64_t *reg;
*reg &= ~x;
In the last statement x is promoted to an int, and then when the logical NOT occurs every bit is set to 1, including the high bit. When it's converted to a uint64_t for the AND, the high bits are also set to 1. So the result is that the final statement clears only bit 2 in *reg.If it promoted to unsigned int, then it would also clear bits 32-63.
extern volatile uint64_t *reg;
*reg &= ~x;
People should stop doing this. What this means is: extern volatile uint64_t *reg;
uint64_t tmp = *reg;
tmp &= ~x;
*reg = temp;
But of course when you write that chances are somebody will point out that you're running in interruptible context sometimes in this function, so that's actually introducing a race condition. Why didn't they say so when you wrote it your way? Because that looked like a single operation and so it wasn't obvious it might get interrupted. *reg &= ~(uint64_t)x;
or, and there’s no elegant way to even write this in C, I might want: *reg &=(uint64_t)(uint8_t)~x;
The fact that I have to write two casts here to undo the damage of the auto-promotion is evidence of how broken this is. *reg &= (uint8_t)~x;
Or: *reg &= ~x & 0xFF;Here's a fun standardization problem I came across recently (nothing to do with C): http://mywiki.wooledge.org/BashFAQ/105
Now, I don't find this reasoning persuasive--it's not that hard to emulate an 8-bit or 16-bit operation--and judging from the history of post-C languages, most other language designers are equally unmoved by this reasoning, but I can see someone in their right mind designing a language that acts like this. Especially if the first architecture they're developing on is precisely such on architecture (PDP11 doesn't have byte-sized add/sub/mul/div).
IMO the real issue is not so much the fact that all shifts of any type < int is treated as if it were an int, it's that the language doesn't force you to acknowledge that in the code. If you got a compilation error when trying to shift a short and had to explicitly promote to int in order to make it through, at the very least it can't lead to an oversight from a careless programmer.
C is trying to be clever but only goes half way, resulting in the worst of both worlds IMO.
UBs are not required, but you need them if you want C to behave as a macro-assembler as well as allowing for aggressive optimizations. For instance `a << b` if b is greater than a's width is genuinely UB if you write portable code, different CPUs will do different things in this situation. Defining the behaviour means that the compiler would have to insert additional opcodes to make the behaviour identical on all platforms.
You may argue that it's still better than having UB but that's just not C's design philosophy, for better or worse.
Your seem to be mixing up implementation defined behavior and undefined behavior. It would have been perfectly reasonable to make this choice if signed integer ovwrflow were implementation-defined, but it is unfortunately not - it is undefined behavior instead. This means that a program containing this instruction is not valid C and may "legally" have any effect whatsoever.
Yeah, hence why I asked this question years ago: https://stackoverflow.com/questions/39964651/is-masking-befo...
Yes, it'd be better if you had to explicitly cast to int or unsigned to perform arithmetic, but that ship has sailed.
For a website meant to educate programmers on C language gotchas, this is a pretty lackluster effort.
Even the initial disclaimer, "assume x86/x86-64 GCC/CLang", is wrong, as the compiler does not have anything to do with integer widths.
Occasionally I'll tutor C and C++ to college students, and I swear that I spend half of my time in such cases correcting erroneous (or absurdly incomplete) info provided by instructors, Stack Exchange articles, etc. It's especially challenging when I get, "but I read it on HN?!"
>>: 2 % -5
-3
But in C: 2 % -5 == 2
If you want to use modulus in C rather than Euclidean remainder, then you have to use a function like this, which does what Python does: long mod(long x, long y) {
if (y == -1) return 0;
return x - y * (x / y - (x % y && (x ^ y) < 0));
}OTOH, the largest "wat" of C-like languages is the following:
> -8 % 5
-3
whyPython here is once again correct:
>>> -8 % 5
2You may as well use {white, blue, black, red, green} to represent congruences mod 5, mathematically speaking it's not wrong, as long as they respect the axioms of a field.
What's "wat" (in the same sense as the js "wat" talk) is that the answer changes if the argument becomes negative. By all reasonable definitions of modular arithmetic, -8 and 7 are in the same class mod 5. Why is then:
> (-8 % 5) == (7 % 5)
false
in C-like languages? I'm pretty sure it's perfectly consistent with all the relevant language standards, but it's a "wat" nonetheless.It's the same in Python though? you just prefer the way it behaves (though to be fair it's generally more useful and less troublesome): Python's remainder follows the sign of the divisor (because it uses floored division), while C's follows the sign of the dividend (because it uses truncated division).
Strictly speaking, neither is an euclidian modulo.
Haskell does it right. Instead of having operators like %, it gives the programmer choice to choose either the modulus or the remainder. Using GP's example:
ghci> 2 `div` (-5)
-1
ghci> 2 `mod` (-5)
-3
ghci> 2 `quot` (-5)
0
ghci> 2 `rem` (-5)
2
And we have both (x `quot` y)*y + (x `rem` y) == x and (x `div` y)*y + (x `mod` y) == xCompletely agree with you, and not just in C. Using esoteric features of any language is equivalent to putting landmines in front of your teammates - bad idea.
Also, outside of production code, using esoteric features is a good way to get familiar with every corner of that language - which is useful so that you can diffuse those landmines when they are accidentally created. i.e know them to avoid them.
Simple MISRA C with all warnings on is already close to language with strict type checking. C compilers give you a choice, use it. If you are programming for money, good static analyzer makes it possible to write safety critical code.
"Although the first edition of K&R described most of the rules that brought C's type structure to its present form, many programs written in the older, more relaxed style persisted, and so did compilers that tolerated it. To encourage people to pay more attention to the official language rules, to detect legal but suspicious constructions, and to help find interface mismatches undetectable with simple mechanisms for separate compilation, Steve Johnson adapted his pcc compiler to produce lint [Johnson 79b], which scanned a set of files and remarked on dubious constructions."
Dennis M. Ritchie -- https://www.bell-labs.com/usr/dmr/www/chist.html
Unfortunately too many think they know better than the language authors themselves.
Would it be possible to have a standard, where undefined behavior is just a compile error? What would we lose - apart from legacy compatibility?
UB started as means to not kick out computer architectures that would otherwise not be able to be targeted by fully compliant ISO C compilers.
Given that C prefers to be a kind of portable macro assembler than care about security, it was only a matter of time until those escape hatches started to be taken advantage for optimizations.
Same applies to other languages, however since their communities tend to prefer security before ultimate performance, some optimization paths are not considered as that would hurt their safety goals.
In what concerns C, C++ and Objective-C, dropping UB optimizations would mean going back to the 1990's in terms of code quality.
No [if you're aiming for something in the same vein as C]. Undefined behavior is ultimately an inherently dynamic property--certain values could make a statement execute undefined behavior, and consequently, virtually every statement could potentially cause undefined behavior. Note that this remains true even in languages like Rust: Rust has loads of undefined behavior, but you do have to wrap code in unsafe blocks to potentially cause undefined behavior.
> What would we lose - apart from legacy compatibility?
In particular, it is clear at this point that if you want to permit converting integers to pointers, you will either have to live with undefined behavior (via pointer provenance) or forgo basically all optimization whatsoever.
int foo(int x, int y) { return x+y; }
compile? After all, this function can be called in a way that causes UB.(This is the case in e.g. C#.)
Unless your program can prove, that the function is only ever called with arguments that don't hit UB.
> Sorry about that -- I didn't give you enough information to answer this one. The signedness of the char type is implementation-defined
...why not have an "implementation-defined" answer button then, because that's what people should know (instead of knowing all ABIs) and what the question is about anyway?
> If these operators were right-associative, the expression [x - 1 + 1] would be defined for all values of x.
That's just wrong, no? If + and - were right-associative, it would be parsed as x - (1 + 1), which is decidedly not "defined for all values of x".
I.e. C standard enforces undefined very sparingly iirc, and most of the corner cases are implementation defined.
I may also be misremembering things: it's been 20 years since I've carefully read C99 (draft, really) for fun :))
If you read the text of C99 carefully, yes, it's implied that INT_MIN % -1 should be well-defined to be 0. However, the % operator is usually implemented in hardware as part of the same instruction that does division, which means that on hardware where INT_MIN / -1 traps (thereby causing undefined behavior), INT_MIN % -1 will also trap. The wording was changed in C11 (and C++11) to make INT_MIN % -1 explicitly undefined behavior, and given the reasoning for why the wording was changed, users should expect that it retains its undefined behavior even in C89 and C99 modes, even on 20-year-old compilers that predate C11.
And here's an exact quote from the C11 standard:
>If the quotient a/b is representable, the expression (a/b)*b + a%b shall equal a; otherwise, the behavior of both a/b and a%b is undefined.
From C99 §6.3.1.1:
> — The rank of long long int shall be greater than the rank of long int, which shall be greater than the rank of int, which shall be greater than the rank of short int, […]
> — The rank of any unsigned integer type shall equal the rank of the corresponding signed integer type, if any.
According to the standard[1], if short and int have the same size (even if not the same rank) both numbers are converted to unsigned int (that is, the unsigned integer type corresponding to int) because int can't represent all values of unsigned short.
The usual arithmetic conversions never "promote" an unsigned to signed if doing so would change the value.
The question would be fine (given that note) if "impementation-defined" was not an option, like for example the question about "SCHAR_MAX == CHAR_MAX".
The last x86 processor I know of to have these sizes is 80286.
The last machine with a word size of 16 bits? The 286 was as well. There were others like the WDC 65816 (the Apple IIgs and SNES CPU).
It just so happens that there simply far fewer 16 bit CPUs then there were 8 or 32 (or “32/16” like the 68k). Also 8 bit CPUs are simply a poor fit for C by their nature and the assumptions C makes. But the numerous ones still relevant today will use 16 bit ints.
The use of Real or V86 mode on the x86 went on for many years after the demise of the 8086. I think it is a somewhat of a joke at this point that they’re teaching Turbo C in some developing countries.
> abs(-2147483648) = -2147483648
I think one's complement is more sensible since it doesn't have this problem, but it loses out because it requires a more complex ISA and implementation.
> The name "ones' complement" (note this is possessive of the plural "ones", not of a singular "one") refers to the fact that such an inverted value, if added to the original, would always produce an 'all ones' number
In fact, for every numerical base there is an equivalent. For example, in base 3 you can represent a number in twos' complement notation (each trit X is replaced with 2-X, so -3 is represented as 212 on three trits, with +3 being 010) is or in three's complement by adding 1 (212+1 = 220). For base 10, you can do nines' complement (-3 on 3 digits is 996) or ten's complement (996+1 = 997).
You should prefer signed types for computing (if not storing) sizes and for pointer arithmetic, since they are more forgiving with underflow.
Using an integer, in the 1000s of for loops I've written, none get even remotely close to the billions - it is optimizing for a 1 in a million case, and if I know something can run into the billions of iterations I'm going to pay more attention to anyway. I've seen 0 occurrences of bugs relating to this kind of overflow.
Using a size_t, it is effectively an unsigned integer that risks underflowing which can easily cause bugs like infinite loops if decrementing or other bugs if doing any index arithmetic. I've seen many occurrences of these kind of bugs.
Any such code submitted to one of our projects would be rejected or fixed to use the proper type.
I've not introduced a security bug in every for loop I've written. What I've written shouldn't be controversial, just take a look at Googles style guide:
"We use int very often, for integers we know are not going to be too big, e.g., loop counters. Use plain old int for such things. You should assume that an int is at least 32 bits, but don't assume that it has more than 32 bits. If you need a 64-bit integer type, use int64_t or uint64_t.
For integers we know can be "big", use int64_t.
You should not use the unsigned integer types such as uint32_t, unless there is a valid reason such as representing a bit pattern rather than a number, or you need defined overflow modulo 2^N. In particular, do not use unsigned types to say a number will never be negative. Instead, use assertions for this.
If your code is a container that returns a size, be sure to use a type that will accommodate any possible usage of your container. When in doubt, use a larger type rather than a smaller type.
Use care when converting integer types. Integer conversions and promotions can cause undefined behavior, leading to security bugs and other problems."
One of these days, the compiler will do something surprising to one of your expressions involving signed integer overflow, like converting x < x + 1 to true. Or it'll delete a whole loop because it noticed that your code is guaranteed to trigger an out-of-bounds array read, e.g.: https://devblogs.microsoft.com/oldnewthing/?p=633
> If you do that you'll be fine.
I would not trust code written based on the methodology you described. However, if you also add -fsanitize=undefined,address (UBSan and ASan) and pass those tests, then I would trust your code.
Are you really, though? I would argue that it's a matter of perspective and/or semantics.
The Linux kernel is built with -fwrapv and with -fno-strict-aliasing, and uses idioms that depend on it directly. We can surmise from that that the kernel must be:
1. Exhibiting undefined behavior (according to a literal interpretation of the standard)
OR:
2. Not written in C.
Either way, it's quite reasonable to wonder just how much practical applicability your statement really has in any given situation -- since you didn't have any caveats. It's not as if the kernel is some esoteric, obscure case; it's arguably the single most important C codebase in the world. Plus there are plenty of other big C codebases that take the same approach besides Linux.
Lots of compiler people seem to take the same hard line on the issue -- "the C abstract machine" and whatnot. It always surprises me, because it seems to presuppose that the only thing that matters is what the ISO standard says. The actual experience of people working on large C codebases doesn't seem to even get acknowledged. Nor does the fact that the committee and the people that work on compilers have significant overlap.
I'm not claiming that "low-level C hackers are right and the compiler people are wrong". I'm merely pointing out that there is a vast cultural chasm that just doesn't seem to be acknowledged.
x and y are unsigned short. The expression `x*y` has defined behavior for...
a) all values of x and y.
b) some values of x and y.
c) no values of x and y.
If short = 16 bits and int = 32 bits (or heck even 17 bits), then x and y will be promoted to signed int. Signed multiplication overflow is undefined behavior, so x*y will be undefined for some values when x*y is too large. In particular, 0xFFFF * 0xFFFF = 0xFFFE_0001, which is larger than INT_MAX = 0x7FFF_FFFF.
If short = 16 bits and int = 64 bits (or even 33 bits), then x and y will be promoted to signed int. The range of x*y will always fit int, so no overflow occurs, and the expression is defined for all input values.
Isn't C fun?
But if int were larger (eg 64 bit) and short remained 16 bits, there’s no overflow and the answer is (A). I think.
I missed that the result would be promoted to signed int, which only most significant bits are set
Personally, I love C but I don’t use it in anything serious. Probably for the best given I apparently cannot remember how C ints work.
As antirez said as well, I don't pride myself in understanding the intricacies of integer arithmetic and promotion well. I try to write clear code by writing around the less commonly understood rules. Nevertheless I wanted to test myself. There are two questions that surprised me somewhat.
Apparently you can't left-shift a negative value even by a zero amount, as in (-1 << 0).
And is it true that the value of "-1L > 1U" is platform dependent? I had assumed that 1U would be promoted to 1L in this expression, even on x86 where unsigned int and long have the same number of bits. According to the following document, long int has a "higher rank" than unsigned int. https://wiki.sei.cmu.edu/confluence/display/c/INT02-C.+Under... . (Edit: according to rules 4 and 5 of "Usual arithmetic conversions", it's not only the rank but also about "the values that can be represented")
This is relevant to the question asking if -1L > 1U.
I think `undefined behaviour` still has its place in C - dereferencing a freed pointer comes to mind as an obvious example - but I think a good proportion of the really unintuitive UB conditions could be made saner without sacrificing portability or optimisation opportunities.
If the hardware traps on bogus pointers, then reading a bogus pointer may trap. But if you read a recently-freed pointer, it may still be valid according to the hardware (e.g. will have valid PTEs into the processes address space) so won't trap. Therefore you won't be able to guarantee any particular behaviour on an invalid read, so I don't think you'd be able to get away with "unspecified" or "implementation defined" behaviour on most hardware.
In practice, compilers aren't actually adversarial. A lot of the discussion around UB is catastrophizing and talks about how the compiler will order you pizza or delete your disk. Some problems are real and I know some people whose graduate school work was very specifically on the problems that this causes for security-related checks but compilers really really do not transform your program into "complete garbage." They transform your program into a program with a bug, which is a true statement about that program.
I'm reminded of the apocryphal story about being asked how a computer knows how to do the right thing if given the wrong inputs. This feels similar.
> On two occasions I have been asked, — "Pray, Mr. Babbage, if you put into the machine wrong figures, will the right answers come out?" In one case a member of the Upper, and in the other a member of the Lower, House put this question. I am not able rightly to apprehend the kind of confusion of ideas that could provoke such a question.
-- Charles Babbage
https://en.wikiquote.org/wiki/Charles_Babbage#Passages_from_...
And further, defining a lot of UB won't actually improve things. Imagine we define signed integer overflowing behavior. Hooray. Now your program just has a different bug. If you've accidentally got signed integer overflow in your application then "did some weird things because the compiler assumed it would never overflow" is going to cause exactly the same amount of havoc as "integer overflowed and now your algorithm is almost certainly wrong."
This is exactly what JF Bastien argues in his hourlong talk: https://youtu.be/JhUxIVf1qok?t=2284
So yeah, defining signed integer overflow isn't going to fix bugs in the vast majority of existing programs.
That being said, enforcing signed overflow wraparound at least makes debugging easier because it's reproducible. This is how it is in Java land - int overflow wraps around, and integer type widths are fixed, so if you trigger an overflow while testing then it is reliably reproducible across all conforming Java compiler and VM versions and platforms.
Defined for all values of x
Defines for some values of x
Defined for no values of x
(chose second answer) Wrong answer
Shifting (in either direction) by an amount equalling or exceeding the bitwidth of the promoted operand is an error in C99.
So according to Wikipedia: https://en.wikipedia.org/wiki/C_data_types
int signed signed int Basic signed integer type. Capable of containing at least the [−32,767, +32,767] range.[3][a]
So a minimum of 16 bits is used for int, but no maximum is specified. Thus, if my C compiler on my 64-bit architecture uses 64 bits for int, this is perfectly allowed by the specification and my answer is correct.
> All other things being equal, assume GCC/LLVM x86/x64 implementation-defined behaviors.
However, you're right to point out that the quiz would be better if you can't make any more assumptions than guaranteed by the basic language standard.
For char, you have 3: signed char, unsigned char and char. It's not specified if char without keyword is signed or unsigned.
You have integer types such as size_t, ssize_t and ptrdiff_t. They may, under the hood, match one of the other standard int types, however this differs per platform, so you can't e.g. just easily print size_t using the standard printf formatters, you really have to treat is as its own type. Also wchar_t and such of course.
Then you have all the integers in stdint.h and inttypes.h. Same here applies as for size_t. At least you know how many bits you get from several of them, unlike from something like "long".
Then your compiler may also provide additional types such as __int128 and __uint128_t.
This has been fixed in C99. For size_t, it's "%zu", for ptrdiff_t it's "%td", for ssize_t it's "&zd" and for wchar_t, it's "&lc".
For example, you can see what (signed long) > (unsigned int) would convert to, under different environment bit width assumptions.
Also, there's a discussion on Reddit: https://www.reddit.com/r/cpp/comments/x4x01f/cc_arithmetic_c...
Still, the most important suggestion I would make here is: Always compile with warnings enabled, in particular:
* -Wstrict-overflow=1 and maybe even -Wstrict-overflow=3
* -Wsign-compare
* -Wsign-conversion
* -Wfloat-conversion
* -Wconversion
* -Wshift-negative-value
(or just `-Wall -Wextra`)
And maybe also:
* -fsanitize=signed-integer-overflow
* -fsanitize=float-cast-overflow
(These are GCC flags, should work for clang as well.)
struct { long unsigned a : 31; } t = { 1 };
What is t.a > -1? What about if a is instead a bit-field of width 32? (Assuming the same platform as in the quiz.)C, like Rome, is a wilderness of tigers. Why ask for trouble?
Which both make signed integers take the expected two's complement behavior, or traps signed integer overflow, respectively.
N.B. left shifting into the sign bit is still undefined behavior... On x86-64 gcc and clang seem to perform the shift as if the number is interpreted unsigned, then shifted, then interpreted signed.
1- If you need arithmetic, use signed; if you need a bitmask, use signed. C is not assembly, and bit shift is no multiplication nor division.
2- Make sure you stay within bound, as in: don't even think you can approach the boundaries. C is not assembly, and overflow flag does not exist.
The fact that the text butts up to the left of the page with no margin is pretty incredible.
#include <stdint.h>
and don't use primitive types, and you will avoid many of these issues.
Most professionals don't worry about these effects, instead choosing to cast and use parenthesis directly. Those instruct the compiler instead of relying upon warnings from the compiler.
Abolish those three concepts and C is a significantly improved language!
Any system that lacks them will have to hack them back in at some point, in one form or another.
Option<types> are not an option for C, and sentinels like NULL work just fine.
Only because the C FFI is the lingua franca of operating systems.
> Option<types> are not an option for C
Disagree. C could absolutely support SumTypes / TaggedUnions.
> sentinels like NULL work just fine.
Spans/Slices are objectively superior. Null terminated strings aren't the biggest source of pain in the universe, but they're problematic.
I'm not arguing that C _should_ add these things today. I'm saying that C made bad choices 50 years ago and we're paying the price to this day. This isn't pointing fingers. It'd be upsetting if the industry hadn't made progress and didn't learn better ways to do things!
No, because they are genuinely useful. How should you best implement Option<type> for pointers? Right, using NULL as a sentinel. And representing sum types explicitly in the type system (instead of ad-hoc using sentinels) would add a huge amount of complexity.
I still like the simplicity of zero-terminated strings. They are self-contained strings that can be passed around as a simple char-pointer. No other format can give you this simplicity. They are still an acceptable default for string literals. And whenever you have a non-trivial usecase, just use the best string representation for that. You don't need to use zero-terminated strings everywhere.
Use a macro like this for a dumb length-delineated string
struct MyString { const char *ptr; int length; };
#define S(stringlit) (struct MyString) { "" slit, sizeof slit - 1 };
Done!This is an issue? Works fine for me.
I'm glad I run clang with -Weverything and use ASan and UBSan.
I limit shift strictly to unsigned integer numbers.
uint16_t x = (...);
uint16_t y = x << 15;
Any arithmetic involving uint16_t will be promoted to some kind of int. If int is 16 bits wide, then uint16_t will be promoted to unsigned int before the shift, and all values of x are safe. Otherwise int is at least 17 bits wide, then uint16_t will be promoted to signed int before the shift.On a weird but legal platform where int is 24 bits wide, the expression (uint16_t)0xFFFF << 15 will cause undefined behavior.
My workaround for this is to force promotion to unsigned int: (0U + x) << 15. https://stackoverflow.com/questions/39964651/is-masking-befo...
Since most languages use the same IEEE-754, you can hate them all.
Good grief, thanks Quiz for reminding me to stay away from that mess of language.