The type of char + char
blog.knatten.org
blog.knatten.org
I know what the type of char + char is. I know that it's either int or unsigned int, depending on the ranges of values supported by types char and int. I know what it is for any given implementation. And I know that it's int, not unsigned int, for every implementation I've ever used or am likely to use.
Implementation-defined features are not some unsolvable mystery. They're just implementation-defined.
[1] int8_t is not required.
[2] char is the fundamental unit of addressability. sizeof char always evaluates to 1, sizeof int8_t must be non-0, char must be at least 8 bits, and int8_t must be precisely 8 bits, therefore sizeof int8_t == sizeof char and CHAR_BIT == 8.
char is an arithmetic type, but it rarely makes sense to treat it as one, because its signedness is implementation-defined. If you want a very narrow integer type, both signed char and unsigned char are arithmetic types, and can reasonably be used that way. (Arrays of unsigned char are also used for raw memory.)
And you should understand how char, signed char, and unsigned char behave when you do use them as arithmetic types.
Promotion to int or unsigned int, depending on the range of the type, can be confusing. The same applies to all integer types with lower rank than int, including short, unsigned short, and intN_t and uintN_t for N==8 (and probably for N==16, and maybe for larger N).
Note also that this:
char c = '0';
++c;
is guaranteed to set c to '1'. (This guarantee applies only to decimal digits, not to letters.) #include <type_traits>
static_assert(std::is_arithmetic<char>::value, "char is arithmetic");(In C++, you of course have operator overloading, that's how std::string concat sugar works.)
'1' + '1' == 'b'. Because 49 + 49 == 98. ASCII '1' == 49, and 'b' == 98.
A multiplication `uint16_t * uint16_t` can still cause an overflow after promotion to signed int, which is undefined behavior! So "unsigned types wrap around" doesn't apply to `uintN_t`, because you can never know for sure whether those types are "smaller than int" and thus get promoted to signed types when you do any arithmetic.
Of course, in practice this just means: every C and C++ program relies on tons of implementation-defined behavior. A `sizeof(int)` greater than 32-bits would break most code in existence (e.g. hash code computations using `uint32_t`).
Last time similar thing bit me was when the platform had 16-bit int... so just adding two int16_t can very well cause int overflow.
You can deduce the width (number of sign + value bits) of the standard integer types from their limits (e.g. INT_MAX, INT_MIN, etc). The problem has been that this is non-trivial if not impossible to do from the preprocessor. The next C standard will include width constants (e.g. INT_WIDTH) for the standard integer types.
char toupper(char c) {
if (c >= 'a' && c <= 'z') {
return (c - 'a') + 'A';
else { return c; }
}In C++, 'a' is a char and the comparison result is a bool, though it doesn't really make a difference in that function.
toUpper(char) -> $ + char. % It's "$ " (c - 'a') + 'A'
contains both (c + 'A') - 'a'
to make this more clear, but I think that's actually UB with signed chars- e.g. for c='a', 'a'+'A' exceeds the range of a signed 8-bit value!Promotion should save us here, but that's a bit too yikes-y for my comfort.
char + char = invalid
char + offset = offset + char = char
offset + offset = offset
char - char = offset
char - offset = char
offset - char = invalid
offset - offset = offset
Given that, (c - 'a') + 'A' is perfectly valid without adding two characters.
edit: formatting
>I think that char – char should definitely be legal. The distance between characters is well defined. Same for char + numeric. Both logically makes sense. I think a good analogy might be floors in a building. Asking what’s the distance between the second and seventh floor makes sense, or what’s two floors above the 4th. But the question ‘what’s the 5th floor plus the 6th floor’ doesn’t make sense.
>Affine space describes these kind of relationships in mathematics. Eg position and disposition in n dimension, or count and offset in buffers, even timestamp and duration.
https://stackoverflow.com/questions/27001604/32-bit-unsigned...
https://stackoverflow.com/questions/39964651/is-masking-befo...
A sadistic part of me would prefer if it was interpreted as a bitwise and... not because that's good or reasonable or smart... but to punish the behavior. But then that backfires when people use it for underhanded code.
Specs don't matter beyond the conventions they inspire.
In a standards committee, the standard is the standard, in practice common practice is the standard.
That said, the int promotion alone may be surprising / nonobvious to some people (it was to me, when I learned about it!).