Checked integer arithmetic in the prospect of C23
gustedt.wordpress.com
gustedt.wordpress.com
bool add_will_overflow(int32_t a, int32_t b) {
uint32_t c = (uint32_t)a + (uint32_t)b;
return (((uint32_t)a ^ c) & ((uint32_t)b ^ c)) >> 31;
}
That produces the following assembly (see Godbolt[2]): lea edx, [rdi+rsi]
mov eax, edi
xor eax, edx
xor esi, edx
and eax, esi
shr eax, 31
ret
In Rust, you can write a.checked_add(b).is_none() which produces the following assembly[3]: add edi, esi
seto al
ret
A fun fact about this code: the overflow flag which is set by the add instruction and then harvested dates back at least to the 8080 (almost 50 years ago) and is not present in vanilla ARM. However, Apple Silicon has it as an extension, to make life easier for Rosetta 2 binary translation[4]. So when you do get to use this shorter code sequence, be thankful of the effort that chip designers put in to make it execute efficiently.I expect the C23 built-in functions will perform as well as Rust here, which is a win both for ergonomics (you can't really consider the current state of "will a+b overflow" to be discoverable) and performance.
[1]: https://mastodon.online/@raph/109535617953722719
[2]: https://godbolt.org/z/17zMsWjYv
Although, recently, I noticed I wanted a case where I wanted checked (u32 - u32) -> i32 and (u32 + i32) -> u32 operations, which even Rust's standard library doesn't provide. (The use case is keeping track of a running delta between two lists of u32 values--the delta can go positive or negative, so it has to be signed, but the values in the lists can never be negative).
Worse, in C or C++ you also need to find a way to do it without undefined behaviour. You can't just do the operation and see if the result matches expectations…
i32::checked_add_unsigned(some_u32) and
i32::checked_sub_unsigned(some_u32)
... which I think are exactly what your parent needs.I really hope stuff like this is added to the standard.
They also give example code
bool add_invalid = ckd_add(&result_add, a, b);
I can see that fits with “most of the time, anything positive means ‘no error’”, for example in malloc, write, read or printf, but these new functions return bool, not int, and the chosen method will require writing a double negation sometimes: if(!add_invalid) { … }
That’s not too bad, but if I were to see if(!ckd_add(&result_add, a, b)) { … }
I would expect that to test for failure, not success.Because of that, I think I would have chosen to return true on success, false on failure. I’m curious as to what arguments led to the choice made.
If only it were so simple. read and write, for example, return a number less than zero on error and a non-negative number on success, and malloc returns zero on error, and nonzero on success.
The general rule for early C seems to be “whatever’s the best way to cram a return value or an error in an int” (probably the correct decision for the time)
Also, these new functions return a bool, which, in C23, gets integer-converted to zero for false and one for true (https://en.cppreference.com/w/c/language/bool_constant. C17 had macros for true and false, with false being zero)
and the reverse, converting to bool similarly has zero fro false (https://en.cppreference.com/w/c/language/conversion#Boolean_...):
“A value of any scalar type (including nullptr_t) (since C23) can be implicitly converted to _Bool. The values that compare equal to an integer constant expression of value zero are converted to 0 (false), all other values are converted to 1 (true).”
Functions that only return an error code like `stat`, `connect`, and the proposed ckd_add, return 0 on success and nonzero on error.
Here, the committee just standardized existing practice, namely the gcc builtins. We just adjusted the call sequence in putting the pointer parameter for the result first.
int rc = func();
if( rc ) { /*handle error*/ }
examples from the stdlib: connect(), stat(), etc. Hell, even main is defined to return 0 on success, nonzero on error. if(ckd_add(&result_add, a, b) == CKD_SUCCESS) { ... }
or alternatively, if(CKD_SUCCESS(ckd_add(&result_add, a, b))) { ... }They felt fine changing the argument order, why stick with the reverse polarity?
Which part of the example implemented using `nullptr` instead of `NULL`, which is also from the future, though?
#include <stdckdint.h>
bool ckd_add(type1 *result, type2 a, type3 b);
bool ckd_sub(type1 *result, type2 a, type3 b);
bool ckd_mul(type1 *result, type2 a, type3 b);
#include <stdckdint.h>
#include <limits.h>
/* ... */
int x;
int a = INT_MAX;
int b = INT_MAX;
if (!chk_add(&x, a, b)) {
/* error! */
}
Other stuff on the table for C23- https://thephd.dev/c-the-improvements-june-september-virtual...
- (PDF) https://open-std.org/jtc1/sc22/wg14/www/docs/n3054.pdf
does it return non-zero on success???
So it is on the programmer to define what happens on error, they could ignore, try to back off by computing the high value bits, `exit` or `abort`.
Also they can be useful to implement bigints.
Signals (kill), signal, raise, sigsetjmp, siglongjmp etc are C's exception handling mechanism. It's not as well integrated into the language as say a try-catch construct but it works well enough for situations like these. See: signal.h and setjmp.h
To the question of "why not raise a signal on integer overflow in C?" - because signals are a terrible way of dealing with this. The signal handler doesn't know what code caused the overflow, and can't really do anything about it. Once the signal handler returns, the code itself has no idea it caused an overflow. Signals are a way for the OS to send signals to your program and not for control flow, after all. That's why `feenableexcept` is a niche extension that nobody uses.
The standard way of checking for fp errors is by calling `fetestexcept`. Personally I prefer this strategy (doing operation, then checking for errors) vs the new proposal for ints (checking for potential errors before doing the operation). But that is a matter of taste.
https://en.wikipedia.org/wiki/Exception_handling
(which has this bit: "C does not have try-catch exception handling, but uses return codes for error checking. The setjmp and longjmp standard library functions can be used to implement try-catch handling via macros.")
Its the same thing as "run time". Pedantically, crt0 exists and therefore C has a "runtime". But it is nothing like what we refer to as a "runtime" today. The literal words are true but the meaning of the words doesn't match expectations.
C is likely considerably older than those 'other programming languages' and I'm still stuck in the past with my terminology.
Signals were a feature of Unix and C from the early 1970s; at least by 1973, if not earlier. Signals were a software-based abstraction over the concept of hardware interrupts, and even today we often say that hardware interrupts are triggered by (among other reasons) an "exception" or "hardware exception".
Similarly, the fact that the concept of interrupts and exceptions are related can be seen in the later (4.2 BSD, 1983) select(2) syscall, where the third fd_set argument is for capturing "exceptional condition(s)". (https://pubs.opengroup.org/onlinepubs/9699919799/functions/s...)
Ada enjoyed the benefit of at least another 10 years of reflection and evolution in computer science when choosing the semantics of their exception mechanism. But the success of Ada's (or Ada-like) semantics hasn't yet completely redefined the concept of exceptions.
The only other realistic options outside of government/enterprise were Pascal, BASIC and assembler, and within the bulk of the work was done in COBOL.
As you should, for any datatype.