My first guess is that it involved some kind of architectural peculiarity of the PDP11.
One's Complement arithmetic wasn't uncommon on computers of the era and is fully symmetric. And, I believe that a PDP-7 uses one's complement.
2) "unsigned" didn't get implemented until about 5 years after C was initially implemented.
Why? Because it's quite likely the result will be used in arithmetic expressions involving other signed ints and/or subtraction. That goes with the kinds of things you'd use normally abs() for.
There will be some cases where abs() is being used by someone on a value that could be INT_MIN. But when your code might have to handle the full range of possible integers (signed or unsigned), there are probably many other things to get right as well. In such code, abs() is the least of your worries.
I think most uses of abs() are in arithmetic expressions of "numbers whose values we don't expect to overflow", and I'm fairly confident if abs() returned unsigned int, there would be more bugs that nobody spotted in the world than with it returning signed int. Not many more, because abs() is rarely used, but a few.
I used to be one of those people who felt it made sense to use unsigned types in C for values that can never be negative. That was tidier and clearer. It stated my intent.
After a few years of coding in C like this, I changed my mind: I spotted occasional little hidden bugs here and there, undetected by the compiler or the programmer, from overzealously using unsigned types just for "showing intent" where signed would have been fine.
It's not like using unsigned types prevents arithmetic bugs in C in practice (unless you really have values exercising most of the unsigned range). The "use unsigned because types should reflect intended range" argument is muddier than it first looks: The range [0..2^B-1] is no more "correct" than [-2^(B-1)..2^(B-1)-1] for almost all quantites, as people often aren't paying attention to correct handling of numbers in the upper end of the unsigned range either. They rely on "practical numbers are small enough that it doesn't matter", same as with signed ints.
If unsigned acted as a range-constrained arithmetic type in C, meaning "this variable can only hold a subset of the default integer type" and "arithmetic with other types is consistent", that would be different.
But it doesn't. It acts more like an unsignedness virus in C arithmetic, adding unnecessary boundary conditions into innocent-looking expressions.
Of course you can be aware of these issues and avoid them. I'm pretty experienced and can avoid such issues easily. But having to be careful for no added benefit, especially with multiple people, just raises risks. So now I default to, and advocate, sticking with signed integer types for values representing arithmetic quantities. If I were designing a language, I'd probably advocate for ranged types instead. That means limited to ranges like [0..2^(B-1)-1] (the upper half of the signed range), and have consistent arithmetic. But C is not like that.
Unsigned types in C are of course completely appropriate for bitwise and modular arithmetic uses, and for holding raw data. I'm a big fan of using them for relative timestamps using modular arithmetic comparisons, as done in Linux. If you're writing compression routines or a database engine they will be very useful.
But for general arithmetic quantities, nowadays my view is similar to this answer: https://stackoverflow.com/questions/51677855/is-using-an-uns.... I'm sympathetic to the "unnecessary discontinuity near common values" and "arithmetic closure is more useful" view. People won't agree on this. They don't agree in answers to that SO question. All I can say is, I used to be zealously "use unsigned for non-negative quantities", and then after some experience it seemed more pragmatic and safe-by-default to stop doing that, and the type-as-intent was misguided anyway when the upper range of unsigned wasn't being used. You still have to be unsigned-aware in C due to size_t especially. So I still use unsigned-by-default for arithmetic dancing around object sizes, memory and offsets in C.
One reason I picked this route is because Rust only lets you index arrays by usize and not isize, so I chose to extend this "unsigned for non-negative quantities" philosophy to C++. Additionally, std::vector::size() is size_t and unsigned, so I decided to follow. Because of the mixture of unsigned and signed indexing (and 32 vs. 64 bit values) across different libraries, I decided to turn on warnings so I know where incompatibilities lie. Rust outdoes C++ because it makes incompatible integer widths hard errors, comes with the equivalent of `-Wtautological-unsigned-zero-compare` out of the box, and has runtime checking (in debug mode) which panics if you decrement an unsigned value past 0.
I can understand "signed by default" even though I disagree and prefer not to use it for my own code. And I think 31-bit integers are a good idea if you don't need the range of an unsigned integer. I wish Rust had a type for "31-bit integer with negative values serving as a niche for enum cases" (though my concern is that a zero-overhead mutation API without runtime checks, combined with niche filling, would be UB since you can decrement 0 to -1 and effectively transmute an enum holding a u31).
Also signed integers aren't fully trouble-free, and can still overflow for very large differences (though that's less likely than 2 - 3). The Stack Overflow post mentions:
> Want to find the "delta" between two unsigned indexes into a file? Well you better do the subtraction in the right order, or else you'll get the wrong answer.
Well naive signed integer subtraction can be UB as well.
constexpr int f() {
return 0x7fffffff - -0x7fffffff;
}
constexpr int x = f();
<source>: In function 'constexpr int f()':
<source>:4:23: warning: integer overflow in expression of type 'int' results in '-2' [-Woverflow]
4 | return 0x7fffffff - -0x7fffffff;
| ~~~~~~~~~~~^~~~~~~~~~~~~
<source>: At global scope:
<source>:7:20: in 'constexpr' expansion of 'f()'
<source>:7:21: error: overflow in constant expression [-fpermissive]
7 | constexpr int x = f();
| ^
Compiler returned: 1
As jepler mentioned (https://news.ycombinator.com/item?id=28983587), if you want to handle arbitrary difference without overflow, you may need to compute the absolute difference and the sign separately. Sadly it's a massive pain to accomplish.I guess this case (large numbers) is very rare compared to accidentally subtracting 3 from 2, and addition also risks overflow. And if you're using subtraction for deltas, then restricting file sizes to 2^31 - 1 or 2^63 - 1 makes differences fit in int32_t or int64_t (64-bit ptrdiff_t).
Sure, but the point is that you have to get overflow due to high magnitude of the involved integers. So yeah INT_MAX - -INT_MAX is UB, but you have to have quantities up around INT_MAX for this to be a problem. For the unsigned case you just have to have quantities around 0.
Now, having quantities around INT_MAX isn't necessarily that unusual, but what seals the deal for me (I wrote that SO answer) is that 64-bit values are becoming more ubiquitous, and there we can guarantee in some sense that many quantities won't reach their 2^63 limits any time soon. So I definitely thing "signed by default" is much more of a slam dunk if it is paired with "64-bit by default".
Note that though I'm a proponent of signed by default, in practice I find it hard to write this way in C and C++ because you are constantly fighting the impedance mismatch between the language, size_t, size(), etc and your own rule. So signed by default is definitely more palatable in languages which made that same decision for their builtins and standard library.
Yes yes and yes. There is nothing that would make me happier than if GCC and Clang gave us the freedom to choose an ILP64 data model. In that case, we wouldn't need prototypes anymore and we could restore much of the original intent behind the design of C.
The only issue with negative indices is that it is no longer possible to have containers-of-bytes as large as the address space, but except for that tiny time window when an usable 4GB address space was actually available on 32 bits systems, it is in practice not a big restriction.
std::string::find is size_t and unsigned too, now guess what it returns in corner cases.
>"showing intent"
I think intent of a signed integer is "it's a number, don't think weird things here".
If you absolutely have to, there are escape hatches like Uint8Array/Buffer/etc which allow you to deal with bytes. There are also ways to force JIT to treat your numbers as integers with tricks like `i = i|0`, but that's barely used.
Obviously compiled languages are different, but to me you could rename `signed integer` to `integer` and `unsigned integer` to `rawdata32` or `byte4`