Unsigned int considered harmful for Java (2014)
nayuki.io
nayuki.io
"Oftentimes it’s a novice coming from C/C++ or C# who has worked with unsigned types before, or one who wants to guarantee no negative numbers in certain situations."
This is a neutral observation first and foremost. By itself there is nothing wrong with it. I would expect newcomers of a language to miss features of where they came from as well. Although it should be noted that this statement is only a gut-feeling without proper statistics to back it up. What irks me is what comes next:
"Said novice C/C++/C# programmer typically does not fully understand the semantics and ramifications of unsigned integers in C/C++/C# to begin with."
Not only does he presume that the newcomer does not fully understand the semantics and ramifications of unsigned values in java. He also asserts that this is the case for the language that they came from. How do you expect a reasonable debate to continue from here on out? Every disagreement is settled by just stating, that the proponent just doesn't understand the problem.
Now what if they understand the semantics and ramifications and still want that feature? Does a world exist for the author in which this is the case?
Moreover, the article is pretentious as well. The semantics of unsigned integers in java hasn't been decided yet, so you either have to argue against each proposed model or in general. Here, a model in which implicit type conversion from unsigned to signed values is argued against. Other possibilities are not considered. At this point I am inclined to just step away, as there is no information about both the semantics that is being argued against (method: find it by reading it) and, whether it is the only model proposed. There is no insight to be had except that the author doesn't want unsigned integers and is content with how it works as is (with an asterisk for bytes).
I do believe that the people in favour of the addition have to make a case for it. But I also thought that we were beyond arguments of the "you're holding it wrong" kind. As silly as it may sound, an argument like this makes me consider the other side immediately. And I say that as someone who never felt the need for unsigned integers in Java.
Never mind the fact that the vast majority of C/C++ programmers don't appreciate the fact that signed overflow in their languages is undefined behavior and can result in anything from correct behavior to complete unpredictable garbage. It's easy to argue for adding unsigned to Java because "I want this feature"; it's harder to argue against it because "I don't want others to use this feature / I don't like how this feature interacts with existing features".
Another example of not understanding types, did you know that Rust's usize can be as small as 16 bits? And that even on a 32-bit system, objects are limited to 2 GiB in size, not 4 GiB? https://stackoverflow.com/questions/32324794/maximum-size-of...
> The semantics of unsigned integers in java hasn't been decided yet
I think they have been decided. Java SE 8 added functions to treat bits as unsigned and perform operations (e.g. Integer.divideUnsigned(), Long.toStringUnsigned()). They might have also stated that it was a non-goal to add unsigned primitive integer types to the language.
For what it's worth, I am aware that the lack of unsigned integers in Java is a very contentious topic for decades, and the loudest voices are definitely on the pro-unsigned side. My favorite example is this thread: https://stackoverflow.com/questions/430346/why-doesnt-java-s...
Who cares? You could ask the same question about operator precedence, and the answer to both is similar: If it's unclear to you or those who might read your code, just be more explicit than strictly necessary. Handle or assert cases of possibly overflow. And by all means, you don't have to mimic the C/C++ implicit casting rules. If something's a potential pitfall, just make it an explicit cast (like `long` to `int` already is).
The status quo has some huge pitfalls of its own. `uint64` is out there, whether certain Java folks like it or not. It's there in protobuf, it's there in your databases, it's there in your FFI. This leaves people who have to interoperate with things from Java with the choice of either using `long` and hoping beyond hope your users read the docs, or using BigInteger and taking the accompanying performance hit, plus more correctness problems if your users try to operate on the value and take it outside its range.
The Java version of unsigned integers could completely disallow casting to and from signed integers and they would still be a very useful addition to the language.
Instead what's going to happen is we're going to get JEP 401 and you'll have a hundred different uint64 primitive classes that are exactly the same as each other except that they have different types. That or we'll get lucky and they'll add one to the standard library, at which point everybody who confidently asserted that adding unsigned integers to Java was a bad idea will have to rationalize to themselves that it doesn't really constitute adding them to the language because they're not in the bytecode, or they're not accessible without an import, or something.
I'd say that in this contest betweeen the author and somebody described as literally not understanding the implications of what they are asking for, the author might have eked out a victory. Now they can move on to the winner's bracket and argue against an intermediate-level engineer.
Finally a short story for the record. In 1968, the Communications of the ACM published a text of mine under the title "The goto statement considered harmful", which in later years would be most frequently referenced, regrettably, however, often by authors who had seen no more of it than its title, which became a cornerstone of my fame by becoming a template: we would see all sorts of articles under the title "X considered harmful" for almost any X, including one titled "Dijkstra considered harmful". But what had happened? I had submitted a paper under the title "A case against the goto statement", which, in order to speed up its publication, the editor had changed into a "letter to the Editor", and in the process he had given it a new title of his own invention! The editor was Niklaus Wirth.
https://www.cs.utexas.edu/users/EWD/transcriptions/EWD13xx/E...
That was an editorial title. The original title was : "A Case against the GO TO Statement"
The proof is in the use. If people create and use libraries that abstract for a primitive that every other compiled language has, then the primitive is missing from your language.
I also sometimes want to have unsigned integer types, mostly for data model correctness sake (aka this value can't be negative anyways, lets use unsigned). But so far I only really missed unsigned bytes in java (which the author acknowledges). And even in other languages I rarely "need" unsigned types.
Only a minority uses those workarounds, because they are cumbersome workarounds. If they were given proper unsigned types, they would use them much more often. I find myself using unsigned types in 99% of cases in Rust. I don't even remember the last time I had to use signed type.
As I work in bioinformatics, the in-memory collections tend to be pretty large. While the arithmetic is usually done with 64-bit integers, it often makes sense to store the numbers in 32 bits to save space. And since the length of a human genome is ~3 Gbp, that means unsigned 32-bit integers. Signed 32-bit integers are just bugs waiting to happen.
And sometimes the integers are stored in bit-packed arrays, where the width could be 29 bits, 33 bits, or something like that. Those are much easier and less error-prone with unsigned integers.
When adding unsigned int as a language feature, don't we get to make up the rules? It seems like we can choose to make the rules not awful; we are not beholden to what C++ has done.
Both are shit and will create bugs if you let them through, doesn't make much a difference.
In fact there are languages which allow negative indices, in which going to -1 is a lot worse, because that's a valid index, just a nonsensical one.
> Automatic signed to unsigned type coercion is extremely bug prone, so better avoid using unsigned integers entirely unless you never want to use them in math formulas
Or you can just not have "automatic signed to unsigned type coercion" in the first place. Or any sort of automatic coercion for that matter.
> You avoid many more bugs by limiting collections to 2 billion than by mixing signed and unsigned integers.
Hell you'd avoid many more bugs by limiting collections to 32000 too!
> I have had a lot of bugs related to unsigned to signed casts.
I've had a lot of bugs related to signed to signed casts.
That doesn't say anything about signed values, that says something about casts.
I think the same thing about strongly typed languages that lack generics. Generics are obviously a complex feature but, if you don't have them, you're just spreading the complexity elsewhere.
I’m not saying a good design exists or would be worth the cost to implement. This post just does not explore the solution space much.
Every time I work with byte in Java I have to cast it to an int with (b & 0xFF).
Utter madness.
> Bytes are a different story however. I argue that signed bytes make programming needlessly annoying, and that only unsigned bytes should have been implemented in the first place
> https://www.nayuki.io/page/javas-signed-byte-type-is-a-mista...
If someone designed a language this way today they would be called a master troll. At least you can require CHAR_BITS == 8 in practice for everything except weird embedded systems - both POSIX and Windows do just that.
Technically C++ now does indeed have a blessed byte type (std::type). But it is not an integral type: it support bitwise operations but not artihmetic ops, so signedness doesn't come into the picture.
Not an 8-bit byte which is the definition most people think of when talking about bytes and the only one that matters once you need to interact with other systems.
> and in practice.
Yes.
> Technically C++ now does indeed have a blessed byte type (std::type). But it is not an integral type: it support bitwise operations but not artihmetic ops, so signedness doesn't come into the picture.
Its signedness does matter for casts to integral types - int(std::byte(255)) == 255 whereas int(char(255)) is implemenation-defined.
byte a1 = 0x40; byte a2 = (byte) 0x80; a1 >>= 1; a2 >>= 1; System.out.printf("0x%02X\n", a1); System.out.printf("0x%02X\n", a2);
And that code (note the operator ">>>" instead of ">>"):
byte a1 = 0x40; byte a2 = (byte) 0x80; a1 >>>= 1; a2 >>>= 1; System.out.printf("0x%02X\n", a1); System.out.printf("0x%02X\n", a2);
Congrats if you guessed right on your first try, because I certainly did not.
I don't want to spoil the ending. I encourage people to run the code on their own after they made their guess to compare.
my model which might not be correct is:
this `byte a2 = (byte) 0x80` is a cast from the integer 0x80 to a byte and when there is a cast from integer to byte it just truncates the bits. when printing out this value it interprets it as 2s compliment so its 'value' is -128. so it just takes 8 bits starting from the least significant bit. so you end up with the bits 1_000_0000. then when you do this `>>=` operator that promotes to an integer and then it does the right shift then casts back to a byte. the promotion from integer to byte is just done by sign extending the byte to 32 bits. so you get 25 1's followed by 7 0's. the shift is a signed shift so you get 26 1's followed by 6 0's. then the cast back to byte leaves you with 11_000_000 which is why you get the 'weird' result of 0xC0 which is larger than 0x80. and the unsigned shift works in a similar way except the intermediate result is 0 followed by 25 1's followed by 6 0s. but this values in the high bits have no effect after the truncation which is why you get the same result.
you need to do `a2 & 0xFF` to clear out the sign bits outside the byte range before applying the shift in order to do the unsigned byte to unsigned int (or signed int?) promotion correctly. but using unsigned types in java is super dangerous.
i think it would be sensible to add unsigned, and disallow implicit conversion (definitely at least narrowing conversions).
but java is a distant world to me.
For example, implicit conversion from a 32bit signed integer into a 64bit signed integer.
I suspect part of the problem is conversions between types is complex and annoying to implement in a compiled language. With the added bonus that's only something the end users care about not the compiler writer. Hence in the wild this stuff tends to be broken or inadequate. Like in C it's way broken. In Java it's way inadequate.
Uint63 can be promoted to whatever is the largest native integer type on most systems, and int53 fits into a JavaScript number. It would be useful to have explicit checks for these values on the protocol level.
Mapping for C type to ByteBuffer call:
uint16_t -> getShort().toUShort()
uint32_t -> getInt().toUInt()
int32_t -> getInt()
It's not the most ergonomic set of calls, but definitely limits the error prone nature of it all since if you try to pass an Int into a function that needs a UInt, you'll get a compiler error> I argue that signed bytes make programming needlessly annoying, and that only unsigned bytes should have been implemented in the first place: Java’s signed byte type is a mistake.
> None of these alternatives is as correct as using uint32_t in C/C++, but I don’t think this situation comes up often enough to matter.
Never mind that the supposed benefits are extremely slim, and mostly due other weaknesses in Java
This was filed under the heading "Straightforward emulation". In the next section titled "Efficient emulation", I show that operating on two 64-bit values in unsigned mode is easily doable.
> other weaknesses in Java
Please explain. Keep in mind that Java is a simpler C++, and consciously chose to remove features that would be confusing or dangerous (e.g. unsigned integers, pointer arithmetic, destructors, multiple inheritance, operator overloading).
> when a file/network format specifies an unsigned data field – how should we represent it in Java? [...] None of these alternatives is as correct as using uint32_t in C/C++
Types like uint32_t don't buy you as much functionality as you think. In the last few projects I worked on, the domain of allowed values was a weird range. For example, 1 <= QrCode.version <= 40 ( https://github.com/nayuki/QR-Code-generator/blob/2643e824eb1... ). For example, 1 <= Png.Ihdr.width < 2^31 ( https://github.com/nayuki/PNG-library/blob/536acb238f9c000e2... ). Neither Java int nor C uint32_t capture these constraints properly.
https://docs.oracle.com/javase/8/docs/api/java/lang/Long.htm...
So in a program that uses this a lot, a reader can never tell reliably if a value typed "long" is signed or unsigned and the compiler can't catch a mistake like using "divideUnsigned" and then printing the result without using "toUnsignedString"...
I don't know how many strange issues I've tracked down that amounted to "this protocol buffer has a uint32 field, and surprise now the value is negative in Java and oops there was a check that cared about that." At least five or six issues.
At least when it comes to serialization, enforce invariants above the serialization layer.
I would love an unsigned byte type in Java though. What a pain.
Typing allows enforcing the boundaries where you go from signed to unsigned, and the bit fiddling is handled in the class
UShort.toInt(): Int = data.toInt() and 0xFFFF
They're defined as inline classes/functions that have no runtime cost so it's equivalent to the "Straightforward emulation" from the articleThen you can have explicit functions for the operation you want: signed_mul(), signed_less_than(), unsigned_greater_than(), add_assume_no_overflow(), etc. (add your favorite syntactic sugar/operator symbols for these).
Assembly is more explicit and clearer to understand in this regard.
The only time signed vs unsigned matters is for comparisons and mul/div, and those have explicit names so you're never surprised that 0xFFFFFFFF is less than 0 when doing signed_less_than().
Due to byte/short/char being promoted to int for any arithmetic, any time you do an operation like char+char→int, you'll still come into contact with signed types and need to continually cast back to char. This problem also occurs in C/C++, and you can typically find that uint16_t * uint16_t → signed int.
Using char to represent numbers (e.g. 16-bit channel values for pixels) instead of text is semantically incorrect. There's no problem if you're just working within Java, but once you get into comparing or porting code across languages, this is a code smell. Just as an example, the range for char in Rust is [0x000000, 0x10FFFF], and this is checked at run time when converting to char; this would be seen as weird from the lens of C/C++/Java where char is just a sequence of bits where all values are allowed.
Worry not - the only time i ever authored java code was when i was making the world's smallest (still) JVM. You have no reaosn to worry about bad smelling java code using char as uint16 from me :D