(Some Rust developers argue that even "size as i32" should be avoided, and "size.try_into()" should be used instead, since it forces the programmer to treat an overflow explicitly at runtime, instead of silently wrapping.)
(Some Rust developers argue that even "size as i32" should be avoided, and "size.try_into()" should be used instead, since it forces the programmer to treat an overflow explicitly at runtime, instead of silently wrapping.)
It's important to still have the option for efficient truncating semantics, though; some software (e.g. emulators) needs to chunk large integers into 2/4/8 smaller ones, and rotation + truncating assignment is usually the cheapest way to do that.
But, importantly, this is a rare case. Most software that does demoting casts does not mean to achieve these semantics.
So I wonder — are there any low-level/systems languages where a demoting cast with a generated runtime check gets the simple/clean syntax-sugared semantics (to encourage/favor its use), while truncating demotion requires a clumsier syntax (to discourage its use)?
edit: a cast will throw an exception if it fails, in case that was not clear from context.
Then safety is free...
I know Erlang's BEAM VM has a "fail jump pointer register", where instructions that can fail have relative-jump offsets encoded as immediates for those instructions, and if the instruction "fails" in whatever semantic sense, it takes the jump.
But most CPUs don't have anything like that.
Would you want it to trap, like with integer division by zero?
CPU traps are pretty hard to handle in most language runtimes, such that most compilers generate runtime checks to work around them, rather than attempting to handle them.
I always thought that overflows should be checked in hardware, I suppose it's not a stretch to extend that to truncation. It's controversial though, and obviously mostly a thought experiment anyway unless we manage to convince some CPU manufacturer to extend their ISA that way.
MIPS does have (optional) trapping overflow on signed add/sub overflow, so at least there's a small precedent for it.
It seems obvious that future Apple CPUs will have hardware support for this, if they don't already.
Given that Apple was one of the original founders of ARM it’s quite possible that their license allows much more latitude anyone else’s.
If AMX is allowed under their license, there is no reason why checked extensions would not be.
let x = size as i32;
... even if what you meant was closer to let x: i32 = size.try_into().expect("We are 100% sure size is small enough to fit into x");
But at least it isn't C or C++ where you might accidentally write x = size;
... and the compiler doesn't even warn you that size is bigger than x and you need to think about what you intended.It's really hard to fix this in C++. Some of the Epoch proponents want to do so using epochs to get there, basically you'd have a "new" epoch of C++ in which narrowing must be explicit, and old code would continue to have implicit narrowing so it doesn't break.
That's not true tho, compiler with reasonable flags set will definitely warn you and if you really don't like this kind of code you can force compiler to issue an error instead
int main(void) {
int some_int = 1234567;
char c = some_int;
return c;
}This is GNU's idea of "all".
Contrast to Clang's -Weverything, which will.
I can see the reason behind it, but I feel that this behavior is something you opt into when you use -Werror.
It feels like the "right thing" here would instead be for the compiler to allow build scripts to reference a specific point-in-time semantics for -Wall.
For example, `-Wall=9.3.0` could be used to mean "all the error checks that GCC v9.3.0 knew how to run".
Or better yet (for portability), a date, e.g. `-Wall=20210720` to mean "all the error checks built into the compiler as of builds up-to-and-including [date]."
To implement this, compilers would just need to know what version/date each of their error checks was first introduced. Errors newer than the user's specifier, could then be filtered out of -Wall, before -Wall is applied.
With such a flag, you could "lock" your CI buildscript to a specific snapshot of warnings, just like you "lock" dependencies to a specific set of resolved versions.
And just like dependency locking, if you have some time on your hands one day, you could "unlock" the error-check-suite snapshot, resolve all the new error-checks introduced, and then re-lock to the new error-check-suite timestamp.
[1]:https://docs.microsoft.com/en-us/cpp/error-messages/compiler... [2]:https://docs.microsoft.com/en-us/dotnet/fundamentals/code-an...
Personally i would vote for "Wall with Werror" means no guarantee for your build.
At least with -Werror on at all times, devs will tend to upgrade before the very-stable CI environment does, and thereby catch the problem at development time (usually less time-pressure) rather than release-cutting time (usually more time-pressure, esp. if the release is a hotfix.)
-----
Mind you, it does work to enable -Werror only in CI, if you lock your CI environment / compiler Docker image / etc. to a specific stable version, and treat that as the thing to re-lock in place of the "error-check suite snapshot version."
This has the disadvantage, though, that you can't take advantage of newly-stable/newly-unstable language features, or of newly-introduced compiler optimizations, without biting the bullet and taking on the work of fixing the errors introduced by re-locking the base-image.
With a separate flag for locking down the error-check-suite snapshot version, you could continue to upgrade the compiler — and thereby get access to new features / optimizations — while staying on a particular build regression "scope."
If you don't want your build to fail on warnings, don't use -Werror. If you want it to only fail on specific warnings, use -Werror=...
> with nobody sure why
Unless they look at the errors in the compiler output. What does it matter if it was brought on by a compiler update or a push?
> At least with -Werror on at all times, devs will tend to upgrade before the very-stable CI environment does
Nothing wrong with -Werror for devs - the problem is when you ship code to others and leave -Werror on by default.
And it is false. My default configuration C++ project created in Clion shows it very clearly, and even pesters to use int32/int64 over int/long.
But as usual the default fallback when you're wrong about C++ is "uh yeah but lotta footguns amirite"
As if there aren't enough that we need to start making them up...
0: https://developers.redhat.com/blog/2018/03/21/compiler-and-l...
Now that the default in the most beginner friendly of IDEs catches it, the goalpost is "my pet source of customization designed with C++98 in mind doesn't catch this"
Of course, even your pet source of customization caught up: https://developers.redhat.com/blog/2021/04/06/get-started-wi...
I mean MSVS uses Clang-tidy too, Clang-tidy integrates style guides provided by Mozilla and Google.
Most C++ Google projects have clang-tidy configs.
Clang-tidy is literally table-stakes for modern C++ tooling.
Github shows 970,000 commits related to setting up clang-tidy
But uh, yeah, let's see where the goalpost skitters to next.
-
The irony is I said above, C++ has enough footguns without sticking your fingers in your ears and ignoring boring, easy to setup, widely well known and well used tooling.
But in the war against C++ no stone must be left unturned.
C++ is a tiny fraction of all the code I've written in my life but it irks me to no end that people can't deal with the idea that language safety can improve, that tooling can be considered part of that safety. Or rather they can... unless they're talking about C/C++
I thought the point I had been making, as had others, is that by default this is an easy footgun.
There are all sorts of things that can be added on to all languages to help - if you know it’s a problem worth solving, etc. which is inevitably after you’ve footgunned yourself with it bad enough you felt the need to research how to prevent it.
Other languages just do the safer thing (or most compilers By default warn at least about common footguns) more - which is the whole point of this thread?
But tooling that is incredibly common, that beginners will run into even if they take the path of least resistance, and experts will use because it enforces standards at the very least, covers it.
Like Js without linters is a minefield, but everyone accepts you should lint your Js. Why does that change when C++ is involved?
The goalpost, since you're insistent on being explicit about it, was whether a C/C++ compiler "with reasonable flags" will catch the implicit wrap. GCC is a very popular compiler, and to be honest, I'm still not sure how to get it to warn on the above code, if doing so is possible.
Edit: Just read the rest of the thread, it's -Wconversion, which I suppose makes sense. Ignore me, point taken.
Unfortunately, over the years people baked the semantics of -Wall into their builds so new diagnostics could not be added to that flag.
And clang’s -Weverything shows how the opposite can fail as well
Also there are some warnings that won't be produced if you compile without optimization, because the needed analysis isn't performed.
-Wconversion
https://gcc.gnu.org/onlinedocs/gcc/Warning-Options.html
This is common knowledge for ages. Any cursory Google search returns countless answers.
Take this post made over a decade ago.
https://stackoverflow.com/questions/1730255/gcc-shouldnt-a-w...
x = { size };
The "as" operator is often considered to have been a mistake. Both because of unchecked casts and because of "doing to much".
So I wouldn't be surprised if in the (very) long term there will be a rust edition deprecating `as` casts (after we have alternatives to all cast done with `as`, which are: Pointer casts, dyn casts/explicit coercion and truncating integer casts, for some we already have alternatives for on stable for other not).
And for all who want to not have `as` today you can combine extension traits (which internally still use `as`) + clippy lint against any usage of `as`.
EDIT: I forgot widening integer casts in the list above ;-).
Imagine there's a suitable narrow::<type>() function introduced which has the same consequence, always narrowing, if your data was too wide it may drop important stuff on the floor, and narrow() just says that's too bad.
Rust 2030 can introduce narrow::<type>(), warn for narrowing as usage and then Rust 2035 can error for as. The Rust 2030 -> 2035 conversion software can consume code that does { x as y } and write { x.narrow::<y>() } instead. This code is not better but it's still working in Rust 2035 and this explicit narrow() function is less tempting for new programmers than as IMO.
// a is a UInt64
let a = Int.random(in: 0..<Int.max)
// causes a fatal runtime error if out of range, halts execution
let b = UInt32(a)
// returns a UInt32?, which will be nil if out of range
let c = UInt32(exactly: a)
// another approach for exact conversion:
guard let d = UInt32(exactly: a) else {
// conversion failed
// handle error and return
return
}
// 'd' is a UInt32 without any bits lost
// always succeeds, will return a UInt32, either clamped or truncated. Truncation just cuts off the high bits.
let e = UInt32(clamping: a)
let f = UInt32(truncatingIfNeeded: a)Not exactly the same thing, but in a related area C++ does this a bit.
In C++ you can always still do a c-style cast `(int) some_var` (and the implicit casts obviously), but in general you're meant to use the C++ style explicit casts like `static_cast` and `const_cast`. These are generally tidy, but the most powerful and dangerous of these casts is deliberately awkwardly named as `reinterpret_cast<int>(some_var)` rather than something terse.
It's always easy to spot during a code review.
We consider it best practice to try and organize the types such that explicit casts are minimized.
I thought about this very, very carefully when designing Virgil[1]'s numerical tower, which has both fixed-size signed and unsigned integers, as well as floating point. Like other new language designs, Virgil doesn't have any implicit narrowing conversions (even between float and int). Also, any conversions between numbers include range/representability checks that will throw if out-of-range or rounding occurs. If you want to reinterpret the bits, then there's an operator to view the bits. But conversions that have to do with "numbers" then all make sense in that numbers then exist on a single number line and have different representations in different types. Conversion between always preserve numbers and where they lie on the number line, whereas "view" is a bit-level operation, which generally compiles to a no-op. Unfortunately, the implications of this for floating point is that -0 is not actually an integer, so you can't cast it to an int. You must round it. But that's fine, because you always want to round floats to int, never cast them.
While some folks aren't too fussed about warnings like this, those folks generally aren't writing secure code like kernels. I'm very surprised that kind of conversion was permitted in the code.
https://rust-lang.github.io/rust-clippy/v0.0.212/#cast_possi...
Adding .into() works though, which is the recommended method if the conversion can be statically guaranteed (otherwise try_into should be used, which will become easier in the 2021 edition as the TryInto trait will become part of the prelude).
All explicit conversions, more so fallible, are "frustrating in practice" especially when coming from a language without those foibles.
But given the semantics of usize/isize, it is perfectly reasonable, nay, a good thing, that they're considered neither widenings nor narrowings of other numeric types.
usize is already on that hook: `0x100000000usize` will fail to compile in 32 bit, so you already risk compile errors when switching archs.
As it is I'm just writing `as u64` which is clearly worse.
We don't even use full 64bit pointers today on x64.
push x // 256 bit constant
LOAD
// top of stack now contains storage[x]You could think of checked memory models where half of the 256 bit address is a 128 bit random key needed to access some allocation, or maybe even a decryption key. Similar things are done with the extra space of x64 as well.
Also, 128 bit numbers are still quite uncommon in Rust. Easy conversion of usize to them wouldn't be that useful if conversions of the other number types don't work.
People have said that about pretty much every memory size in the history of computing. The argument for 64-bit is not "2^64 bytes ought to be enough for anyone"; it's "you couldn't use more than 2^64 bytes even if you wanted to". Writing a full register worth of data every clock cycle at 4GHz works out to 32GB/s. (2^64B / 32GB/s) is just over seventeen years, to fill up a 64-bit address space, assuming you're doing no actual computation. Few computers even work for seventeen years without replacement, much less run single processes continually with no reboots.
It would have cause more bugs if they defined it the other way and people used usize to store addresses.
class TimestampedObject implements Comparable<TimestampedObject> { long timestamp; int compareTo(TimestampedObject other) { return (int)(timestamp - other.timestamp); } ... }
ErrorProne catches this: https://errorprone.info/bugpattern/BadComparable
This os why in C it is a good practice to enable all compiler warnings and to have the compiler treat warnings as errors.
If you write for C compiler Foo 8, there's a decent chance Foo 9 will raise a warning which didn't exist before. Now you have to handle "why doesn't this compile" issues and distributions have to patch your sources to do future releases. And that's ignoring bugs like GCC in the past where in some versions you could not satisfy specific warnings.
But never ever leave -Werror enabled for building in Release mode. You'll be preventing your code from building as soon as a new compiler version goes out. Maintainers or code archeologist will have a much worse time than if this option was simply disabled to start with.
IMHO one should always be using C99 types instead of int, but Linux predates that.
Also, shouldn't that implicit conversion cause a compiler warning?
But of course getting C programmers to use these integer types rather than the ones they grew up with isn't easy.
That would lead to a hole in the type sequence (char <= short <= int <= long <= long long) for 64-bit targets (where int is the 32-bit type while size_t is 64 bits).
> IMHO one should always be using C99 types instead of int, but Linux predates that.
On the other hand, on Linux "long" has always been defined as the same size as size_t, so using "long" instead of "int" everywhere could also be an option.
"Always" is a strong way of putting it, there are often times where it makes sense to use the platform's "natural" word sizes (which is the entire point of having `int` `long` `long long` etc.)
In those cases we probably don't care about the full range of the larger types, so it doesn't hurt to use the smallest type for a range of expected values. If it does make a difference, the program will behave differently when compiled on a different arch or even a different compiler.
But maybe "generally" instead of "always". OTOH even I am guilty of using an int to loop over an array.
ILP64 causes a lot of problems, most notably needlessly-increased memory usage and, in C, the inconvenience of requesting a 32-bit type when int is 64-bit. It's rather uncommon to actually need the extra 64-bit range except when describing pointer addresses and memory/disk sizes, both of which benefit from an explicit intptr_t/size_t type for readability if nothing else.
In Haskell I would use an exception, and mark the function as unsafe, but the stdlib seems to disagree with me here.