IIRC a secret 'other' field (or '__non_exhaustive' or something) is actually how we did thing before non_exhaustive was introduced.
IIRC a secret 'other' field (or '__non_exhaustive' or something) is actually how we did thing before non_exhaustive was introduced.
> The word “other” means “not mentioned elsewhere”, so the presence of an Other logically implies that the enumeration is exhaustive.
In Rust, because all enums are exhaustive by default and exhaustive matching is enforced by the compiler, there is no risk of this sort of confusion. And then the fact that his proposed solution is:
> Just document that the enumeration is open-ended
The non_exhaustive attribute is effectively compiler-enforced documentation; users now cannot forget to treat the enum as open-ended.
Of course, adding non_exhaustive to Rust was not without its own detractors; it usage for any given enum fundamentally means shifting power away from library consumers (who lose the ability to guarantee exhaustive matching) and towards library authors (who gain the ability to evolve their API without causing guaranteed compilation errors in all of their users (which some users desire!)). As such, the guidance is that it should be used sparingly, mostly for things like error types. But that's an argument against open-ended enums in general, not against the mechanisms we use to achieve those (which, as you say, was already possible in Rust via hacks).
It's just that with #[non_exhaustive], you must specify a default branch (`_ => { .. }`), even if you've already explicitly matched on all the values. The idea being that you've written code which matches on all the values which exist right now, but the library author is free to add new variants without breaking your code - since it's now your responsibility as a user of the library to handle the default case.
https://doc.rust-lang.org/rustc/lints/listing/allowed-by-def...
Let's say I am using a crate called zoo-bar. Let's say this crate is not using non-exhaustive.
In my code where I use this crate I do:
let my_workplace = zoo_bar::ZooBar::new();
let mut animal_pens_iter = my_workplace.hungry_animals.iter();
while let Some(ap) = animal_pens_iter.next() {
match ap {
zoo_bar::AnimalPen::Tigers => {
me.go_feed_tigers(&mut raw_meat_that_tigers_like_stock).await?;
}
zoo_bar::AnimalPen::Elephants => {
me.go_feed_elephants(&mut peanut_stock).await?;
}
}
}
I update or upgrade the zoo-bar dependency and there's a new enum variant of AnimalPens called Monkeys.Great! I get a compile error and I update my code to feed the monkeys.
diff --git a/src/main.rs b/src/main.rs
index 202c10c..425d649 100644
--- a/src/main.rs
+++ b/src/main.rs
@@ -10,5 +10,8 @@
zoo_bar::AnimalPen::Elephants => {
me.go_feed_elephants(&mut peanut_stock).await?;
}
+ zoo_bar::AnimalPen::Monkeys => {
+ me.go_feed_monkeys(&mut banana_stock).await?;
+ }
}
}
Now let's say instead that the AnimalPen enum was marked non-exhaustive.So I'm forced to have a default match arm. In this alternate universe I start off with:
let my_workplace = zoo_bar::ZooBar::new();
let mut animal_pens_iter = my_workplace.hungry_animals.iter();
while let Some(ap) = animal_pens_iter.next() {
match ap {
zoo_bar::AnimalPen::Tigers => {
me.go_feed_tigers(&mut raw_meat_that_tigers_like_stock).await?;
}
zoo_bar::AnimalPen::Elephants => {
me.go_feed_elephants(&mut peanut_stock).await?;
}
_ => {
eprintln!("Whoops! I sure hope someone notices this default match in the logs and goes and updates the code.");
}
}
}
When the monkeys are added, and I update or upgrade the dependency on zoo-bar, I don't notice the warning in the logs right away after we deploy to prod. Because the logs contain too many things no one can go and read everything.One week passes and then we have a monkey starving incident at work.
After careful review we realize that it was due to the default match arm and we forgot to update our program.
So we learn from the terrible catastrophe with the monkeys and I update my code using the attributes from your link.
diff --git a/src/main.rs b/src/main.rs
index e01fcd1..aab0112 100644
--- a/wp/src/main.rs
+++ b/wp/src/main.rs
@@ -1,3 +1,5 @@
+#![feature(non_exhaustive_omitted_patterns_lint)]
+
use std::error::Error;
#[tokio::main]
@@ -11,6 +13,7 @@ async fn main() -> anyhow::Result<()> {
let mut animal_pens_iter = my_workplace.hungry_animals.iter();
while let Some(ap) = animal_pens_iter.next() {
+ #[warn(non_exhaustive_omitted_patterns)]
match ap {
zoo_bar::AnimalPen::Tigers => {
me.go_feed_tigers(&mut raw_meat_that_tigers_like_stock).await?;
@@ -18,8 +21,12 @@ async fn main() -> anyhow::Result<()> {
zoo_bar::AnimalPen::Elephants => {
me.go_feed_elephants(&mut peanut_stock).await?;
}
+ zoo_bar::AnimalPen::Monkeys => {
+ // Our monkeys died before we started using proper attributes. If they are hungry it means they have turned into zombies :O
+ me.alert_authorities_about_potential_outbreak_of_zombie_monkeys().await?;
+ }
_ => {
- eprintln!("Whoops! I sure hope someone notices this default match in the logs and goes and updates the code.");
+ unreachable!("We have an attribute that is supposed to tell us if there were any unmatched new variants.");
}
}
}
And next time we update or upgrade the crate version to latest, another new variant exists, but thanks to your tip we get a lint warning and we happily update our code so that we won't have more starving animals. diff --git a/wp/src/main.rs b/wp/src/main.rs
index aab0112..4fc4041 100644
--- a/wp/src/main.rs
+++ b/wp/src/main.rs
@@ -25,6 +25,9 @@ async fn main() -> anyhow::Result<()> {
// Our monkeys died before we started using proper attributes. If they are hungry it means they have turned into zombies :O
me.alert_authorities_about_potential_outbreak_of_zombie_monkeys().await?;
}
+ zoo_bar::AnimalPen::Capybaras => {
+ me.go_feed_capybaras(&mut whatever_the_heck_capybaras_eat_stock).await?;
+ }
_ => {
unreachable!("We have an attribute that is supposed to tell us if there were any unmatched new variants.");
}
But what was the advantage of marking the enum as #[non_exhaustive] in the first place?Each option has its place, it depends on context. Does the creator of the type want/need strictness from all their consumers, or can this call be left up to each consumer to make? The lint puts strictness back on the table as an opt-in for individual users.
#[derive(serde::Serialize, serde::Deserialize)]
pub struct SomeEnum {
AValue,
BValue,
}
My customers use that and all is well. But I want to add a new enum value, CValue. I can't require that all my customers update their version of my Rust client before I add it; that would be unreasonable.So I add it, and what happens? Well, now whenever my customers make that API call, instead of getting some API object back, they get a deserialization error, because that enum's Deserialize impl doesn't know how to handle "CValue". Maybe some customer wasn't even using that field in the returned API object, but now I've broken their code.
Adding #[non_exhaustive] means I at least won't break my customers' code when I add a new enum value.
This allows you to do something like:
#[derive(Clone, Copy)]
#[repr(u8)]
#[non_exhaustive]
pub enum Foo {
A = 1,
B,
C,
}
impl Foo {
pub fn from_byte(val: u8) -> Self {
unsafe { std::mem::transmute(val) }
}
pub fn from_byte_ref(val: &u8) -> &Self {
unsafe { std::mem::transmute(val) }
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn conversion_copy() {
let n: u8 = 1;
let y = Foo::from_byte(n);
assert!(matches!(y, Foo::A));
let n: u8 = 4;
let y = Foo::from_byte(n);
assert!(!matches!(y, Foo::A) && !matches!(y, Foo::B) && !matches!(y, Foo::C));
let n2 = y as u8;
assert_eq!(n2, 4);
}
#[test]
fn conversion_ref() {
let n: u8 = 1;
let y = Foo::from_byte_ref(&n);
assert!(matches!(*y, Foo::A));
let n: u8 = 4;
let y = Foo::from_byte_ref(&n);
assert!(!matches!(*y, Foo::A) && !matches!(*y, Foo::B) && !matches!(*y, Foo::C));
let n2 = (*y) as u8;
assert_eq!(n2, 4);
}
}
This lets you have a simple fast parsing of types without needing a bunch of logic - particularly in the ref example. Someone else sent you data over the wire and is using a vendor defined value, or a newer version of the protocol that defines Foo::D? No big deal, you can igore it or error, or whatever else is appropriate for your case.If you want to define Reserved and Vendor as enum attributes, now you have to have logic that runs all the time - and if you want to preserve the original value for error messages, logs, etc - you can't Repr(u8) and take up more memory, have to do copies, etc.
#[non_exhaustive]
pub enum Foo {
Undefined =0,
A = 1,
B,
C,
Reserved(u8),
Vendor(u8),
}
impl Foo {
pub fn from_byte(val: u8) -> Self {
match val {
0 => Foo::Undefined,
1 => Foo::A,
2 => Foo::B,
3 => Foo::C,
4..=127 => Foo::Reserved(val)
128.. => Foo::Vendor(val)
}
}
}
You also need logic to convert back to a u8 now too.It's not strictly necessary, but it certainly makes some things far more ergonomic.
Apologies for my pre-coffee brainfarts.
#[non_exhaustive] is most popular for the variants of an enumeration but is permissible for published structure types (it means we promise these published fields will exist but maybe we will add more and thus change the size of the structure overall) and for the variants of a sum type (it means the inner details of that variant may change, you can pattern match it but we might add more fields and your matches must cope)
For a network protocol or C FFI you probably want a primitive integer type not any of Rust's fancier types such as enum, because while you might believe this byte should have one of six values, 0x01 through 0x06 maybe somebody decided the top bit is a flag now, so 0x83 is "the same" as 0x03 but with a flag set.
Trying to unsafely transmute things from arbitrary blobs of data to a Rust type is likely to end in tears, this attribute does not fix that.
You can practically use it today by gating on a nightly-only cfg flag. See https://github.com/guppy-rs/guppy/blob/fa61210b67bea233de52c... and https://github.com/guppy-rs/guppy/blob/fa61210b67bea233de52c...
I think half of it is developers presuming to know users' needs and making decisions for them (users can make that decision by themselves, using the default case!) but also a logic-defying fear of build breakage, to the point that I've seen developers turn other compile errors into runtime errors in order to avoid "breaking changes".
You have to opt into it but it's nice that it's available.
If I have a parameter in my public API that has enumerated options, I should be able to add a new option without needing to bump my semver major version number since downstream existing code obviously isn't going to use it yet. If downstream was using my public api's enum for some of their own book keeping and so matched on my enum, I want to reserve the right to to say that that is non-public use of my enum, hence the idea that exhaustiveness in enums is a separate decision on to what is included in a public API or not.
On the other hand, if I introduce a new variant in a return value and existing code will get it and need to actually do something with it, then it should probably be breaking. Errors are somewhat of an exception to this since almost all error enumerations need a general, "unknown error" category anyways and introducing a new variant is generally elevating one case out of that general case. Obviously authors can make mistakes.
The alternative, when you cannot mark non_exhaustive, is to introduce stringly typed catch alls, which is much less desirable for everyone.
Swift allows a ‘default’ enum case which is similar to other but you should use it with caution.
It’s better to not use it unless you’re 110% sure that there will not be additional enums added in the future.
Otherwise, in Swift when you add an additional enum case, the code where you use the enum will not work unless you handle each enum occurrence at it’s respective call site.
The issue with traditional “default” cases is that they shadow warnings/errors about unhandled cases, but you’d still want to have some form of default case for forward compatibility.
Separate compilation is a technical implementation detail that shouldn't have an impact on semantics. Especially since LTO (link time optimisation) is becoming more and more common; 'thin' LTO is essentially free in Rust at least in terms of extra build time. LTO blurs the lines between separate compilation units.
On the flip side, Rust can use multiple codegen units even for the same crate, thus introducing separate compilation where a naive approach, like in classic C, would only use a single one.
In other words, you want to ensure that you have the most appropriate behavior for whatever values are currently known, and a fallback behavior for the future values that by definition you can’t possibly know at the present time. Of course, this is more or less only practical in languages where the interface version you compile against is only updated deliberately, while the implementation version at runtime can be any newer compatible version.
Swift also has two versions of a `default` case in switch statements, like you described. It has regular `default` and it has `@unknown default`. The `@unknown default` case is specifically for use with non-frozen enums, and gives a warning if you haven't handled all known cases.
So with `@unknown default`, the compiler tells you if you haven't been exhaustive (vs. the current API), but doesn't complain that your `@unknown default` case is unreachable.
#[non_exhaustive]
pub enum Protocol {
Tcp,
Udp,
Other(u16),
}
It allows you to still match on the unrecognized case (like `Protocol::Other(1)`, which is nice), but an additional enum variant may eliminate that case, if our enum gets extended to: #[non_exhaustive]
pub enum Protocol {
Tcp,
Udp,
Icmp,
Other(u16),
}
Even though we can add additional variants in a semver-nonbreaking way due to `#[non_exhaustive]`, other people's code may now be broken until they've changed `Protocol::Other(1)` to `Protocol::Icmp`.Having had this in the back of my head for quite some time, I think instead of an `Other` case there should be two methods, one returns an `Option<Protocol>` and the other one returns the `u16` representation. Unless there's a match on one of your expected cases your default branch would inspect the raw numeric type, which would keep working even if that case is added to the enum.
F# I believe is similar wrt discriminated unions and pattern matching
By default Rust expects you to handle every enum variant. Not doing so would be a compile error.
An example - my library exposes enum Colour with 3 variants - Red, Blue Green. In your application code you `match` on all 3. So far so good. But now if I add a 4th colour to my enum, your code will no longer compile because you are no longer handling every enum variant. This is a crappy experience for the user of the library.
Instead, the library writer can make their intent clear - with the #[non_exhaustive] attribute. On such an enum it's not enough to handle the 3 colours of the enum, you must add a wildcard matcher that matches any variants added in future. This gives the library writer flexibility to make changes, while protecting the application developer from breakage.
Thanks for taking the time to explain!