Rust for JavaScript Developers – Pattern Matching and Enums
sheshbabu.com
sheshbabu.com
I've seen a lot of attempts on these teams to ensure these mistakes aren't made (special parsing libraries, using default fallback values, etc.), but someone always manages to sneak in an enum somewhere where they shouldn't and then something blows up 6 months later when a new value is introduced. Enthusiastic engineers will spend a lot of time trying to figure out how to add an over-strict enum, because hey, who doesn't want type safety?
Sorry to say but JavaScript is the only ecosystem I've seen where people blithely `JSON.parse` anything they get over the network and directly access the JSON structure expecting everything they want to be there without checking.
Enum types are supposed to be strict. Decoding raw data into enums can always fail but the failure should be handled at a higher level.
The way protobuf handles enums, it recommends having an "UNKNOWN" as value 0, as when deserializing from the wire any unknown enum value gets turned into value 0.
So it's recommended to have something like
enum FailReason { UNKNOWN_ERROR = 0, ERROR_1 = 1, ERROR_2 = 2, etc }
But someone long ago made the enum I was dealing with look like this
enum { SUCCESS = 0, ERROR_1 = 1, ERROR_2 = 2, etc. }
But this means I can't easily add a new error case, as it'll get interpreted as a success unless I carefully roll out the change to the receiving servers first.
But if I wasn't dealing with a distributed system I would like to have exhaustive checking.
I guess one option is to replace the endpoints that use the enum, but that seems heavy handed if it's wide spread.
Also, I'm not sure if you understand Rust enums. They should probably be called union types, because an enum in Rust signifies that the value can be one of a number of different types. This is extremely useful for code which has to handle a few different things differently. It's much more useful than the standard concept of an enum in languages such as Java and Rust, where it's just a dressed up integer flag.
Imagine you have two servers communicating with
enum Packet { Ping, Pong, }
Where the semantics are that if you receive a ping, you say pong, and if you see pong, you say ping.
Suppose we deploy this out to several servers, with code built on 2020-7-12. Now on 2020-07-13 you commit to head a new enum value `Stats`, which returns a new kind of reply.
If you send outgoing requests in the same 2020-07-13 build, you're going to have a bad time, as your fleet is going to start at 100% 2020-07-12 and switch over to 100% 2020-07-13, but during that switch you'll have some 50% 50% split. If a new server sends a request to an old one, you'll have a crash.
Instead you have to add receivers in 2020-07-13, but not use it yet. Roll it out, make sure it's past how far you might roll back, then you can deploy a 2020-07-20 build which makes calls using the new enum.
If you skip that step, you're going to have a bad time. Maybe you want to instead have a default case handler instead of crashing, but that means each of your enums has to know about "UNKNOWN" and you're deserialization error has to be future proof.
I wrote the Haskell implementation. It maps unions to variants (what Rust calls enums), BUT it always generates an extra 'unkown' variant, which gets used whenever the variant is one not recognized by the generated code.
In this case the value proposition of variants/enums, including exhaustiveness checking, is still really useful -- it can not only deal with the possibility of more variants added in the future, it forces you to handle that possibility, which imo is the best of both worlds.
But yes, in general, when modeling data that comes from the outside world, you need to about overfitting.
Vaguely related: https://lexi-lambda.github.io/blog/2020/01/19/no-dynamic-typ...
I deal a lot with this, because I interface with software outside my control. In the case of Rust I find super easy to solve, just use this litte trick:
https://serde.rs/attr-flatten.html
#[derive(Serialize, Deserialize)] struct User { id: String, username: String,
#[serde(flatten)]
extra: HashMap<String, Value>,
}
Serde is very good. Near all you want is there.Why does Rust destructuring require the struct name? The example uses:
let Person { name, city } = person;
This seems a little verbose since I would expect the type system to know that person is of type Person already.Is this because destructuring syntax is trying to be compatible with the pattern matching syntax?
It isn't in the language yet due to a combination of subtle issues that need to be considered before a final design can be reached, trying to keep the grammar clean, accepting that explicitness and verboseness now is better than hastily moving towards something that seems better but that can introduce technical, documentation or understandability issues.
> Is this because destructuring syntax is trying to be compatible with the pattern matching syntax?
It's not trying to be compatible, it is pattern matching syntax. If the language changed to, for example, allow the following
let _ { name, city } = person;
Would also allow match opt_person {
Some(_ { name, city }) => {}
None => {}
}
[1]: https://internals.rust-lang.org/t/pre-rfc-struct-constructor...[2]: https://internals.rust-lang.org/t/pre-rfc-implied-type-for-p...
[3]: https://internals.rust-lang.org/t/pre-rfc-anonymous-struct-a...
fn takes_person(p: Person) {
fn takes_person(Person {name, city}: Person) {
But I wouldn't inherently say that's the reason, it's also because structs are nominally rather than structurally typed.I also appreciate the lesson of converting discriminated unions in JS and to rust - that’s my favorite pattern from the TypeScript community.
enum Foo {
StructVariant { named: i32, fields: String, with: (), types: usize },
TupleVariant(i32, String, usize), // only types and position information
UnitVariant, // just a tag, no other days being carried
}
The struct variant is useful when you are carrying enough different fields that you want to name them.With WASM becoming mainstream and the growing maturity of the Rust ecosystem, it appears to be a solid choice for backend systems and potentially emerging frontend use cases.
WASM is not needed for the backend, so that is not the reason either. For frontend, I can see it for some pages (I used it myself), yes, but there are also other languages that can target WASM.
You also mention maturity of Rust, but other languages in its domain are more mature, so that cannot be it either.
Mixing languages does not require WASM in native code either. It has always been done through the C ABI.
I don't know how to import a ES6 module using C ABI.
Then you cannot use any modern computer in your company. Much less hypercomplex systems like Node, v8 and JITs in general.
> compared to the multi-million risks
What "risks"? What standards are you following? Insurance?
> WASM execution can be optimized a lot as it goes past MVP.
Citation needed. There is no magic.
> I don't know how to import a ES6 module using C ABI.
So that is why you want WASM?
I don't know what magic you're talking about - we're simply waiting for WASM to standardize GC, SIMD, tail calls etc, you can read more here[0]; and to get the execution optimized like V8 did over the past 10 years[1] (WASM on V8 gets most of these gains by default as a bonus).
I already said why I want WASM: because it provides a good sandbox and opens most ecosystems through safe and typed interop. I don't care about ES modules in particular - I care about ES modules, Ruby modules, Python modules, Prolog modules, Fortran modules, C modules, Rust modules, AssemblyScript modules, COBOL modules, Erlang modules, F# modules, Java modules, and so on - at once. Finally a future where I can use and trust any library and not get limited to the language I use (if I don't want to spend most of my time fighting and securing C ABI interop, if at all humanly possible) is there.
When I decided I wanted to move away from Ruby, I didn’t even consider getting deeper into Python. Why would I? They’re so similar that the cost/benefit ratio is off, at least for me back then.
YMMV of course.
Oh and finally, not every bit of writing has to be extremely useful or hit a wide audience. Sometimes people write things because they want to.
The reason is that otherwise you always try to map things back to your previous knowledge and expertise, which can make you biased or give you tunnel vision.
"If you have a hammer..."