Quite a few times I was surprised that Rust breaks with some old patterns that were copied over and over in the last 50 years or so. For example: "match" instead of "switch", or the same if/else regardless if it is a statement or a value. These are small touches, but they show attention to detail.
These are good ideas for sure, but they are bread and butter to anyone who is familiar with functional programming. The Lisp and ML families of languages have had them for many decades.
The fact that more languages _aren't_ like this is what's surprising to me, given how effective they are. I love that Rust is bringing them to the masses, but what on earth took the industry so long to accept them?
I guess the answer is that algebraic data types (inductively defined data) tend to be pitted against object-oriented programming (coinductively defined data), and object-oriented programming has dominated the industry for the past 3 decades. Some languages like Kotlin have tried to combine them, but personally I'd rather just embrace the former and relegate the latter to a seldomly-used design pattern, not a programming paradigm hardcoded into the language.
I'm currently designing my own toy language and writing the compiler (to LLVM IR) in Rust.
Representing the AST with Rust's sum types is so simple. Visiting that AST through pattern matching is great. But the "enum" keyword still bugs me.
The way you define product types (tuples, records, empty types) and then their implementation, just awesome. But the "struct" keyword still bugs me too.
It feels "high level" with some quirks.
Then you have references, Box, Rc, Arc, Cell, lifetimes etc... It feels (rightfully) "low level".
Then you have traits, the relative difficulty (mostly for Rust newbies like me) of composing Result types with distinct error types, etc...
It feels "somewhat high level but still low level".
Sometimes you can think only about your algorithm, some other times you have to know how the compiler works. It seems logical for such a language, but still bugs me.
The one thing I hate though, is the defensive programming pattern. I just validated that JSON structure with a JSON schema, so I KNOW that the data is valid. Why do I need to `.unwrap().as_string().unwrap()` everywhere I use this immutable data?
[0]: https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-va...
{
"type": "object",
"oneOf": [{
"required": ["kind", "foobar"],
"properties": {
"kind": {"enum": ["foo"]},
"foobar": {"type": "string"}
}
}, {
"required": ["kind", "barbaz"],
"properties": {
"kind": {"enum": ["bar"]},
"foobar": {"type": "number"},
"barbaz": {"type": "string"}
}
}]
}
Or am I wrong? #[derive(Serialize, Deserialize)]
#[serde(tag = "kind", rename_all = "lowercase")]
enum X {
Foo { foobar: String },
Bar {
#[serde(skip_serializing_if = "Option::is_none")]
foobar: Option<f64>,
barbaz: String
}
}
[0] an enum of numbers would be an issue for instance, though I guess you could always use a `repr(C)` enum it might look a bit odd and naming would be difficult.I think that schema in particular could be represented, though, as:
enum Thing {
foo { foobar: String },
bar { foobar: Option<f32>, barbaz: String },
}Also, JSON schemas allows you to encode semantics about the value not only their types:
{"type": "string", "format": "url"}
That's something I like about Typescript's type system btw: type Role = 'admin' | 'moderator' | 'member' | 'anonymous'
It's still a string, in Rust you would need an enum and a deserializer from the string to the enum.That kinda sounds like you just launched the goalposts into the ocean right here.
> Also, JSON schemas allows you to encode semantics about the value not only their types:
JSON schemas encode types as constraints, because "type" is just the "trival" JSON type. "URL" has no reason not to be a type.
> in Rust you would need an enum
Yes? Enumerations get encoded as enums, that sounds logical.
> a deserializer from the string to the enum.
Here's how complex the deserializer is:
#[derive(Deserialize)]
#[serde(rename_all = "lowercase")]
enum Role { Admin, Moderator, Member, Anonymous }
And the second line is only there because we want the internal Rust code to look like Rust.I come from highly dynamic languages, and even when I was doing C/C++ 10 years ago, I would do more at runtime that what could be considered "best practice".
Right, well, since they're validators anyway, might as well represent them as a defunctionalized validation function or something. Agreed that this is more-or-less past the point where the type system helps model the values you're validating, though a strong type system helps a lot implementing the validators!
> It's still a string, in Rust you would need an enum and a deserializer from the string to the enum.
Yep, though if you really wanted it to be a string at runtime, you could use smart constructors to make it so. The downsides would be, unless you normalized the string (at which point, just use an enum TBH), you're doing O(n) comparison, and you're keeping memory alive, whether by owning it, leaking it, reference counting, [...].
Thankfully due to Rust's #[derive] feature, the programmer wouldn't need to write the serializer/deserializer though; crates like strum can generate it for you, such that you can simply write:
use strum::{AsRefStr, EnumString};
#[derive(AsRefStr, EnumString, PartialEq)]
enum Role {
Admin,
Moderator,
Member,
Anonymous,
}
fn main() {
assert_eq!(Role::from_str("Admin").unwrap(), Role::Admin);
assert_eq!(Role::Member.as_ref(), "Member");
}
(strum also has also a derive for the standard library Display trait, which provides a .to_string() method, but this has the disadvantage of heap allocating; EnumString (which provides .as_ref()) compiles in the strings, so no allocation is needed, and .as_ref() is a simple table lookup.)Nit: it's more constraining but serde can deserialize to an &str, though that assumes the value has no escapes.
Ideally `Cow<str>` would be the solution, but while it kind-of is, that doesn't actually work out of the box: https://github.com/serde-rs/serde/issues/1852
But yeah, I tend to do more work at runtime than compile-time, which is not really the way to go in Rust.
Everything else is named so well and designed so well that this really sticks out.
foo.expect("already validated") foo.expect("There was a problem and I had to crash")
That _reads_ as "I expect there to be a problem and have to crash". But it _means_ something totally different. It means: There should not be a problem, but if there were a problem
I would crash and print out this message.
And the way to write that in readable code is foo.unwrap_or_else(|| panic!("There was a problem and I had to crash"))
I'm sorry if I'm misunderstanding -- if I am can you explain to me how? Because I feel like my view is diametrically opposed to yours but it's really not that subjective a matter; only one of us can be right.> That reads as "I expect this to be already validated". Is that what you intended it to mean?
Yes
> foo.expect("There was a problem and I had to crash")
> That _reads_ as "I expect there to be a problem and have to crash". But it _means_ something totally different. It means:
Absolutely. That's why I'd phrase it foo.expect("foo invariant upheld")
thread 'main' panicked at 'foo invariant upheld'
And everyone, except those who are inured to the illogic via experience, will read that as thread 'main' panicked because 'foo invariant upheld'
especially new users, less experienced programmers, and non-Rust programmers seeing the panic message.Now if the Rust language made that error message say
thread 'main' panicked at failed expectation for 'foo invariant upheld'
that might be cool. But it does not do that. It feels like you're bending over backwards trying to excuse what is clearly a wart in a beautiful language. I'm happy to continue the debate though, I am fascinated to know how I can be wrong here :)So that's my main point, that your strategy doesn't work due to the error message. But a secondary point is that Rust documentation encourages people to use it opposite to what you suggest: I mean, just look at the first example in the Rust book. It's completely nonsensical:
let f = File::open("hello.txt").expect("Failed to open hello.txt");
I know there was some controversy about this early on and somehow it didn't get reverted; it would be more honest if the docs would at least admit that it's "a little strange".https://doc.rust-lang.org/book/ch09-02-recoverable-errors-wi...
Playground if you want to remind yourself of the psychological experience of seeing that error message at run time with minimal effort.
https://play.rust-lang.org/?version=stable&mode=debug&editio...
(a) If you write the message the way you suggest the printed error message is misleading, and
(b) The Rust community does not encourage using .expect() in the way you suggest; in particular the canonical Rust documentation source instructs us to use it as
foo.expect("A thing I do not expect") let hello_1 = echo_hello.output().expect("failed to execute process");- Pattern matching via "match"
- if being an expression (among most things)
Those are both found in Scala.
[1]: https://pingcap.com/blog/rust-compilation-model-calamity
They are awesome but they make compile-time arbitrarily bad. The D compiler is roughly as fast as the Go one. However, D has macros though and that makes it very slow to compile sometimes.
The alternative to macros are code generators. Works fine for bigger stuff like a parser generator but not for smaller stuff like a regex.
It's possible that Zig's approach helps here -- since the metalanguage is just the language, you can take some of your intuition about performance along to compile time. And the macro language is not weirdly restricted, so you can write something you're more used to, with similar idioms.
In the limit this is clear: I'm using several Python code generators C++ in https://www.oilshell.org, and it's easy to reason about the performance of Python. Doing the same in C++ metaprogramming would almost certainly be a lot slower. It would also give me less indication of code bloat; right now I simply count the lines of output to test whether my code generator is "reasonable".
e.g. I generate ~90K lines of code now, which is reasonable, but if I had 1M or 10M lines of code, that wouldn't be. But there's not that much of a sanity check on macros (except binary size).
I think of one as metaprogramming with the parser and the other as metaprogramming with an interpreter. (And the C preprocessor is metaprogramming with only a lexer. Code generation is the kind of metaprogramming that every language supports :) )
Although maybe you're saying Zig doesn't have the functionality of Rust macros, and that could be true; I haven't played with it enough.
However I do think there is overlap as Zig implements printf with compile-time evaluation and Rust does it with macros:
https://ziglang.org/documentation/master/#Case-Study-printf-...
https://andrewkelley.me/post/zig-programming-language-blurs-...
Anyway if you want to know why your compile time is slow in a given zig project, you can probably get pretty far by grepping for calls to that builtin.
The question is whether you do the meta programming/code generation within the language or as external tool. Code bloat is an issue in both cases.
Macros make it easy to write code which generates code which generates code which... This enables some wonderful use cases, often around generating type declarations. If done with an isomorphic language, you can also reuse the same code for compile time and runtime implementations with just thin wrappers.
External code generators however will build faster because the build system takes care of reusing the intermediate code.
Not according to this benchmark:
https://github.com/nordlow/compiler-benchmark
dmd performed around 1.5x to 4x faster than go.
My only experience with Rust so far has been trying to learn it by writing applications in it and also use 3rd party CLIs, but quickly loosing interest because the "change <> try out change" cycle has been too slow and cumbersome, and installing/compiling dependencies take fucking forever, even on a i9-9900K.
An important thing though, if you aren't doing this already, is to not wait for a full build to know if your types check out. You can use cargo-check if you prefer (https://doc.rust-lang.org/cargo/commands/cargo-check.html), but really I recommend using an editor with immediate feedback if at all possible. rust-analyzer (an LSP) is one of the best, and should be available even if you're on Vim or something.
Using Rust without snappy editor hints is fairly miserable because of how interactive the error feedback loop tends to be. If you don't rely on a full build for errors - just for actual testing - I find the build times to be perfectly livable (at least in the smallish projects I've done).
[1] https://bevyengine.org/learn/book/getting-started/setup/#ena...
Maybe you're comparing apples to oranges here. I've worked professionally with both Rust and Go for years now on a variety of real world projects, and I've never seen a similarly-sized Go codebase that compiles slower than a Rust one. If you're comparing incremental Rust compilation to first-time Go compilation, maybe they could be competitive, but... Rust is incredibly slow at compilation, even incremental compilation.
Yes, using lld can speed up Rust compilation because a lot of the time is often spent in the linker stage, but... that's not enough to make it as fast as Go.
YMMV, of course, but... my anecdotal experience would consider it disingenuous to say that Rust compile times are an advantage compared to Go, and I'm skeptical that Rust compile times are even an advantage compared to the notoriously slow webpack environments.
Rust is good at many things, but compilation speed is not one of them. Not even close, sadly. "cargo check" is tolerable most of the time, but since you can't run your tests that way, that's not actually compilation.
It is absolutely apples to oranges, but if you just care about the everyday local workflow and ability to iterate and test, it's close enough most of the time to not be much of a problem in either case.
I'm sure this experience isn't guaranteed for all codebases, and it certainly helps that I make heavy use of crates, which would minimize the work required during incremental compilation. Though I'm not actively going out of my way to optimize for incremental compilation really, beyond the config linked above.
Plus, Kubernetes is sitting at 5 million lines of Go code, by my count. Try compiling a 5 million LoC Rust code base... I won't wait around.
I would assume that a controller written in Rust bypasses all the complex legacy of the Kubernetes code base, and that's why it can compile faster. If someone made a similar project in Go[0] to write Kubernetes controllers in Go without depending on the mega-Kubernetes code base, I'm sure it would be incredibly faster at compiling than the Rust version.
Beyond that, I've heard that the Kubernetes codebase is internally just a nightmare of basically untyped `interface{}` stuff floating around everywhere, which would make the development experience subpar. I don't know how much this is exposed to custom controllers.
So, if Kubernetes is your only experience with Go... I'm sorry you've had to experience that. It's a product that people seem to agree is functional and works most of the time, but I can't remember hearing any positive experience from people working on it. From what I understand, it was originally prototyped in Java, and then hastily rewritten into Go before public release, and I'm sure that didn't help things.
[0]: conceptually, maybe something like this? https://github.com/ericchiang/k8s or an updated fork of it like this: https://github.com/karlmutch/k8s No idea how well either works, if at all.
About on par with C/C++ and Go. In that you don't have one and don't want for one. REPL driven development is difficult with languages like Rust, both to implement and use.
I think there are some projects floating around out there, but I personally don't see a purpose for one. It's not python or matlab.
I was thinking of a REPL in the sense of common Clojure usage (https://vvvvalvalval.github.io/posts/what-makes-a-good-repl....) not the basic "write lines into a separate program and then copy-paste it into your source code" that Python offers.
Notebooks work better than raw REPL's for a language that's so heavily based on static typing, but they're idiomatically quite similar.
On the other hand iterative development with rust analyzer going and all the dependencies already built is pretty painless. By the time you run your tests or program, it's likely to take only a few seconds to build.
That said you can write terribly long to compile rust. Usually there's some trades you can make to weigh compile time as more important than flexibility or static code paths.
I've written a sizeable ad server in rust with maybe 25kloc. Where I was the sole developer and sys ops person, it cost me a little upfront time but saved me many more hours in operations work.
In my experience, you write some code, the rust analyzer (which is easily embedded in an IDE like VSCode or IntelliJ - which has its own "analyzer" I think) gives you immediate feedback, so you know immediately if things compile or not (there are a few edge cases the analyzer might miss, so when you actually run "rustc" something doesn't compile, but it's pretty rare)... you then write a little test , and Rust has many ways of letting you do that (unit tests right into the same file as the code being tested, integration tests which let you use the code as if from another crate, and even doctests, which are like unit tests but embedded in the documentation of your code)... running the tests is a matter of pressing a button and waiting a few seconds (compilation + test runtime) normally, unless you change dependencies between runs as that requires downloading/compiling your code AND the dependencies, which can be very slow (dozens of seconds)... which is the same problem as with a fresh build, which will almost certainly run in the minutes because of the necessary local compilation of all dependencies... but the experience is not very different from something like Java or Kotlin IMO (but definitely a slower cycle than Go, for example).
But once they're there, the result is extremely advanced tooling that other language ecosystems have taken years to develop. Things like IDEs, debuggers, static analyzers, superoptimizers, fuzzers, verification tools, build systems, language interop adapters, and code generators, are all examples of things that can take advantage of advanced type systems to develop strong capabilities extremely easily.
I think Rust is starting to show signs of such benefits, and I think in 10 years we'll all be looking back at it with surprise that anyone ever doubted it.
I'm curious, could you elaborate?
foo(x : int)
Therefore, one would expect to annotate the return type as,
foo(x : int) : string
Since the pattern is showing foo applied to x. The Rust syntax is actually confusing for both Haskell/ML programmers (where the arrow comes from) and mainstream programmers. It's too small an issue to change now though.
Rust's support for proper "algebraic data types" is very good and gives it an advantage over languages like C++. However there are some small surprises, such as forcing all enum constructors/fields to be public (one must therefore wrap it to make an abstract data type).
Every language has its warts and these are particularly minor ones.
Haskell doesn't use that notation either, it uses -> both for the parameter list and for the return type, and separates argument names (arguably, these are poor choices, since currying is not an efficient CPU-native operation and not intuitive so distinguishing between multiple arguments and returning closures is useful, and argument names are useful for documentation).
But foo(x : int) is a string! It literally reads "foo applied to x". In the function definition, it appears to be used as a left-hand-side pattern which is "matched". The definition is written as if to say, whenever the term foo(x) is encountered, use this definition here. At least, that was my expectation.
> Haskell doesn't use that notation either
OCaml does and Haskell once had a proposal to add it. Haskell type signatures are normally written separately, but it does support annotating patterns with the right extensions.
fn foo(x: int) -> ReturnType {
body
}
The ast here splits into Function {
name: foo
signature: (x: int) -> ReturnType
body: body
}
I.e. the arrow binary op binds more tightly than the adjacency between foo and x: int. And the type of foo is a function, not a string.A "better" way to write this (in that it breaks down the syntax into the order it is best understood) might be
static foo: (Int -> ReturnType) = {
let x = arg0;
body
}
Or to put it another way. Reading foo(x: int) as "foo applied to x" in this case is a mistake, because that's now how things bind. You should read that "foo is a (function that takes Int to String)". It's a syntactic coincidence that foo and x are beside eachother, nothing more.It's mixing up assigning a global variable, specifying that variables type, and destructuring an argument list into individual arguments, in one line. I've played at making my own language, and this is one part that I've never been satisfied with.
Personally I'd probably at least go with a `foo = <anonymous function>` syntax to split out the assigning part. But that's spending "strangeness budget" because that's not how C/Python/Java do it, and I can understand the decision to not spend that budget here...
Can you say so decisively for a language with first-class functions?
foo :: (Int, Int) -> Int
foo (x, y) = ...
On the other hand the enum thing is certainly surprising.
In your example, why bother naming the inner "x" variable for the function param? It cannot be used on the right-hand-side (definition of "bar"). For that reason, the notation is not exactly "clear". In OCaml the annotation would be:
bar( x : int -> string )
Ocaml's syntax is more consistent, I agree, but its colon operator has different precedence than in Rust, so I am not sure its rational applies to Rust.
> foo(x : int) : string
> since the pattern is showing foo applied to x.
kind of like the C/C++ "declaration mirrors use" thing, i.e.
int *foo;
which means"the result of dereferencing `foo` is of type `int`"
which is not the same as saying
foo : Ptr<int>;
because in the former, you're kind of describing what `foo` is a without actually saying it, if that makes sense.i find that way of specifying types counterintuitive in both C/C++ and MLs.