[1] unless you liberally use `union`s everywhere which is very unidiomatic C++, but can be done if desired.
For what it's worth, Python's type hints are also Turing complete (https://arxiv.org/abs/2208.14755, discussed on HN at https://news.ycombinator.com/item?id=32779296)
https://developers.redhat.com/blog/2017/11/16/speed-python-u...
- nobody uses it in the ecosystem. As outlined in the article, a lot of value of Option/Result is derived from their pervasiveness in the Rust ecosystem. C++ is far from this
- the ergonomics of it are terrible: no pattern matching, structural variants instead of named variants (yes you can emulate that with wrapper types but meh), lambda-oriented matching means you cannot as easily do things like early returns, statement-oriented language limits the usefulness anyway, lack of combinators, and for error handling specifically, lack of `?` operator
- performance is dubious. I had very steep and unexpected performance cliffs when lambda inlining started to fail for some reason. Having sum types be a language construct guarantees we're not relying on things like lambda optimisation here.
Error types are probably the most popular incarnation of that. There’s several libraries available and they will be part of the C++ standard.
This is a case of the Rust community overselling a minor feature as a game-changing novelty.
I think you misunderstood me though if you say I haven’t used them properly. I just didn’t need them or find them that useful because it rarely happens that I want to model a type which has several states.
Specific variant types like std::optional or the Result-equivalents are useful, but not that game changing either. They’re nice I suppose.
I wonder what kind of software you write that such low-level coding idioms make a big difference to the end result. Or what do you mean by game changing?
You can't use what you don't have, so you adapt to the tools you do have. In my C++ time, the team would often write types that logically held several variants. However they were expressed as product types, so with space overhead and error-prone, unergonomic use.
Sum types are game-changing. Look at any Rust project you'll find enums with data everywhere. They're just a building block to model problems, like product types are. A language that miss them is as strange too me as a language without product types. After 10 years of mostly C++11 and C++14 I would never go back to it for this reason alone (although as outlined in the article there are other reasons too).
> Error types are probably the most popular incarnation of that. There’s several libraries available and they will be part of the C++ standard.
Again, with what performance and ergonomics?
In matcher.rs, the enum is used as poor man’s OOP. I saw several match expressions which then call the same function on each matched type.
searcher/mod.rs is an error definition.
glob.rs contains a sort of policy enum which can be implemented once again with OOP or as a policy template.
json.rs usage can be modeled as a single class.
core/app.rs is more involved, but can be modeled as a series of structs with a map from enum -> any. Or as am std::variant. Or using OOP.
I looked at all instances and didn’t see anything game-changing. It’s a nice syntax and it should have really good performance, but such idioms are way too low-level to change any game.
As should be patently fucking obvious to anyone who has been commenting on a technology web site for as long as you have, sum types are a tool. They are a tool for expressing clearly and concisely the idea that a value can be exactly one of several possible options. That tool then interacts with the rest of the language based on that invariant, sometimes providing things like exhaustiveness checking and pattern matching. Put all this together, and you have a very succinct and very clear way of representing certain kinds of values in a program.
Your comment might as well go through ripgrep and talk about how functions aren't needed. "They could have just used goto here and there."
> I looked at all instances and didn’t see anything game-changing. It’s a nice syntax and it should have really good performance, but such idioms are way too low-level to change any game.
Sum types were game changing to me when I learned about them over a decade ago. Since then, they have been a significant factor in how I think about and structure data in programs.
I have zero interest in trying to convince someone like you that you should think it's game changing. That's not the point. Maybe you could do some perspective taking and realize that others might just think differently than you.
Sum types are not novel (Standard ML had them 40 years ago), but I think they are indeed a game changer.
Sum types reify control flow into an object from which said control flow can be retrieved. Compiler checked sum types remove the possibility of retrieving inconsistent control flow.
The transform is equivalent to callback to future.
It's one of the things that looks unimportant until you use it. After that, the absence is repeatedly experienced when working with C++. We don't use tagged unions much because the ergonomics are terrible.
The choice to provide exceptions everywhere as a error handling means C++ is obliged to admit that your std::variant may not have a value at all. Which blows up all of your type safety. In Rust I can say that this Pet is either a Dog or a Cat, and it cannot be neither, but in C++ std::variant of a Cat and a Dog might nevertheless be valueless_by_exception anyway. What can you do about that? I guess you could throw an exception...
In Rust if X is either A or B, and Y is either C or D, and Z is either E or F, then a structure with X, Y and Z has only eight possible states. In C++ this structure has 27 possible states because of exceptions.
Wrt type safety we had more pressing issues with C++ (use after move, implicit conversions, ...)
You also don't have to handle the exception, in which case you won't access the variant again anyway. Or it gets handled where the variant is teared down. It's very unlikely that it gets handled where the variant is constructed or assigned to, making it a non-issue.
Granted, it's more awkward when you consume 3rd party libraries, but you can still wrap them in noexcept interfaces and be fine with terminating when an exception is actually thrown or do something else. Not much different to a panic.
(And C++ std::optional<T> is somewhat similar to Rust's Option<T> type.)
https://github.com/facebook/folly/blob/main/folly/Expected.h...
namespace our_project {
class Status { ... };
template <class T>
using Expected = folly::Expected<T, Status>;
}
absl::StatusOr<T> looks a lot like our_project::Expected<T>. This pattern of providing your own error type and aliasing Result is also somewhat common in Rust, I believe. const number = try parseU64(str, 10); // returns the error
const number = parseU64(str, 10) catch 13; // default
// more complex stuff
if (parseU64(str, 10)) |number| {
doSomethingWithNumber(number);
} else |err| switch (err) {
error.Overflow => {
// handle overflow...
},
// we promise that InvalidChar won't happen (or crash in debug mode if it does)
error.InvalidChar => unreachable,
}https://en.cppreference.com/w/cpp/utility/expected
Unfortunately it does not come with nice syntax to return on error.
> Unfortunately it does not come with nice syntax to return on error.
Yeah. We use a C preprocessor macro for this, but it's a little more verbose than question-mark operator.
It's less awkward with Rust's enums for sure though. And pattern matching as in Rust is far more expressive (and legible) than what std::variant gives you.
https://gist.github.com/erinok/c823af95db408653c7e42ab189307...
That was removed in C++11. The rule now is:
> Absent default member initializers ([class.mem]), if any non-static data member of a union has a non-trivial default constructor ([class.default.ctor]), copy constructor, move constructor ([class.copy.ctor]), copy assignment operator, move assignment operator ([class.copy.assign]), or destructor ([class.dtor]), the corresponding member function of the union must be user-provided or it will be implicitly deleted ([dcl.fct.def.delete]) for the union.
i.e., you need to provide an explicit version of the special function for the union if any member has a nontrivial implementation.
So you could have a struct or class with one std::variant field and some methods which can match on the type of the variant. But it would be kind of clunky.
In C, you have to wrap the union in a struct in order to add the tag, and that pattern is fantastically common. Let's have a little geometry-inspired example:
typedef enum { SHAPETYPE_RECTANGLE, ... } ShapeType;
typedef struct { ... } Rectangle;
typedef struct { ... } Circle;
typedef struct { ... } Triangle;
typedef struct { ... } Polygon;
typedef struct { // Outer struct, not a union at this level.
ShapeType type;
union {
Rectangle rectangle;
Circle circle;
Triangle triangle;
Polygon polygon;
} // This can be nameless in new(ish) C, which is nice.
} Shape;
Then you'd create a value like this, maybe: Shape rect = { .type = SHAPETYPE_RECTANGLE, .rectangle = { 0, 0, 20, 10 } };
Of course wrapping the initialization in a function would make it nicer.The above union-in-a-struct wrapping is my mental model for how enums work in Rust, but I still find it jarring. :)
So for your example you put ShapeType type in each of Rectangle, Circle, Triangle etc. and then you can union all of them, and the language promises that shape.circle.type == Rectangle is a reasonable thing to ask, so you can use that to make a discriminated union.
That doesn't sound right to me. Do you have a source? Is that in the standard?
11.5.1 [class.union.general]
[Note 1: One special guarantee is made in order to simplify the use of unions: If a standard-layout union contains several standard-layout structs that share a common initial sequence ([class.mem]), and if a non-static data member of an object of this standard-layout union type is active and is one of the standard-layout structs, it is permitted to inspect the common initial sequence of any of the standard-layout struct members; see [class.mem]. — end note]
So yeah, it's possible to store the tag in the common initial sequence of all the union members.
This is quite niche and rarely used. Most of the time it makes more sense to have the tag outside. Sometimes the tag in the common initial sequence allows the whole data structure to pack better than with a tag outside of the union.
C++ unions can actually have methods, although this isn't used very much. However C++ enums can't have methods, even C++ 11 scoped enums ("enum classes") can't have methods, I have no idea why that restriction seemed like a good idea.
† Unions are special because they're crazy dangerous, which is why they're not usually covered in material for learning Rust - you can't fetch from them safely. You can store things in unions safely because the process of storing a value in a union tells the compiler which is the valid representation - the one you're storing to, but fetching is unsafe because you might fetch an inactive representation and that's UB. However Rust does have a particularly obvious union right in the standard library - MaybeUninit - and sure enough MaybeUninit implements Copy and has a bunch of methods.
It is possible that Oracle holding the patent[1] to methods on enums is the blocker, rather than any technical restriction.
I don't think programmers often consult 20 year old "inventions", so it seems pretty obvious on its face that the supposed benefit of patents, that something is _only_ locked up for 20 years, is quite pointless in software.
Anyway, for loops are safe, unless the for loop is over the elements of a linked list. Then you need to wait until next year: https://patents.google.com/patent/US7028023B2
Patents really have a chilling effect, even if a particular one might not be enforceable.
Having said that this is the first time I ever heard of methods on enums being patented, what a ridiculous patent. It's a good thing then that C++ doesn't have methods, it has "member functions" :).
Also C++ allows user defined operators on enums, which feels somewhat adjacent.
In the case of MPEG the result is MPEG LA, a US company which you need to pay to implement certain important standards. In the case of JPEG the result was a little different, since only the improved Arithmetic Coding of JPEG was patented, people just don't implement the actual standard, they cut out the patented part, so all the world's JPEGs (well, mostly JFIF files, which are slightly different but we call them "JPEGs" anyway) are a little bigger than they need to be for no reason except patents.
So no, I don't buy that "ISO is also very averse of patents" in a sense that would restrict this unless you can show that's a new stance.
enum X {
A,
B,
}
fn foo() -> X::A {
X::A
}
which helps a lot when composing state machines. In the meantime, you can indeed do what you propose: enum X {
A(Foo),
B(Bar),
}
struct Foo;
struct Bar;
fn foo() -> Foo {
Foo
}Seriously though, calling Algebraic Data Types an Enum was a stroke of genius.
[0]: https://dr-knz.net/rust-for-functional-programmers.html
Dynamically-typed languages often have adhoc sum types.
Suppose I say that cars without seatbelts are "unfit" to use on public roads. Clearly I can't mean it's impossible to use such cars, people used to do it all the time, but perhaps I mean it's a bad idea to use them, and that's harder to argue with which is why we got laws saying you need seatbelts.
It's one of those patterns that just comes up again and again and it's wonderful when the language offers a good solution.
Since 2017, there is std::variant, which is a sum type template that _technically_ allows for language-level exhaustiveness checking via std::visit. It's not pretty, but it gets you there without requiring compiler support.
Rust's version looks way better.
Most modern C++ codebases do compile-time polymorphism via templates, though. std::variant is a niche use for things like serialization, when you need to make type choices at runtime.
It's just a pain to use because of how much template magic is used.
A "Rust enum" is a sum type.
Also, I talked about ASTs being a good example use case for a sum type. Here's a "real" version of an AST for a regular expression in all its complex glory: https://github.com/rust-lang/regex/blob/a9b2e02352db92ce1f6e...
You can see that it starts out as a sum type at the top level. And inside each variant is all sorts of product types and other sum types.
But do note in practice that sum types aren't limited to specialized things like ASTs. In practice, they come up everywhere. For example, here's a small little state machine used to implement a simple "unescape" routine. e.g., converting the string 'a\xFF\t' to the byte sequence 0x61 0xFF 0x09: https://github.com/BurntSushi/ripgrep/blob/44fb9fce2c1ee1a86...
> i have almost never needed to use such types
Well sure. When is an abstraction "needed"? Before Fortran, nobody "needed" a programming language either. So is a programming language necessary?
That's the funny thing about the word "need." It has very narrow application, and it's precisely why I didn't mention the word a single time in my top level comment.
If you're asking to compare tagged unions and unions, then...
A union is not a tagged union. A union is part of a tagged union. A union on its own is just a region of memory. What's in that memory? I dunno. Who does? Maybe something about your program knows what it is. Maybe not. But what if you need to know what it is and some other aspect of your program doesn't tell you what it is? Well, you instead put a little bit of memory next to your union. Perhaps it an integer. 0 means the memory is a 32-bit signed integer. 1 means it's a NUL terminated string. 2 means it's a 'struct dirent'. The point is, that integer is a tag. The combination of a union and a tag is a tagged union.
An abstract syntax tree is a classic example of a tagged union.
If instead you're asking to compare tagged unions and sum types, then...
Tagged unions are one particularly popular implementation choice of a sum type. Usually sum types have additional stuff layered on top of them that a simple tagged union does not have. For example, pattern matching and exhaustiveness checking.
In that case, the pattern matching and exhaustiveness means that the value the sumtypes in rust offer is in combination with those other language features, not the mere usage of the sumtypes themselves then, I think that makes more sense.
The discussion around "sum type" the concept is to disentangle it from the implementation strategy. Tagged unions are the implementation strategy, and they don't come with pattern matching or exhaustiveness checks.
Of course, many of these terms are used interchangeably in common vernacular. But if you look at the comment that kicked off this annoying sub-thread:
> so, um, unions? i have to say that in many years of programming, i have almost never needed to use such types.
Then this is clearly trying to draw an equivalence between two different terms, but where no such equivalence exists. There's no acknowledgment here of the difference between what a "union" is (nevermind a tagged union) and the "sum types" being talked about in my original top level comment. And the last part, "i have almost never needed to use such types" is puzzling on multiple levels.
no - one of the advantages of OO programing is that class hierarchies (should you feel the need to use them, which mostly i do not) can be expanded. or indeed contracted.
> Or variables where certain values have special case meaning?
very, very rarely (i would say never, in my own code) but of course we have the null pointer as a counter-example. to which i can say: don't use raw pointers.
It doesn’t make OOP obsolete, though, at all. Certain other areas are better expressed as hierarchies, e.g. GUI nodes.
https://fsharpforfunandprofit.com/series/designing-with-type...
By the third entry in the linked series, we've motivated the use of a sum type with three variants in the example domain model.
Enums + switch is the other, and far less powerful.
Switch expressions were also implemented, and they can exhaustively match on the given sum type. Pattern matching is quite limited as of now (only usable with records), but it is coming.
The problem with Java was always that it has moved so slowly in acquiring nice things. The actual core runtime is amazingly well engineered.
Relatedly, storing booleans is a smell, imho typically an enum or sum type is always better in languages that have concise syntax for these. True and False are meaningless without context, and so can easily lead to errors where they are provided to a different context or generally misused due to a misunderstanding about the meaning.
i agree completely with your second point - passing around or storing booleans is usually horrible.
No fanciness needed, just plain old sum types. It is certainly possible to express those invariants directly in languages with a dependent type systems or refinement types like in liquid haskell - see https://ucsd-progsys.github.io/liquidhaskell-tutorial/Tutori.... It's typically much easier to reason about and use sum types, though.
Of course these examples are trivial and silly, but I see instances of these patterns all the time in big co software, and of course usually the invariants are far more complex but many could be expressed via sum types. I've seen loads of bugs from constructing data that invalidates assumptions made elsewhere that could have been prevented by sum types, as well as lots of confusion among engineers about which states some data can have.
> Field X is only set if field Y is true
Original gnarly C style pattern:
struct TurboEncabulatorConfig {
// When true, the turbo-encabulator must reticulate splines, and 'splines'
// must be non-null. When false, 'splines' must be null.
bool reticulate_splines;
struct Splines *splines;
};
Rust (let's ignore pointer vs value distinction): enum TurboEncabulatorConfig {
NonReticulatingConfig,
ReticulatingConfig { splines: Splines },
}
> If field X is non-null then field Y must be null and vice-versa.Original gnarly C style pattern:
struct TurboEncabulatorConfig {
// When non-null, lunar_waneshaft must be null.
struct Fan *pentametric_fan;
// When non-null, pentametric_fan must be null.
struct Shaft *lunar_waneshaft;
};
Rust: enum TurboEncabulatorConfig {
PentametricTurboEncabulator { pentametric_fan: Fan },
LunarTurboEncabulator { lunar_waneshaft: Shaft },
}If we cannot run it at runtime now we need tests to hit that path or to test it manually to actually know if it works.
I'm just nitpicking but Rust is 7 years old and 7 doesn't feel almost 10...
Edit: Rust hit 1.0 7 years ago. Now I feel silly.
First spike at Google Trends for "Rust (Programming Language)" happened back in late 2013/early 2014 (https://trends.google.com/trends/explore?date=all&q=%2Fm%2F0...) so not at all impossible for someone to have ~10 years experience with it. Although that's probably pretty uncommon.
As a bonus, here is a very old (by internet time standards) HN submission titled "Mozilla releases version 0.1 of the Rust programming language (mail.mozilla.org)" - 236 points | Jan 23, 2012 | 82 comments - https://news.ycombinator.com/item?id=3501980
This is my first public commit: https://github.com/BurntSushi/quickcheck/commit/c9eb2884d6a6...
I didn't write any substantive Rust before that point. So I'm at over 9 years.
Edit: Oh uh you're burntsushi. Should've checked your username.
Rust has been around for well over a decade - the first “stable” release was 7 years ago. It’s always been opensource - any random hacker could download it and start using it since it was available. I think you will find kind quite a few people on this site that took it for a spin when it was still under development.
I have nowhere near the credits of BurntSushi, and even I was dabbling with rust more than 7 years ago.
Don’t be a reply guy.