Union types ('enum types') would be complicated in Go
utcc.utoronto.ca
utcc.utoronto.ca
Additionally, just having actual enumerations, Pascal / Algo style, not ML style, would already be an improvement over the iota/const hack.
var foo interface {
struct { A int } | struct { B string }
}
It currently fails with the following error:"cannot use type interface{struct{A int} | struct{B string}} outside a type constraint: interface contains type constraints"
(I'm the author of the linked to article.)
In cases where there is a lot of variance in the size of the different interface implementations, separate allocations could actually be more memory efficient than a tagged union. In any case, I'm not sure that memory efficiency is the main reason that people miss Rust-style enums in Go.
I agree with you about the overall motivation for Rust-style enums. I just think it's surprisingly complex to get even the memory efficiency advantages, never mind anything more ambitious.
I'm not sure what you're referring to with 'anything more ambitious'.
The solution to this should be trivial. You just have to extend the gcshape concept to account for the enum discriminator.
Honestly interfaces with unexported methods are 90%+ of what people want. It's just not spelled the way they expect. And if you're not going to be happy except at absolutely 100%, a position I can and do respect, there's no point waiting for Go to get any better because I can guarantee you no Go proposal for sum types will fix that you will be forced to have a "nil" value in the sum type, so there's no point in waiting.
If interfaces are used for union types then the obvious zero is nil, not any constituent. Nil is the zero value of interfaces.
It’s perfectly consistent and in line with the rest of the langage.
Even if you have to manually define separate named structs, you still have the benefit of exhaustivity checking in type switches. That's arguably the other 10% that people want.
var foo interface {
struct { A string } | struct { B string }
}
Eg in Rust that would be 'Result<String, String>', where your success happens to be a String and your errors happens to be a String error message.This doesn't cover everything people might want to do with unions, but it covers the billion-dollar mistake and doesn't run against the grain of the entire language (as far as I know)
Not an expert, but my gut says maybe it runs against zero values? As in, "what's the zero value for a non-nullable reference?" Maybe the answer is something like "you can only use this type for parameters", but that seems very limiting.
What is missing is the ability to have pointer variables and have the compiler ensure that it will be never nil. I believe this was a design choice, not some technical limitation.
That seems better than not having algebraic data types at all.
It's basically a non-issue IMO. Just stop explicitly ignoring the errors returned from constructors.
Some of the most common areas that infest the code with nullable pointer types are when you have to deal with de-serializing data a lot. This is due to lack of a common built-in Optional type (I know you can define one easily, but you can't force libraries that you rely on to use that type).
The best we have now is https://github.com/uber-go/nilaway and it's improving with time. It does static analysis, but it has a very difficult job to do, so right now it's super slow to run, and is prone to having false-positives.
Union types and enum types are not the same thing, and this misunderstanding invalidates the entire article. An enum type includes a marker which indicates which value it contains. The garbage collector would be able to read this tag value and know.
What Rust calls enums are called tagged unions in most other languages. Enum usually refers to a simple collection of named values, what Rust calls fieldless enums or unit-only enums.
Not sure about that. Swift also calls them enums, in Java and Kotlin (and maybe Scala? I forget) they're sealed classes/interfaces and in PL theory and many typed FP languages they're called "sum types".
Enumerations are a sequence of values.
In the sense that it first points to the type, and then to the value, similar to how other newer types are already implemented?
type whatever enum {
value1 iota
value2
}enum would just be a similar shorthand to uint or whatever you want to identify it uniquely.
I don't see much problem implementing this apart from the methods in the GC that have to be touched to trace the object tree correctly.
The comparison handling must be touched in either case, so I don't think an additional pointer to the enum's definition there counts as much work.
To put it another way, a Rust `enum` is not the same as C's untagged unions but it is still a union.
> let's ask why we can't implement such a union type today using Go's unsafe package to perform suitable manipulation of a suitable memory region.
They are aware that you could add union types to Go - their point is that it would require garbage collector modifications which may be difficult:
> The corollary to all of this is that adding union types to Go as a language feature wouldn't be merely a modest change in the compiler. It would also require a bunch of work in how such types interact with garbage collection, Go's memory allocation systems (which in the normal Go toolchain allocate things with pointers into separate memory arenas than things without them), and likely other places in the runtime.
Metadata. The same way shapes of functions let Go know which generic 'profile' to use for a given thing are metadata. Just like when a reference to something is taken that type is compile time metadata.
Structure. The precise layout of structures and other data fields is also compile time metadata. They don't even need to remain the same between versions of go, or even builds if somehow they're randomized. That isn't how programmer's think (at least any who were also trained in assembly???). When I lay out a struct I do expect undersized fields to get padded, but I expect every field in order, and I'd prefer some way of forcing the issue for precise padding.
However, 'union' of types is just syntax sugar. Give the programmer the above basics and add one more: builtin.*reshape()*. reshape() would allow any similarly shaped structures to replace the type of the reshaped item. E.G. reshape({x, y, z uint64},{x, y, z int64}, A, B) would convert a 192bit chunk of 3 ints of one type to the other. It could also convert anything else similarly.
That's a trivial example, what about some private structure from a library? I'd think the unsafe package's version should allow violation of the private field space, but the normal safe version might force the unexported fields to 'pad' (inaccessible) space. I don't think this would alter garbage collection, as that process likely has to keep it's own track of regions of memory and places that point within them. That's runtime (maybe compiler time sometimes?) metadata which reshape() would have to work with.
Generics even _sort_ of do this already when prefixed with ~type in the list of allowed types; the compiler's allowed to use the passed thing as that type of value and the return it back the same as input... and I really don't see why a reshape function couldn't do the same thing.
There is something else though: reshape() likely needs to consume / claim the resource, since it wouldn't initialize anything. So maybe it needs to return the recast value to be assigned or passed somewhere and further invalidate usage of the variable past that use. Alternately it could take a single value variable and modify the type as part of it's call (also providing the value as it's return would be useful sometimes too).