Notes on the Go translation of Reposurgeon (2020)
gitlab.com
gitlab.com
I would also add that Enums + Exhaustive Switch is a very very weak area in Go that would really benefit the language a ton. I've used those features in other languages and that's one of the things I miss the most, especially when dealing with a ton of web API's that have a defined set of values for properties.
Able to have the confidence that we're checking for every situation that could occur on an enum type across the codebase is one less thing I need to worry about.
golangci-lint (https://github.com/golangci/golangci-lint) is an absolute must, and includes https://github.com/nishanths/exhaustive which will check this for you.
I'd rather just use go vet + staticcheck. Way simpler, no complex configuration, and no license concerns :)
and then you can create a file like the following and add whatever analyzer provides value for your specific case (including this exhaustive), easily. Example: https://github.com/FiloSottile/mkcert/blob/master/analysis.g...
e.g. #![allow(dinosaur::nonsense)] in Rust tells the tools that you know you're not supposed to do whatever dinosaur::nonsense might be, but you want to do it anyway in the following code and the hypothetical dinosaur linter shouldn't bother you about that.
When a maintenance programmer is staring at this block of code later, the fact you explicitly intended to do dinosaur::nonsense is right there, documented where it happens, and they can decide if the proper course of action is to leave that as it is, fix the code to not be dinosaur::nonsense, allow raven::stupidity because the new Raven linter is better but now warns about the same problem under a different name or what.
It's a halfway house to a language like Swift or Rust where none of the assignments have value. So, in a context where you don't want a value, arguably i++ is a better choice by eliminating a footgun. I don't hate it.
The bits about exceptions and typing remind me of an open question I have about Go: to me it mainly looks like a language for well-understood problems. Static typing and the lack of ability to do broad, high-level exception catching seem to me to be well-matched for problems where you already understand the solution pretty well. E.g., a port like this.
But how is it for poorly understood problems? E.g., You're doing a startup where the technical risks are low and the major unknowns are about user needs and the correct experiences to deliver. Or you're doing exploratory work for an art piece.
In those contexts, I've felt much more effective starting with vague/implicit typing and some broad, high-level exception-catching blocks. That lets me avoid a bunch of questions about the right types and structures until a more solid domain vocabulary emerges. At that point you can refactor/rewrite toward clarity. But I don't see a good way to be willfully vague/casual in Go.
If you don’t care so much about costs and scale on the server side then something like Ruby on Rails might give you that extra productivity boost you need to verify the concept.
And yes, the point of verifying the concept is key for me in this. Before product-market fit, I just don't have much confidence in any domain model. Once we have demonstrated that particular people are excited to pay for a particular thing, that changes. Then we can understand how those people think, and how we think about their behavior, such that we can do real domain-driven design. At that point I'm much more willing to lock down types.
For me, having strict typing helps with unknown problems because it gives me greater confidence that refactoring wouldn’t introduce subtle regression bugs.
For me the cost isn't just redefining a type. It's the continuous work of reifying types that could be implicit for a while and possibly forever, because concepts don't last long enough to need it.
For me, having written a lot of code in Java and Scala but also in Ruby and Python, typing is a nice adjunct to unit tests and operational monitoring, but not my fundamental way of ensuring correctness. And correctness often isn't the highest value in practice, especially early on in a startup's life.
The exception handling looks onerous, and is a pain at the start. But it doesn't take long to get used to it, and then (for me anyway) it becomes second-nature. Everything returns an error, and you have to handle that error (even if only passing it up again).
Static typing is more "fun" when you're exploring new concepts, but interfaces are the key. Defining the interface at the beginning ("what do I need this thing to do?") and then implementing it with a struct/whatever, feels like the right way to do this.
I did get used to it, but all the
if err != nil {
return nil, err
}
sometimes taking up half or more of the vertical space of a function is still an eyesore to me, even after writing a quite substantial amount of Go.Maybe the Go ? is allowed within any function whose last return value is an error on any function call whose last return value is also an error. Call it with ? at the end and accept all but the last return value. At runtime, if the err is not nil, then it returns from the function, supplying the err value and nil for any other return values.
I don’t think it went anywhere, unless I missed it.
First,
..., err := process(...)
if err != nil {
return nil, err
}
is an unfortunately common antipattern. Errors should always be annotated, e.g. ..., err := process(...)
if err != nil {
return nil, fmt.Errorf("process: %w", err)
}
The extra lines carry no significant cost -- it's not like reading them imposes a burden versus parsing a single line dense with semantic information. They expose the `return` keyword, which clearly signals a control flow point that is hidden by method chaining and `?`. And it doesn't grant the error control flow special status! These are virtues. I don't see this as unuseful at all. It's fine if you do, of course! But there's not an objective ruling, here.Passing an error manually up several levels when it may only be a theoretical concern is a ton of expressive duplication. If I'm in the mindset, that would feel like valuable work, even though it isn't actually making things better for users.
It reminds me of one team I dealt with years ago. They were Java experts used to doing enterprise stuff; their overlords had put them on a scrappy, startup-like thing that was intended to be open source. As was in fashion for Java at the time, they had written things in many layers, the goal of which was to provide scaling cut-points. However valuable in theory, the project ever needing to scale was uncertain. What really mattered was finding something that served user needs. But the layer-by-layer duplication increased the cost of change significantly, lowering the odds we'd find the right product.
I don't miss _exceptions_ in Go but after using Rust for a while I've come to really love Rust's `Option` and `Result` types. They're more ergonomic and expressive than Go's errors-as-values, and the use `Option` also eliminate `nil` entirely, which has been a regular source of runtime panics in every large Go codebase I've worked on.
In Rust itself, writing stuff where I actually care whether my structure is 16 or 20 bytes because I need to fit hundreds of millions of them into RAM, I like that Sum types ensure Option<NonZeroUsize> is the same size as usize, by reasoning that 0 isn't a valid NonZeroUsize and so it can signal None. However on a language like Go I don't miss that - what I do miss is the fact that in Rust I can't mistakenly end up with None(actually_something) or both an error and the result that shouldn't be there if there was an error.
Good idea :) `?` hides a control flow that Go takes great pains to make explicit. Adding it to the language would be a disaster.
In conditions of high volatility, that's not always the case. If I create some throwaway prototype to test a hypothesis based on talking to a few users and then test it on a dozen more users, I don't want to think a ton about the domain. The sheaf of possibilities is often quite large, and narrowing down the possibilities requires better understanding of the users and the things we're creating for them. That understanding is only available in the future, and we only get to that future by making something.
I really do love designing clear, expressive type systems. But the beginning of the project is when we know the very least about what will happen, so it's the worst time to invest in those type systems. It can be fine anyhow if the domain is stable and well-understood. Which I gather is Go's sweet spot, and I'm fine with that.
I think that's what Go was designed for, as a language. To be readable, usable. It's uncaring for your personal programming philosophies. I think that's why it's been successful.
Go's philosophy (which clearly flows from its creators being C-enthusiasts) is that the only thing that matters for reading, writing, and understanding a program is what it concretely does, i.e. what structures are created, where values are stored, how computations are performed etc. If that's also your philosophy, then of course it's going to jive with you.
But plenty of people also have different philosophies. Maybe you think the main thing that's important in crafting programs is developing a rich domain vocabulary that expresses concepts and how they interact. Maybe you think that what's important is formal proof of both logical and concrete correctness. In those cases, Go's rigorous opposition to abstraction (coming from its philosophy that what's important is concrete operations) will probably irritate and slow you down.
I couldn't say exactly why it got popular. I'd guess that some significant segment of programmers also share its philosophy, but I have no evidence to back that up. Certainly any reasonably uncontroversial language with a large suite of libraries backed by Google is bound to have some level of popularity.
Probably because it has the backing of Google.
You can do the same work in Java but you can't statically link the JVM. You can sort of do these in Python, but the compiler story is murky at best, and the language isn't as type safe.
That's not quite accurate; there is a runtime, it just gets statically linked into the binary instead of needing to be externally installed
In particular, I don't understand how Go is more Java-like than Dart.
Feature | Java | Dart | Go
-------------------+------+------+----
jit compilation | yes | yes | no
inheritance | yes | yes | no
classes | yes | yes | no
nominal subtyping | yes | yes | no
native binaries | no | no | yes
static artifact[0] | no | no | yes
static typing | yes | opt | no
value types | no | no | yes
What other features do Java and Go have in common that they don't also share with Dart?[0]: For sanity's sake, we'll assume this means "are static artifacts common/default" and not "is it technically possible to produce a static artifact" because for some sufficiently broad definition of static artifact the answer can be yes for any language (e.g., Docker images).
Go doesn't particularly value readability or usability. For example, the short variable name convention makes it harder to read code you're unfamiliar with, and there are a number of noticeable usability shortcomings, some of which are mentioned in the article.
I think Go's design goals were really to (1) reduce compilation time, which explains why Go has human programmers do work that compilers do in other languages, (2) be statically typed and compiled, so you can use it conveniently for microservices, and (3) have syntax somewhat similar to Python.
I think of Go as the successor to Java. Go is to Python as Java was to C++. That, plus the integration with many libraries is why it's taken off in some niches.
There is a forest full of code hiding a trees worth of business logic, always.
Go is one of the least expressive languages I've ever used.
That is by design, and when it comes to "programming in the large" - a winning formula. I keep repeating this response: I worked on a Perl codebase with a medium-sized team. Perl is very expressive, and my teammates did not hold back. I can tell you that is a nightmare to debug or add a new edgecase to a "clever" Perl 1-liner, usually it involved making the code "less expressive". So, I'll take Go over the more expressive languages in a team setting any day.
My needs are pretty basic: I like code that is easy to understand and easy to change more than writing code that leaves a smug smile on my face. I read more code than I write, so YMMV.
The fewer surprises, the better for me, and so far, the collaborative codebases I've encountered the least number of surprises have consistently been in Go (the other languages I've been paid to work with are Javascript, Perl, Python, Java, and Scala).
But we still have to write code that these people understand. It makes no damn sense.
Sometimes, sure, people write horribly complex code, but sometimes it's just developers who have stopped learning. They see something that isn't immediately familiar and discard it as too complex and make no attempt at trying to learn.
What I don't get is why we have to pander to these people.
Huh? What do you consider syntactic bureaucracy? Go has like 25 keywords and no sigils -- the least "syntactically bureaucratic" language I'm aware of!
> Go is one of the least expressive languages I've ever used.
This is definitely true. Of course expressiveness is not strictly a virtue!
I would assume "amount of syntax required to express a given concept," with the use of the word "bureaucracy" implying that some concepts require too much syntax relative to their complexity (something that depends on your values). The classic example being mapping over a slice.
Here's some Rust code
pub fn read(&self) -> u64 {
self.counts.values().fold(0, |acc, x| acc + x)
}
Here's some analogous Go code func (w *Whatever) Read() uint64 {
var total uint64
for _, v := range w.values {
total += v
}
return total
}
The former is certainly fewer characters than the latter. But to me it represents _more_ syntactic bureaucracy, not less. There are more sigils, more language concepts I need to understand, more _types of syntax_ to express the same thing. It's 20% of the SLoC, but parsing it requires more implicit knowledge, and takes no less time, versus parsing the latter.YMMV, of course. None of this is objective.
Similarly, even though I'm not a big fan of Rust, I would personally prefer to encounter the Rust snippet. The way I read code, I'm already building up mental models of things in my head, so adding more (e.g. what fold is) isn't that big of a deal to me. I think I'm a person who is able to look at a function call and not have the desire to dig into its source, though, which I don't think is the way everybody (and most certainly not the designers of Go) feels. Plus once I understand the concept, even if it's a lot less universally applicable than fold, I can reuse my understanding of it throughout the system and possibly throughout multiple systems.
Like you said, I think this is largely a subjective thing. I just find it objectionable when people, on either side of the fence, come in and say "abstracting over concepts and possibly making them first class is always better" or "...always worse."
Now, there is a problem that all of the struct fields have to have sane zero values for this to work, but I think that is a problem which would naturally arise with keyword arguments as well.
Would it? Only if you allow default values, which is a separate discussion. And even then only if your only mechanism for default values doesn't let the definitions specify what, exactly, the default value is.
Keyword args practically require default values, I think? In any case, while this feature (like every feature) definitely delivers value, it's the considered position of the Go authors that this feature, over time, has a net negative impact on program maintainability.
Isn't a collection of parameters to a function an intent?
Why? What's the utility? Fewer characters?
If you take
func NewServer(addr string, certKey, certCA []byte, ...) (*Server, error)
and coalesce the input parameters into type ServerConfig struct {
Addr string
CertKey []byte
CertCA []byte
...
then that type is still a meaningful domain concept in your program. There is no minimum scope requirement for discrete, named entities, is there?Right, but a struct means "a collection of named parameters", not "a concept in your domain model" although a struct can be used to model the latter.
> A discrete, named entity (a struct)
Structs don't have to be named. E.g., `var person struct { Name string; Age int }`.
> It's the same thing when "Go has no set type" comes up and people say "map[T]struct{}!" You might implement a set using a map, but they're fundamentally different concepts
It's not the same thing at all. In your analogy, a map is not a set but the advice says to use the map instead of the set anyway. But we're not talking about using one thing as another, we're talking about using a struct as a struct--one such (common) use for a struct is passing named values into a function.
That's true. But IMO a language primitive can carry a wide variety of semantics.
[0]: https://www.johndcook.com/blog/2011/07/19/you-wanted-banana/
Using structs to bundle function parameters doesn't preclude or inhibit using them for domain modeling.
I'm not sure which came first, but structs are a ~60 year old concept. That ship has sailed.
Keyword arguments and structs are both named collections of values. Structs just happen to have broader uses. You can argue that we should use the more specialized one, but that's pretty silly--do we implement a general "add(a, b)" function or do we use more specialized add1(b), add2(b), add3(b) functions? Of course we use the more general tool even though a more specialized tool could exist.
Currently the workflow is svn-all-fast-export (SVN to Git) -> git filter-repo (trim content from the repository) -> Git LFS (store large files out-of-band), where each stage has to be manually tested and written, and then hopefully put into a shell script or Bash history or something.
Unfortunately, I have yet to get Reposurgeon working for us at all. The documentation was inaccurate, in that it referred to things which only existed in the Go version, and the Go version went OOM (on a server with 384 GB of RAM) on every repository I've tried it on. This includes just reading from an existing, pruned svndump file, so it's not resource contention either.
Basically, the tool is great conceptually, but it absolutely does not work for our use case. We had to stick with our existing, messy, multi-component solution, which has worked surprisingly well over the last two years, and since I'm the one doing most of the processing, having kind of a messy system is relatively acceptable for the time being.
How many commits does your repository have? The largest one we’ve successfully converted from SVN to Git ourselves was 287k commits.
1. I'll try doing some profiling and see what happens
2. The latest repository I tried to import has 671,190 revisions
You can also force it to write all file content (blobs) out to temporary files rather than keeping them in memory by reading the stream from standard input:
reposurgeon "read -" … <dump.svn
That will be slower and use more disk space, but maybe it will fit.And yes, some of our repositories are huge. This one is about 240 GB, and comprises multiple projects/products.
Our "most important" repository is only 313,000 revisions, but those revisions take up 93 GB, so it's a not-insubstantial repository as well.
There's a few other differences, strings cannot be directly mutated or have the address of their elements taken.
I like Kotlin, but I really think of Go as the most probable outcome of someone saying "I'd like a staticly typed, concurrent/parallel Python", and then mumbled "but I hate exceptions". Go is very close to python in many respects.
I guess maybe they were worried about the lack of GC, but my experience has been that's it's generally quite easy to port code from dynamic languages like Python/JS/PHP to Rust. So I suspect that may have gone into their "Expected problems that weren’t" section if they'd tried it.
This is exactly the impression I have gotten from the crates system every time I've looked. It happens for things as foundational as mmap. Also terrifying: The plethora of highly-recommended crates that have not yet committed to a stable API (ie, are still semantically on version 0.X).
My view is that when you develop a product, you have to own everything. The users don’t care if the bug comes from code you wrote, or code in a library, or the language’s standard library, or the OS. You have to fix the bug no matter what caused it. It doesn’t matter if the code came from a vendor or a language designer or stack overflow; you are the one responsible for fixing it if something goes wrong. Non est salvatori salvator, etc.
With that perspective, I don’t think that 23 crates is terrifying. I’m going to look them all over and either pick one, or write the 24th crate myself. The result is the same either way.
I do agree with you about APIs though. It is nice to find a crate where the author has had the confidence to stabilize their API and declare the version to be 1.x instead of 0.x. But at the same time I recognize that getting to that point requires some real software to use the crate, to create the feedback loop. If nobody ever used a 0.x crate, no crate would ever get that feedback.
So then using a programing language that has a big standard library and a more rich and stable ecosystem is a big advantage. If I want to do a network request, parse a json file, parse or format a date or some other trivial thing I prefer to have a good enough built in way to do it then evaluate 15 packages or write it myself. Otherwise you get in a shitty situation where you inherit some projects and it has 20 dependencies with vulnerabilities, 20 dependencies abandoned, some dependencies that are incompatible with some new version of the language/ecosystem causing issue if you would like to upgrade.
Also, don’t forget that languages with big standard libraries, like Python or C# or Java, always end up with oodles of stuff in there that is deprecated or incompatible. That’s as bad a sign as anything.
But I do agree with you, a lot of the time you have inherited the mess rather than writing it yourself. Not much you can do about it, except to clean it up and then dodge better next time. Going back to Rust specifically, Cargo gives you some tools that make the mess easier to clean. The Rust language also helps, because you can safely use multiple versions of your dependencies, when it turns out to be necessary.
This does not happen that much, and in Java and .Net world if some functionality is deprecated there is a replacement, I think I only remember some possible unsafe or ineeficient functions were deprecated so you use the better ones. I don't have experience with Rust , only with node/npm and is a hell , for some reason is decided to split stuff in super small pakcagtes of various quality, today I sepnd hours debugging a npm freeze caused by some package , in the end the cause is probably a shit npm/node implementation that crashes when some specific git version is installed , but debugging this I discovered that the unit testing packages the project uses(I inherited them) instead of beeing one or few packages are a few, and for some reason one of them had a weird dependency that it should not have(a dev only despondency ) and this dependency was also just pulling directly from GitHub ...shit in a few years when GitHub is gone lots of things will stop working (I also had issues with scripts because someone changed master into main because American politics...).
I personally prefer the Java or .Net model, a big standard lbirary and then for most important things you have a small number of options since the community did not wanted to do CV driven development, and most of the time you find all you need in a single library with no or few dependencies. But Rust does not have the Java or .Net money to hire devs to work on the boring stuff of correctly implementing standards like json,XML, date&time and keep maintaining it - (not sure how Python did it) and Rust community seems to be inspired byt node community a lot and this is bad.
In a fine-grained ecosystem, the problem is transitive. You aren't just picking one from 20+ for each of your direct dependencies, you are also trusting that they in turn did just as much diligence as you did in picking their direct dependencies and so on. Curation and pruning would at least help mitigate the scope of the problem.
> ... or write the 24th crate myself.
This is a good example of the xkcd joke about standards proliferation. If I found myself in this situation, I'd prefer to keep the 24th private to my own package. DRY has its limits. Sometimes you've just got to specialize.
I'm aware of the irony in saying this in the context of mmap. But in this case, a thin wrapper that reflects the system call's semantics ought to be available in a common `posix` crate that others may freely rely upon to build their higher-level services. There's no good reason for e.g. ripgrep to have to pick and choose among a dozen flowers for it.
Absolutely. You definitely have to own the whole stack. Ultimately, you may have to fix bugs at every level, from your own code all the way down to the OS. Big tech companies do this explicitly, with kernel development teams. Small companies do this by occasionally upgrading to the latest version of Ubuntu and hoping for the best.
> But in this case, a thin wrapper that reflects the system call's semantics ought to be available in a common `posix` crate that others may freely rely upon to build their higher-level services.
It is available. If you want to directly call any function in the C standard library, use the libc crate. Here’s the documentation: https://docs.rs/libc/0.2.109/libc/fn.mmap.html
All the other crates on crates.io that you find when you search for “mmap” are higher level abstractions over the raw syscall. They have specific purposes, like file io, or creating circular buffers.
Incidentally, I helped with the port to Go, and I spent some time polishing the code, and finding and fixing performance problems. When we were helping GCC convert their SVN repository to Git (a repository with 287k commits, btw), we reduced the memory usage by 50% (from over 250GB to under 128GB), and the run time by quite a lot as well (down to just around 2 hours to read in the SVN repository and convert it to a basic Git repository).
Now that we’ve done that work, Reposurgeon spends 50–60% of its cpu time scanning the heap for garbage. There is often garbage to find, but just as often there is not. GC is useful, but for Reposurgeon it has become a bottleneck.
My preliminary work on a Rust port shows that it is around 4× faster than the Go version. I personally think that Rust is the future, but I haven’t been able to put as much effort into the port as I would like.
> He did investigate Rust to a certain extent, but bounced hard. His chosen program didn’t really show off Rust’s strengths, because he started by trying to call select and write essentially the same program that he would have written in C. That’s a pretty painful way to go.
Is this stating that he was writing the Rust version similarly to how he would have written it in C because that was the most natural way to do it in Rust or simply because it was the first path he went down? In other words, was your point that Rust is a fundamentally poor fit for this problem, or was it that his lack of experience with the language took him down the wrong path?
Thanks!
What I really missed was generic map-function-over-slice, which could be
handled by adding a much narrower feature.
If one graded possible Go point extensions by a figure of merit in which the
numerator is "how much Python expressiveness this keeps" and the
denominator is "how simple and self-contained the Go feature would be",
I think this one would be top of list.
So: map as a functional builtin takes two arguments, one x = []T and a
second f = func(T)T. The expression map(x, f) yields a new slice in
which for each element of x, f(x) is appended.
I think this is an excellent way of evaluating new features, and it represents a real missed opportunity for Go to explore new PL territory. Instead, we're getting full-blown user-defined generics, which increases the "denominator" far, far more than it increases the "numerator."btw, by "new PL territory" I mean "reifying a small set of container operations, without supporting user-defined generics," which to my knowledge is not a position taken by any mainstream language.
A mapping for loop doesn't pollute scope with incidental variables:
results := make([]Result, len(input))
for i := range input {
results[i] = callback(input[i])
}
^ This only adds `results` to scope, which is the same as `results := map(input, callback)`. In the for loop example, the loop variable `i` is scoped to the loop.Moreover, if you don't care about terseness, you can always pull this out into a well-named function or annotate it with a comment.
> Everyone knows what map/filter/reduce do
In isolation, but for complicated chains of map/filter/reduce (especially with error handling logic in languages which return errors rather than raising them as exceptions) it's much easier for me to read the corresponding for loop equivalent. Even my colleagues at a Python shop had limits on the complexity of list comprehensions beyond which point they were required to rewrite into a for loop because while packing that complexity into a single expression is elegant and clever, it's not particularly readable or easy to understand.
I guess my view can be summarized as: for very simple cases, map/filter/reduce are a bit clearer, but for those same simple cases a for loop is still easily understood and a for loop's readability scales better with complexity.
Just like with monomorphisation, Go just stubbornly rejects all PL research younger than 40.
That's my point.
> Just like with monomorphisation, Go just stubbornly rejects all PL research younger than 40.
That's flamebait.
But Go was not intended to explore new PL territory. Is that not OK?
Me too, and mapping a map, a channel, etc. With generics, things will be easier, we can have iterators, even lazy iterators, but there will still be no type inference in lambdas for arguments and return values and no easy currying, so things will still be awkward, but better.
> Keyword arguments
You can kind of do that by passing a struct as an argument, and represent the optionalness by having a pointer in the struct. That's still awkward, you can't easily make a pointer to an integer or string literal, you need a separate helper function for each primitive type. Overall it could be nicer.
> Annoying limitations on const
I think the Go constants are good, the fact that you can perform calculations on them in compile time also. But I wish there was some way to express immutability (and non-nilability) in the language. Programs are simpler with less moving parts.
> 14KLOC -> 21KLOC
That's to be expected, Go could be nicer if it was more concise, obviously that's a hard thing to balance, because you can end up with too obscure code, I just wish that there were some small improvements from time to time.
> Absence of sum/discriminated-union types
Yeah, I refer to it as Algebraic Data Types, where you can for example have your tree type to be a Node with children or a Leaf with value. And a function can accept a Tree that can be one of these. IIRC, the Go people claim that you can do kind of something like this with interfaces, but it's not very nice to use interfaces in this way and has some drawbacks.
> Catchable exceptions require silly contortions
I don't like exceptions, they disrupt the flow of the program - any function call could "return early", so you need to be always careful and account for that possibility. In C++ this was solved with RAAI, and it was a source of so many bugs that some companies just disallow exceptions internally.
> Aesthetic doubt
Yeah, it's there, some things in Go just come out ugly.
> Absence of iterators
Yes. Hopefully generics will enable implementing them.
I think using a builder is a better option in Go if you have more than 2-3 arguments.
> Yeah, I refer to it as Algebraic Data Types, where you can for example have your tree type to be a Node with children or a Leaf with value. And a function can accept a Tree that can be one of these. IIRC, the Go people claim that you can do kind of something like this with interfaces, but it's not very nice to use interfaces in this way and has some drawbacks.
Except you can't really do this with interfaces. If you need to get back to the concrete implementation of the interface, there's no way to know what all possible implementations might be. This is especially true since any type can implement an interface, not just the types you create initially. So someone else could implement `Tree` and you would not know to account for that in your function that accepts the `Tree` interface.
I don't think there's any substitute for proper sum types in Go.
But then you need to write all the setters, and something that was supposed to be simple, a function, becomes a whole contraption with a bunch of auxiliary code. Personally I would avoid that, and I definitely wouldn't make it a rule of thumb to use builders in every function with more than 2-3 arguments (maybe you meant something else).
Yeah, it's tedious, but that's Go for you ;)
Rust has the same issue, but there are macros libraries that will create the builder for you based on the struct definition, so it's _very_ trivial to create these builders.
You could do the same with codegen for Go, but this always feels much worse to me than using macros for a number of reasons like having to install a separate tool, checking in generated files, making people run `go generate` after some (but not all) changes, etc.
Apparently they decided that it would be too confusing to have variant types alongside interfaces for some reason, and they claim that interfaces handle a lot of the use cases of variants:
Rust has both (traits are somewhat like interfaces) and I don't feel like it's too painful, though using the trait type in function signatures is a lot more involved than in Go, and I still don't fully understand all the nuances. But the complexity isn't because of any overlap between traits and sum types.
type A struct {
B int
C int
}
myfunc(A{B: 1, C: 2})e.g.
func myFunc(cx context.Context) {
ctx, cancel = context.WithCancel(cx)
defer cancel()
for item := range iterator(ctx) {
}
}As a rule of thumb, the amount of work transferred by any concurrency primitive should significantly exceed the cost of the concurrency primitive itself. I do have a couple of uses of this pattern where what is on the other side of the channel is something reading off a network and parsing lines of JSON into internal structs, in which case the overhead of the channel isn't necessarily too bad. (In one case, it even chunks the lines of JSON into a slice of several parsed structs, reducing channel overhead even more.) But it's a terrible solution in general; iterating over an array and doing any sort of very fast "thing" to each element that only costs a handful of assembler instructions, a very common use case, has terrible overhead.
It's a real pity, because the semantics of that solution are pretty close to the right answer. But it's a huge performance trap. Something as basic as iteration needs to not be a huge performance trap.
[1]: Relative to Go, anyhow. I haven't timed it directly but I wouldn't be surprised that a channel-based iterator like that would be comparable to Python's general iteration speed, or at least not off by a very large factor. It's just that "Python's normal level of performance" is "atrocious Go performance".
If you need raw performance and you are just doing some minimal operations on a slice then yeah you’d want to use a simple for loop instead.
I typically use this pattern in situations where you have a ton of data coming back where you want to avoid storing all of that data in memory.
func intSliceIter(ints []int) func() (int, bool) {
i := 0
return func() (int, bool) {
if i < len(ints) {
ret := ints[i]
i++
return ret, true
}
return 0, false
}
}
iter := intSliceIter([]int{0, 1, 2, 3, 4})
for x, ok := iter(); ok; x, ok = iter() {
fmt.Println(x)
}
Of course, there's not much benefit to a SliceIter; this is a contrived example, but you can apply this pattern in more complicated cases as well. Similarly, you can define an iterator as an interface (which is similar to bufio.Scanner and a few others in the standard library—a closure is an object is a closure): type IntIter interface {
Next() (int, bool)
}
type IntSliceIter struct {
Cursor int
Ints []int
}
func (isi *IntSliceIter) Next() (int, bool) {
if isi.Cursor < len(isi.Ints) {
ret := isi.Ints[isi.Cursor]
isi.Cursor++
return ret, true
}
return 0, false
}It should be noted that the writeup is a year old, which (understandably) means that the comments on generics are a little out of date.
Just the other day, I was able to condense over 15 lines of golang code into 3 lines (could also have been 2 lines) in a Python-like syntax, both reducing code length, and substantially increasing readability as it would make the underlying logic clearly stand out instead of having several loop and map constructs.
There's some temptation to think that terse==readable, but in practice this rarely extends beyond the simplest cases.
The real issues (IMO) with Go error handling are the clunky `errors.Is()` and `errors.As()` functions for catching specific error values and types respectively as well as annotating errors to get the requisite context (this may already have a solution which just hasn't yet become idiomatic across the ecosystem). In the meanwhile, I just wrap errors with extra context which isn't super satisfying but largely does the trick, e.g., `return fmt.Errorf("writing to database: %w", err)`. Rust seems to have similar problems with respect to different patterns and libraries for error handling, despite having a standard Result type.
A quick glance at any array language (APL, J, K, Q, etc.) confirms this. ;)