Go range iterators demystified
dolthub.com
dolthub.com
func (s Slice) All() func(yield func(i int) bool) {
return func(yield func(i int) bool) {
for i := range s {
if !yield(s[i]) {
return
}
}
}
}
”So, for an “easy enough” example correctly, you have to write func five times in order to, if I understand this correctly, wrote a function returning a function that takes a function as an argument?
For comparison, C# does that this way (https://learn.microsoft.com/en-us/dotnet/csharp/language-ref...)
IEnumerable<int> ProduceEvenNumbers(int upto)
{
for (int i = 0; i <= upto; i += 2)
{
yield return i;
}
}
Yes, that introduces “magic” where the runtime figures out that ProduceEvenNumbers won’t continue, but why give such functions the flexibility not to listen to such requests (in the golang version, is forgetting the if and just yielding instead ever useful?)Go is far from perfect and certainly has a lot of flaws, but IMO the iterators design fits well into the language and does not deserve the criticism.
The range only needs a function, so what is the point of calling a function that returns a function?
This is enough to make a range iterator, you don't have to call `All` manually:
func (slice Slice) All(yield func(string) bool) {
for _, s := range slice {
if !yield(s) {
break
}
}
}
https://go.dev/play/p/7AYOXzlHU4t?v=gotipIn Go, the expression `object.method` already returns the method as a function, with `object` being already bound in the scope of the function. The syntax is so that if you have this definition:
type T string
func (t T) f(a int) int {
return a
}
Then:- `T.f` returns `func(t T, a int) int`
- `T("whatever").f` returns `func(a int) int` where `t` is already set to `"whatever"`
- `T("whatever").f(42)` calls the function and returns `42`
I don't tend to use that convention though, because most of the time when I use this pattern, I'm returning a function that uses a parameter I pass in. You can see this in the predicate example. I prefer keeping my call sites consistent and always using functions that return a function when invoked, rather than sometimes invoking them and sometimes just passing the func reference. YMMV.
Is it ever useful to be able to do something there before returning? If so, wouldn’t it be cleaner to implement this as an interface with two methods, one that simply has to yield for every item produced and an optional one that gets called when the runtime knows it won’t ever run the first one anymore?
Also there are many cases where you might need to have some special logic after the last iteration, so simply killing the underlying function is not an option.
I give it 10 more years of such special cased improved, for Go not to be any better than the languages the community regularly complains about, while Go is "perfect".
Nice thing about this go approach is you can have just functions as shown in examples. One practical use is functions building range functions on top of existing interfaces. Or as the filters examples shown, you can create filter generators on regular slices/maps.
And yet it's specifically one thing rsc did not want. Further issues described in the rangefunc proposals:
- it would require the desugaring to run off of method-set analysis of userland types, something which does not currently exist
- it severely complicates resource management around the iterator, as you need a 3-step iteration for resource-bearing iterators (acquire iterator, defer cleanup, perform iteration)
And one not actually listed explicitly: for the limited amount of optimisations the Go compiler does, internal iteration is a lot easier to optimise as it pretty much inlines down to a `for` loop, the termination of which is much easier to analyse than bouncing through a bunch of pull calls.
Not only that, but `for range` works off of underlying type, so this is already valid Go:
import "fmt"
type Foo []int
func main() {
f := Foo([]int{1, 2, 3})
for _, v := range f {
fmt.Println(v)
}
}Hopefully it will put the common fallacy of "Go or Rust" to rest as the weight class and capabilities are on opposite ends of spectrum, with much closer comparison being "Rust or C# or Swift or Kotlin" if one looks for a Rust alternative that makes a tradeoff of not forcing many small decision to reduce decision fatigue by conceding to a reasonable extent certain areas Rust excels at.
In any case, for its touted simplicity Go sure doesn't look like a simple and straightforward to follow language anymore.
The 3-step approach to iteration is also well solved[0] in C#, and works even in rather complex cases like `File.ReadLines(...)` where the line iterator internally handles IO, file handle acquisition and disposal. Just `foreach (var msg in File.ReadLines("messages.jsonl"))` and you won't be able to make a footgun out of it.
This also applies to the usage chained with filter/map/etc.
var messages = File
.ReadLines("messages.jsonl")
.Select(line => JsonSerializer.Deserialize<Message>(line))
.ToArray();
[0]: https://sharplab.io/#v2:EYLgtghglgdgPgAQEwEYCwAoBAGABAlAOgEk... note .GetEnumerator() and .Dispose() callsAnd if you do implement interface-based iteration with an Iterator interface, it's not hard to also add a ClosableIterator interface and have the loop handle the auto-close as well.
The only point of methods in Go is to implement interfaces, and that doesn't really work out.
This has been known since the generics proposal.
Of course, other than generic methods, this could also be supported by just supporting universal function call syntax. That is, the compiler could simply take f(x, a) and x.f(a) to be perfectly equivalent, regardless of whether f is a method of x's type or a free floating method. There is some minor complication because of backwards compatibility, but that's easily fixed (the syntax you use can prefer the function f vs the method f if there is any ambiguity).
On the other hand, generic methods can be extremely useful in their own right, for other reasons as well. Having generic methods in an interface, such that a type has to have a generic method to implement that interface, is perfectly reasonable as a feature request - it wouldn't contradict anything in the spirit of Go. Of course, the implementation can have problems and trade-offs, I'm not claiming this is an easy feature to implement. But I don't think it's excluded.
As per my own experience since the mid 1980's.
However, exactly because Go had such an history of language evolution to learn from since FORTRAN came to be in 1958, maybe some of the early decisions could have been done better, instead of Apple style, "we are not doing X, for years later, X is actually something we want".
But it looks the Go core team never care about the consistency problem. The old built-in generics are also incompatible with the new custom generics.
We could also just use an Iterator pattern for such use cases too.
Go used to be such a simple language. I wonder what's driving them to keep adding complexity to the language itself. It makes me respect Harelang's goal of "freeze the language once 1.0 is out" a lot more.
There is no standard way to iterate over a sequence of values in Go. For lack of any convention, we have ended up with a wide variety of approaches. Each implementation has done what made the most sense in that context, but decisions made in isolation have resulted in confusion for users.
In the standard library alone, we have archive/tar.Reader.Next, bufio.Reader.ReadByte, bufio.Scanner.Scan, container/ring.Ring.Do, database/sql.Rows, expvar.Do, flag.Visit, go/token.FileSet.Iterate, path/filepath.Walk, go/token.FileSet.Iterate, runtime.Frames.Next, and sync.Map.Range, hardly any of which agree on the exact details of iteration. Even the functions that agree on the signature don’t always agree about the semantics. For example, most iteration functions that return (T, bool) follow the usual Go convention of having the bool indicate whether the T is valid. In contrast, the bool returned from runtime.Frames.Next indicates whether the next call will return something valid.
When you want to iterate over something, you first have to learn how the specific code you are calling handles iteration. This lack of uniformity hinders Go’s goal of making it easy to easy to move around in a large code base. People often mention as a strength that all Go code looks about the same. That’s simply not true for code with custom iteration.
[0] https://github.com/golang/go/discussions/56413This is why things like ‘close(channel)’ are magic builtin functions, not keywords (more complicated) or a method like ‘channel.Close’ (works with interfaces and consistent with files and such, so not simple).
It's understandable - because unfortunately people judge languages by very shallow metrics. Several times I've seen people use "number of keywords" as a proxy for language complexity.
However, that's completely misguided. `static` in C++ (and, IMO, `for` in Go) demonstrate that overloading a keyword to mean multiple things is harder to understand than having a larger number of more meaningful keywords.
That pointers declared with * in Go are more like references (&) and that there are no true pointers (I think) does not really help.
If you mean pointers arithmetic, that can be achieved with unsafe package.
Being able to do arithmetic is orthogonal to that.
In fact check a language like D, C#, Swift, with references, pointers, and pointers with arithmetic.
The fact that you can store `yield` somewhere allows for more flexible design of iterator functions, e.g. (written in a hurry as a proof of concept so will panic at the end): https://go.dev/play/p/QpVYmmC6g5b?v=gotip
Those hairy details may be hard to remember (or even decide), but they won't matter for most of the users - most users will just use `yield` in the simplest way, without storing them or calling them from another goroutine.
Nothing special. `yield` is just a normal function. Once you realize this, it actually is very easy to reason about. I just think the naming is confusing. I think about it as `body`.
// Only care about the iteration count
for range aContainer { ... }
// Just the values
for v := range myChannel { ... }
// Indexes and values (or keys and values for a map)
for i, v := range mySlice { ... }My understanding is that internal iteration makes it easier to write iterators (producers) but harder to write the consuming code. That's why Go needs to re-write the body of each `for` loop as a function body, including special handling for `break`, `return` etc.
External iteration OTOH makes it harder to write producers but easier to write consumers. Python and C# therefore allow external iterators to be written via coroutines/generators.
Wouldn't Go's goroutines make the coroutine approach to external iterators straightforward? Whereas the re-writing necessary for internal iterators seems convoluted?
Compare that to yield return, .Select and .Where methods in C#, or Filter and Map in many popular languages - it's not a good look.
Or when writing manually, compare it to
static IEnumerable<T> Filter<T>(
IEnumerable<T> source, Func<T, bool> predicate)
{
foreach (var item in source)
{
if (predicate(item))
yield return item;
}
}
Could also compare to how easy it is to use for extremely common patterns in general purpose code: var numbers = Enumerable.Range(0, 10);
var even = numbers.Where(n => n % 2 is 0);
var strings = even.Select(n => n.ToString());They just take the function they decorate, and return another. It does not get simpler yet more flexible than this.
This isn’t a particularly difficult concept, but manipulating functions in this way does feel unusual in such heavily-procedural languages. This has two downsides:
* Programmers who only use those languages are unfamiliar with the concept, and
* The languages’ syntaxes aren’t designed to make it particularly clear.
https://gist.github.com/thwarted/a6e5e7ca5ce552311a7d5ece13d...
To use channels as iterators efficiently, we need find a way to let the functions creating the channels return without creating new goroutines.
An unbuffered channel is really a scheduler abstraction. Consuming from an unbuffered channel blocks, the thread can enter and immediately begin executing the go routine that was blocking on producing. The go routine is acting like a closure around the channel state.
I had some further experiments interleaving these iterators, but didn't clean it up at the time before I had sufficiently convinced myself it was possible and I got distracted with other things.
There’s nothing stopping you from doing this but it does mean you are introducing the requirement of thread safety in your code, in the case where the iterator is stateful.
I would argue anything that needs a range func beyond the simple functional things like filters is probably a stateful iterator (or generator if you’d like), and as such having range funcs is a great way to write code that doesn’t go wrong due to parallelism.
Now you could add two way communication to your channel iterator (or any other locking mechanism) for safety but honestly I think range funcs perfectly solve this use case, and have already used them to keep my code more readable and correct.
All this said, while I’m still a fan of Go and have used it regularly since 0.9 as well as contributed to the language, I will agree with the other comments that sometimes the language design bends over backward to be purist at the cost of having to add more footguns in user land.
The requirement is a limit of the current design and implementation. Channels can be enhanced to avoid the requirement.
Re mutexes: one of Rob Pike‘s Go Proverbs (https://go-proverbs.github.io/) was "channels orchestrate, mutexes serialize“. Both have their places.