Go: Redefining For Loop Variable Semantics
github.com
github.com
functions = []
for i in range(10):
functions.append(lambda: print(f'Hello {i}'))
for fn in functions:
fn()
# Hello 9
# Hello 9
# Hello 9
# Hello 9
# Hello 9
# Hello 9
# Hello 9
# Hello 9
# Hello 9
# Hello 9
I find it one of Python's biggest warts because it's silent, hard to troubleshoot (especially the first time!), and like in Go the most straightforward fix looks like a mistake (i=i): for i in range(10):
functions.append(lambda i=i: print(f'Hello {i}'))
I guess this is a lesson in designing language semantics that match people's intuitions, and learning from previous languages' mistakes. for (var i = 0 ; i < 10 ; i ++ ){
someElement.addEventListener('click',function(){
console.log(i);
})
}
i would always be equal to 9. With let instead of var, i is properly scoped and the executed script display each increment correctly.The solution before let was to introduce a closure in the for statement body
for (var i = 0 ; i < 10 ; i ++ ){
(function(i){
someElement.addEventListener('click',function(){
console.log(i);
}))(i)
}
in order to capture i value. let can also easily isolate the scope of a variable so that it doesn't popule the global scope {
let foo = "bar";
var baz = "qix";
}
// foo is undefined here, while baz is defined.
which removes the need for self invoking functions.If we expand out the loop to be manual we would have a script like:
functions = []
# expanded loop
i = 1
functions.append(lambda: print(f"Hello {i}"))
i = 2
functions.append(lambda: print(f"Hello {i}"))
...
i = 9
functions.append(lambda: print(f"Hello {i}"))
Now at this point if we were to: print(f"Hello {i}")
What would the expected output be? I would posit that anything other than "Hello 9" would be wrong, both logically and intuitively.So by extension a loop of effectively "print(f"Hello {i}")" 9 times should just print "Hello 9" 9 times IMO. Anything else is counter-intuitive and definitely surprising.
def foo(i, j):
i[0] += 1
j += 1
a = [0]
b = 0
foo(a, b)
print(a) # [1]
print(b) # [0]
So people think the body of the loop like a function call.I don't think it's a bad expectation, in fact I think it's quite a natural expectation—in particular if you've programmed functional languages where modifying values is the exception, not the rule—which is why it surprises people. It's just not the one way Python chose.
Which ever makes the most sense depends on the context of a program which is why in languages that have both reference and closure semantics you can choose how the variables are captured. When that isn't the case you need to pick for someone, and it gets weird.
No, this comes down to “should a loop control variable be scoped to the block—or in python’s case function—the loop is in and updated with each iteration or a fresh variable scoped to each loop iteration that happens to share the same name.”
It seems like we're iterating through something to build up a computation we may execute later. I can conceive of a situation where you might want to do that, or at least consider it, but in general I say just do the work now and build a list of results.
It is rare, if only because defining functions (lambdas or otherwise) in a loop is somewhat uncommon.
That's not very helpful when you hit that issue though.
Here's another one, and the way I think about it:
def _loop(i):
functions.append(lambda: print(f"Hello {i}"))
_loop(0)
_loop(1)
_loop(2)
print(i)
# Error!
I highly prefer lexical scoping, where the variables are bound to the block they were declared in, like Javascript's `let` vs the old `var`. This avoids shadowing and general namespace pollution.I know it's not how Python operates, but I think it's how it should. Though I'd argue the syntactical similarities between functions and loops nudge users towards this second model.
>>> functions = []
>>> def _loop(i):
... functions.append(lambda: print(f"Hello {i}"))
...
>>> for i in range(10):
... _loop(i)
...
>>> for f in functions:
... f()
...
Hello 0
Hello 1
Hello 2
Hello 3
Hello 4
Hello 5
Hello 6
Hello 7
Hello 8
Hello 9
> lexical scoping, where the variables are bound to the block they were declared inIn Python that would only work if a "block" included comprehensions. For example:
>>> functions = [lambda: print(f"Hello {i}") for i in range(10)]
>>> for f in functions:
... f()
...
Hello 9
Hello 9
Hello 9
Hello 9
Hello 9
Hello 9
Hello 9
Hello 9
Hello 9
Hello 9
You could fix this by defining a _loop function as above and forming the closure inside it; but changing variable scoping to be lexical in "blocks" wouldn't fix this case unless the list comprehension itself counted as a "block", which is not how Python defines blocks.Since I was confused I did some internet searching and found this:
https://stackoverflow.com/questions/51604346/does-python-sco...
Which I think suggests the thing you are finding confusing isn't lexical scoping in Python, but rather the environment mutability.
Ultimately, I think I better understand what you are highlighting and can't say I entirely disagree. I just personally find how it is today ergonomic, but that could also just be bias as I am fairly comfortable in Python (warts and all).
So, at least in my book, no, Python has "global scope", "function scope" and probably one or two more scopes (I think there's a "class scope" as well).
Here's some code in Python, and some equivalent code in Go.
def foo(a_list):
print(f"list is {len(a_list} elements")
for element in a_list:
print(element)
print(element)
And here's the equivalent Go code: func foo(aList []int) { // Let's use ints...
fmt.Printf("list is %d elements\n", len(aList))
var element int // Notice this declaration! This is ensuring that element is declared outside the lexical scope of the for loop
for _, element = range aList {
fmt.Println(element)
}
fmt.Println(element)
}Python has lexical scoping, but it does _not_ have block level scoping.
https://en.wikipedia.org/wiki/Scope_(computer_science)#Block...
Which, now that I have done enough reading I think crystallizes the confusing thing for others being highlighted here (for me). Depending on preference, the lack of block scoping can be surprising for someone. Which also explains my bias, I started with Python which probably plays a large part in why I find function level scoping without block scoping ergonomic.
For example, equivalent code in Elixir:
i = 0
list = []
list = list ++ [fn -> i end]
i = i + 1
list = list ++ [fn -> i end]
i = i + 1
halfway = list
list = list ++ [fn -> i end]
i = i + 1
list = list ++ [fn -> i end]
IO.inspect Enum.map(halfway, fn f -> f.() end)
IO.inspect Enum.map(list, fn f -> f.() end)
Would produce the functional-intuitive result of
[0,1]
[0,1,2,3]Because there is no mutable state. Those repeated assignments to i and list are exactly equivalent to the scenario where each i was actually i1, i2, i3, etc.
They are both the behavior that’s intuitively obvious. And that’s despite them both being the same behavior.
As we can see, it’s important to not use the simple/obvious implementation because it’s so unintuitive it’ll need to be changed even if breaking (as in C#)
I wouldn't call this a design mistake but more of a misunderstanding of how scope works.
for i in range(10)
The statement literally declares or sets a variable named i in that scope. When the loop exits, i still exists in the scope with the value 9. If you call a function that was given a reference to i, the value will be 9 as expected because the function was called after the loop exited.
Don't see any quirk or mistake here.
For instance the behaviour of a Python loop varies drastically depending on the size of the iteration:
def loop(n):
for i in range(n):
pass
print(i)
loop(10) # 9
loop(0)
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
File "<stdin>", line 4, in loop
UnboundLocalError: local variable 'i' referenced before assignment
That Python works this way is specific to Python. And a language which doesn't have the issues this implies would be "getting it right", in the sense of avoiding sharp corners and edge cases. for f in (lambda: print(i) for i in range(0, 10)): f()
Unfortunately, Python is simply inconsistent in this regard. For example, list comprehensions leak the variable for back-compat reasons, so if you substitute (lambda: ...) with [lambda: ...] above, you'll get a bunch of 9s.But, backwards compatibility aside, the language could change to make for-loops behave like sequence comprehensions wrt scoping.
No, they don’t. They did in Python 2—list comps were introduced in 2.0, genexps in 2.4, and set/dict comps in 3.0 but also included in the later 2.7 release—but that’s been non-current for more than a decade, and conpletely out of support for two years. Let it go.
> But, backwards compatibility aside, the language could change to make for-loops behave like sequence comprehensions wrt scoping.
Sure in Python 4, but after 2->3, not sure many people are looking forward to that.
Python 3.10.5 (tags/v3.10.5:f377153, Jun 6 2022, 16:14:13) [MSC v.1929 64 bit (AMD64)] on win32
Type "help", "copyright", "credits" or "license" for more information.
>>> [f() for f in [lambda: i for i in range(0, 10)]]
[9, 9, 9, 9, 9, 9, 9, 9, 9, 9]
>>> [f() for f in (lambda: i for i in range(0, 10))]
[0, 1, 2, 3, 4, 5, 6, 7, 8, 9]If you instead fully evaluate the generator expression before calling any of the functions (for example, by passing it to the list constructor), you get the same behavior as the list comprehension case:
>>> [f() for f in list(lambda: i for i in range(0, 10))]
[9, 9, 9, 9, 9, 9, 9, 9, 9, 9]Side note: I think that commenters above didn't quite understand what I meant by "leaking", because there's more than one scope boundary here. Roughly speaking, any comprehension or loop can be desugared into something that looks like a C-style for-loop:
/* scope 1 */
for (/* scope 2 */) {
/* scope 3 */
}
Scope 1 is outside relative to the loop. Scope 2 is specific to the loop but shared by all its iterations. Scope 3 is specific to one loop iteration. The "leaking" I referred to above is from scope 3 to scope 2. I think other commenters took it to mean leaking from scope 2 to scope 1 - i.e. the ability to use the variable outside of the comprehension; that is, indeed, something that changed between Python 2 and 3.Bug:
var all []*Item
for _, item := range items {
all = append(all, &item)
}
Fix: var all []*Item
for _, item := range items {
item := item
all = append(all, &item)
}[1]: https://stackoverflow.com/questions/2295290/what-do-lambda-f...
[2]: https://docs.python.org/3/faq/programming.html#why-do-lambda...
In Go example the issue happens because item variable is per-loop, in your python example the issue is not related to loops at all, it's just because functions capture the value of global variables at execution time.
And the cherry on top is that the solution is also similar looking(i=i), but working with a different mechanic underneath(default argument assignment).
Anyway, this was my perspective that led me to interpret this as satire. A bit disappointed haha
def fun(initial_empty_list=[]): where initial_empty_list is a reference captured at function definition time, not a new value initialized on each call to the function.
Edit: looks like Rust does it right: https://play.rust-lang.org/?version=stable&mode=debug&editio...
fn main() {
let mut functions = vec![];
let mut i = 0;
functions.push(Box::new(move || println!("Hello {i}")) as Box<dyn Fn()>);
i = 1;
functions.push(Box::new(move || println!("Hello {i}")));
i = 2;
functions.push(Box::new(move || println!("Hello {i}")));
for func in functions {
func();
}
}
But the most equivalent code would be the following which doesn't compile: fn main() {
let mut functions = vec![];
for i in 0..10 {
functions.push(Box::new(|| println!("Hello {i}")))
}
for func in functions {
func();
}
}It would be better if Python could support block-scoping too. (Maybe it is time for inventing a "strict mode" for Python)
by requiring captured variables to be final it removes a lot of ambiguity around what a variable name refers to. I like that local variables can only be changed locally. If you do want crosstalk between the inner and outer scopes you have to be more explicit and introduce a reference to talk through.
I love python, but I basically avoid this construction and use a single element list if I need it. I can never remember exactly how it works.
the fix was usually adding x := x
before go func() { do something with x }
`go` evaluates everything but the final function call in the context of the caller. So
go func (x int) { … } (x)
Will do the same, and is easier to extract to a named function. go func (x int) { do_work(x) } (x)
seems like a very roundabout way to say go do_work(x)
The "x := x" solution too suffers from this problem but slightly less: both idioms look like they are no-ops (while they are actually not) but at least "x := x" is weird enough to look like it was a deliberate choice, not some vestige from refactoring.Your first version is in fact a worse way to write the second one.
What's even the point of capturing the variable itself? To allow for writing inline callbacks that could sneakily mutate loop-local variables?
That’s how you’d do it using a [=] lambda in c++ or a move closure in rust.
The issue is not restricted to variables declared in loop headers, so the proposed loop change for Go might only be the start.
Technically it’s not but practically it’s by far the most common way for this to unexpectedly arise.
The other cases like closing over a variable and then modifying it before the closure is invoked are a lot less common to hit unexpectedly, and a lot harder to fix nicely (short of Java’s big hammer).
AFAICT it doesn't simplify javac much, if at all. It still needs to synthesize closure objects with fields to store the closed-over values. It's just that those fields can be final.
I think Java did this to avoid programmer confusion. I think it was the right choice.
void foo() {
int a = 1;
Runnable r = () -> System.out.println(a);
r.run(); // prints 1
}
Can get turned into something like this: class r_closure implements Runnable {
final int a;
r_closure(int a) { this.a = a; }
@Override
public void run() { System.out.println(a); }
}
void foo() {
int a = 1;
Runnable r = new r_closure(a);
r.run(); // prints 1
}
The local a is a perfectly normal local, and the field a is a perfectly normal field.What would have to happen if the variable was mutable? For example, if you wanted to write this:
void foo() {
int a = 1;
Runnable r = () -> System.out.println(a);
a = 2;
r.run(); // prints 2
}
You have to transform it to something like this: class r_closure implements Runnable {
int a;
r_closure(int a) { this.a = a; }
@Override
public void run() { System.out.println(a); }
}
void foo() {
int _a = 1;
r_closure r = new r_closure(_a);
r.a = 2;
r.run();
}
Where there is no local, and where the method looks like it's accessing a local, it's actually reaching into the closure and mutating its field!Now think about doing this if you've captured a variable in two closures, or a variable number of closures in a list. The wheels come off this approach.
Instead, you would have to promote the shared mutable variable to its own object, like this:
class int_box {
int i;
int_box(int i) { this.i = i; }
}
class q_closure implements Runnable {
final int_box a;
q_closure(int_box a) { this.a = a; }
@Override
public void run() { System.out.println("q = " + a.i); }
}
class r_closure implements Runnable {
final int_box a;
r_closure(int_box a) { this.a = a; }
@Override
public void run() { System.out.println("r = " + a.i); }
}
void foo() {
int_box a = new int_box(1);
Runnable q = new q_closure(a);
Runnable r = new r_closure(a);
a.i = 2;
q.run(); // prints 2
r.run(); // prints 2
}
Now you've taken a simple local variable which just need to be copied, and turned into its own thing on the heap!Then again, Python has the same syntax for assignment and initialization.
(Note that C# at least had the excuse of not having closures in the first version, which makes scoping of "foreach" moot - the problem only showed up in C# 2.0. But Go had lambdas from the get-go, so this interaction between loops and closures was always there.)
Still looking forward to the day they will discover Pascal enumerations.
Not just a "general rule" that document also specifically talks about precisely this issue (for loops) and resolves that Go will not fix this.
Change is the only constant, we should design systems with the expectation that they'll need to adapt over time or they will be replaced by something which can. With this mindset, Go should have solved the for loop problem years ago, just as C# did. This could have been a story about how once upon a time Go had these very silly for loop semantics, but that hasn't been true for many years.
This necessity of change is why I think the decision not to take Epochs for C++ 20 was much more consequential than things like rejecting the "Goals and priorities" paper which had immediate effects (in that particular case spurring the Carbon experiment).
Just like we'll rarely pick exactly the right behavior, we'll rarely pick exactly the right time to fix the broken behavior.
I'm interested in the logical arguments that would support this. I can't think of any.
No?
> Otherwise it'd been "Arguably if Go had done this years ago when go.mod did not exist, this would have had an even bigger impact" or something like this?
That’s a worst take on the same sentiment? I find your version a lot clunkier. The original sets a hypothetical stage and from that its conclusion, I think it flows better.
The stage is "years ago", and conclusion is "go.mod would not have existed and this would have an even bigger impact". My version re-arranges the sentence so that "go.mod not existing" bit is a part of the premise, not of the conclusion.
Tense agreement in English subjunctive is hard. Especially for non-natives such as me: I do parse the original statement like that and just can't bring myself to understand it otherwise.
If the comma had instead been the word "when", as suggested, this would parse the other way. It still would have been a bit awkward but would make sense.
IF loop variable syntax had been redefined years ago
THEN
go.mod would not exist
AND
impact would have been larger
It seems from this thread that what they meant to say was: IF loop variables had been redefined BEFORE go.mod existed
THEN impact would have been largerBut a language more comfortable with change implies that change is more likely - not less likely. Given pretty much every modern language has some kind of dependency management utility - I'd be surprised if Go didn't end up with one
How did they plan to make Epochs compatible with headers being copy&pasted at the #include site? And what about template instantiation?
https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2019/p18...
Similar changes were made in newer versions of ES so that the for loop in this article works out of the box, like C#.
Slightly off-topic to this article: I wish "do while" loops had the "while" condition in the inner scope, not the outer scope. So many times I have wished that I could access the inner scope... I end up using a while(true) with an if { break; } at the end instead in 99% of cases where a do while could've been the perfect thing...
As an example:
x = [];
for (var i of [1, 2, 3]) {
x.push(() => i);
}
x.map(f => f()) == [3, 3, 3]
x = [];
for (let i of [1, 2, 3]) {
x.push(() => i);
}
x.map(f => f()) == [1, 2, 3]
What’s being proposed for Go is instead a breaking change.given the context, it might actually be a fixing change...
It sounds like you only get the breaking change in modules that require above a certain version of Go. So this should not break old code. It's perhaps more analogous to the way that "use strict" in JavaScript 'breaks' parsing of octal constants.
Then this seems fairly reasonable, given also how rarely things will depend on one way or the other.
It’s a breaking change in the same sense that the `”use strict;”` semantics of JS were: it’s not actually a breaking change, because you have to opt in.
for (var i=0; i<3; i++) setTimeout(() => console.log(i))
Would print 3,3,3.
So in ecmascript 6, blocked-scope variables were introduced, but the semantics of old-style "var" declarations was not changed. You can now write for (let i=0; i<3; i++) setTimeout(() => console.log(i))
Which prints 0,1,2 as expected.And you don't have to go open a `go.mod` file to know what the code you are reading does.
Luckily caught my mistake and fixed it within 15 minutes, but it's always easy to overlook.
The gradual breaking (of fix depending on your point of view) with explicit opt-in looks great to me.
The bar is quite low.
- zero values instead of sum types
- nil, especially on interfaces
- goroutines sharing memory
- goroutine panics not being raised in the spawner by default
(For me, the link you posted does a 302 redirect to https://www.uber.com/pt-BR/blog/data-race-patterns-in-go/ which gives me a 404 error page. It's a bit insane that whether the link you posted works or not depends on your locale, and unfortunately this is not the first time I've seen this kind of baffling redirect misbehavior.)
in this code:
var all []\*Item
for _, item := range items {
all = append(all, &item)
}
When &item is the same for all iterations, that means that it's pointing to the same memory address. Is each item in items copied to this address prior to each iteration body invocation? This seems strange as this copy could potentially be very expensive. What am I missing?Most of the time you're iterating a slice of pointers, though, so only the address gets copied. And in those cases, this bug doesn't exist(unless of course you're going from * to ** for some reason).
for i := range items {
ptr := &items[i]
...
}
In this case of course you also get different semantics. The ptr variable is bound to the address of each of the original Item values in the slice, whereas in the code in your comment, &item is the address of a single heap-allocated Item variable. var item Item
item = items[0]
all = append(all, &item)
item = items[1]
all = append(all, &item)
...I suppose it's the natural way to do it when implementing a language and not thinking about it too much. It makes a for-loop equivalent to a simple while loop with the loop vvariable initialized outside of the loop.
TL;DR is that both "for" and "foreach" scoping fixes would be breaking changes, but "foreach" was easier to justify because it was already a C#-specific construct syntactically, unlike "for" which uses the same exact syntax as C, Java etc, and they were very sensitive to backwards compatibility at the time (esp. since the tooling didn't have the ability to target various language versions within the same project easily). At the same time, "foreach" represented the vast majority of breakage when they looked at existing code, perhaps because the scoping in classic "for" is more obvious due to the fact that variable mutation is explicit there.
>>> def foo():
... return [lambda: x for x in range(10)]
...
>>> [x() for x in foo()]
[9, 9, 9, 9, 9, 9, 9, 9, 9, 9]A common sentiment when Go was unveiled was that its designers had ignored the 35-odd years of language research and experience since C.
>>> [f() for f in (lambda: x for x in range(10))]
[0, 1, 2, 3, 4, 5, 6, 7, 8, 9]
List comprehensions have been around for longer, and so behave the way they do for backwards compatibility reasons.It became an issue as lambdas and other lambda-type constructs (which implicitly keep a reference on the loop variable) became more common, and a bunch of languages got caught in it. Later languages switched to the “inner scoping” mechanism to avoid it.
Go did not follow the switch because Go.
nums := []int{1, 2, 3}
for _, num := range nums {
fmt.Printf("%p\n", &num)
}
This will print the same address three times. If you add "num := num" as the first line in the loop, it will print three different addresses. The proposal is to make this the default behaviour.Yes it is.
> The variable is scoped to the loop.
Maybe “loop body” would make the comment clearer?
> The issue is that even if you take the address of the variable, each iteration of the loop will have the same variable with the same address.
Hence the variable not being scoped to the loop (body), one variable is shared between all iterations of the loop.
I think "loop iteration" might.
for x in range(3):
foo(x)
print(x) # prints 2
or the pre-C99 style: int i;
for (i = 0; i < 3; ++i)
foo(i)
printf("%d", i); // prints 3
This is not the case in Go. I don't think talking about variable scopes accurately describes the issue (because there's nothing special about loop scopes here: the same "escape" can happen from any scope), and changing "loop" to "loop body" doesn't improve this. The term "loop iteration" at least identifies the dependency between different iterations of the loop as the issue.I don't understand what the thought experiment about non-lexical scopes has to do with this.
for _, elem := range elems {
elem := elem
... &elem ...
}
Nothing beyond regular lexical scoping and Go's ordinary assignment semantics are necessary to see how this works. The second 'elem' has a narrower scope than the first (it is limited to the loop body). Abusing Go syntax, you can think of the current semantics as follows: {
var elem Elem
for _, elem = range elems {
... &elem ...
}
}
Here 'elem' scopes outside the loop body, and so is reassigned on every iteration of the loop (and &elem evaluates to the same address on every iteration).>the thought experiment about non-lexical scopes
It's not just a thought experiment. There are languages with dynamically scoped variables (e.g. global vars in Common Lisp).
I guess taking the address "promoted" the shared variable to enable it to survive past the function?
1. Java is also supposed to do that, and clearly way less successful
2. Escape log is part of the standard baseline tools (it’s just a flag), so it’s not considered anything hidden or arcane
3. Go tends to log escapes, meaning it’s baseline is to stack allocate
So it’s a lot more than, say, common subexpression optimisation.
The default is heap, and only when the compiler can be sure it's safe, things go on the stack.
initVariable()
while checkVariable():
doSomethong…
updateVariable()
So the variable is shared by all iterations at first place. And that isn't a issue until we have lambda or something similar that can capture reference of loop variable.Later languages found it is problematic when use with reference capturing features and changed to something else (one variable per iteration)
But the problem should have been known when the language was designed. In CommonLisp, i.e., quite a mighty but old language, it is the same: the capture is on a loop variable that is destructively updated:
(setf refs (loop for i in '(1 2 3) collect #'(lambda () i)))
-> (#<function) #<function> #<function>)
(mapcar #'(lambda (f) (funcall f)) refs)
-> (3 3 3)
The same done with a fresh function parameter works. And since this is the usual style in Lisp, I suppose the problem will not be that obvious like in Go: (setf refs (mapcar #'(lambda (i) #'(lambda () i)) '(1 2 3)))
-> (#<function> #<function> #<function>)
(mapcar #'(lambda (f) (funcall f)) refs)
-> (1 2 3)Complexity increases rapidly when you combine constructs in a programming language. You get some feature interactions which are hard to get right, and also to predict.
Here you can see that the i from the previous iteration gets copied first and i++ applies to the next iteration:
for (let i = 0; i < 10; i++) {
setTimeout(() => {
console.log(i);
}, 1000);
}
// prints 0...9 (as expected?)
I always thought of the i++ happening at the end of the previous iteration, but that's wrong as it produces a different result if written explicitly: for (let i = 0; i < 10; ) {
setTimeout(() => {
console.log(i);
}, 1000);
i++;
}
// prints 1...10 (as expected?)
(Also, these don't change if I remove the curly braces, so the let is not scoped within curly braces as I thought...) for (let i = 0; i < 10; ) {
setTimeout(() => {
console.log(i, foo);
}, 1000);
let foo = "baz-"+i;
i++;
}
// "1 baz-0", "2 baz-1, etc
It's behaving as each loop iteration is creating its own closure, and the inner function is referring to those variables by reference. So any changes you make to them inside the body of the loop will end up visible to the inner function.I think to then understand why my two examples produce different results, you have to know that the i++ in the header happens to be executed after that copying has been made (an arbitrary choice?), while the i++ within the body will be (naturally) executed before the copying.
I suppose the way it works can be intuitive, but it can also be confusing if you think of the i++ as the last action of each iteration in both of my example cases.
Actually, there's a third way to write the example loop, but who can guess which result it gives?
for (let i = 0; i++ < 10; ) {
setTimeout(() => {
console.log(i);
}, 1000);
}
... it will print 1...10! So even within the loop header, one part is run before the copying and another part after the copying. How is this intuitive and where is this documented apart from the language spec?EDIT: My bad, of course the iteration condition has to be checked at the beginning of the iteration, so in this third example i == 1 during all of the first iteration).
Here's the relevant spec section 14.7.4.3 ForBodyEvaluation: https://tc39.es/ecma262/multipage/ecmascript-language-statem...
There we can finally see that on the first iteration, an environment is created (and iteration variables copied) before the test ("step 2"). After that, a new environment is created (and iteration variables copied) towards the end of each iteration, before the increment step ("step 3.e" and "step 3.f").
Variations of such perceived weirdness exist in many other languages with complicated "object models" as well to be fair. Delphi has some strange adressing stuff going on as well. Python has this weird "default list" thing. Most object languages don't let you take the address of something (like the Go example shows) at all, but have only object references which I find unergonomic.
Maybe that's my C/C++/Rust experience.
If append() needs the thing, not just an immutable reference which expires soon, its signature would demand we move one into it, and we don't have one so we'd need to e.g. make one with Clone.
That's what go does under the hood in 95% of cases. If I do:
func foo(a int) *int {
return &a
}
func main() {
fmt.Println(*foo(3))
}
It does the right thing. It's just that in this case (and in lambdas) they got it wrong.- experience with the issue in languages with wider scoping e.g. it’s a common issue in JS, as well as Python (though slightly less so) - Rust’s iterators were originally internal so that was pretty natural
(1) is also why for(let and for(const have different scoping than for(var in JS: `var` has function scoping, `let` and `const` were introduced with block scoping, and for loops they were specifically specced with “inner” (per-iteration) scope.
It's simply a very bad idea that provides no use yet creates many bugs.
Before the early aughts, closures were mostly really common in functional languages which tend towards immutable bindings (and immutability in general), and very closure-focused languages closured everything so didn’t hit that issue (e.g. you wouldn’t hit it in Smalltalk because your counter would be a parameter to a block, so closing over that was no issue).
It’s really in the 00s with the explosion of callbacks-pile-javascript (and more generally the functionalisation of imperative languages) that the problem became a serious concern: you loop over a thing, you start some sort of async operation (network request for instance), and you find out that despite the request being correct the entire thing goes wonky (then again things commonly went wonky which didn’t help).
It's a problem with references in general as this shows.
I also don't feel it's easier to implement at all.
One can either rewrite:
for $id:var in $exp:iter { $code:body }
to: { let $id:var;
while(True) {
let result = $exp:iter.next();
if(result.is_none()) break;
$id:Var <- result.extract();
$code:body
}
}
Or while(True) {
let result = $exp:iter.next();
if(result.is_none()) break;
let $id:var = result.extract();
$code:body
}
The latter implementation is as far as I see easier, not more complex. Obviously all the code to create scoping already exists in the compiler and for-loops over an iterator work with a syntactic rewrite to an infinite loop with a break.Most languages don't have references, and in those that do before the issue was understood, the explicitness made it a much smaller issue.
> I also don't feel it's easier to implement at all.
> One can either rewrite:
Now try lowering to bytecode or assembly instead of high-level pseudocode.
But Go and C++ do, where this issue arose with or without closures.
> Now try lowering to bytecode or assembly instead of high-level pseudocode.
It doesn't matter, because as I said, all that is already in the compiler.
It would be needlessly complex and error-prone for compilers to hardcode custom code generation for such abstractions; it's transformed to something else the compiler already understands at a far higher level. I know for a fact that in Rust, for-loops already desugar to a simple infinite loop construct with a break at the H.I.R. level and all further optimizations only happen from there.
The only explanation I see is that they really gave it no thought at all whatsoever and it wasn't a tradeoff but simply not thinking clearly.
[Edited: I tried to explain what's going on here, but I don't think my explanation was helpful so I've just left the surface]
Yes Rust’s ownership rules make it rather complicated to reproduce the faulty behaviour, as it’s about sharing mutable state which Rust intensely dislikes. You’d need to wilfully share (and update) internally mutable structures (cells, atomics) which is pretty noticeable and not something you do by mistake.
I'm not sure prefixing your comment with "WTF" and your 10-years-ago dismissal helps the discussion here. Yes, as we've learned, this was probably the wrong decision, but it's not hard to see why it was done that way originally (C# made the same decision), and now they're having a reasonable technical discussion to try to solve it. And -- even though I've been bitten by this several times myself -- it's not a terribly common occurrence.
This was basically Dijkstra's point in his BASIC considered harmful post... I think in 2022 it should be C considered harmful for the same reason. C is not the base truth of computation. It isn't even very good. A language smart enough to analyze taking pointers and notice it can't put something on a stack and simply take care of it is, in my opinion, the one that is not catching you off guard... specifically, the "guard" that one must take in C around what is stack versus heap.
That history makes it feel like "taking the address" is a really trivial operation - returning a numerical value that the compiler had access to at that point anyway. Here it's adding a reference to the object in some sense, and maybe even changing how it's allocated earlier in its lifetime (on the heap rather than the stack). I don't use Go and I agree that using &x for that operation feels a bit wrong as an outsider.
It also occurs in langages of category (2), specifically C++ lambdas where i think it can cause UAF/UB. I assume it also happens in C with the block extension (is that still Apple specific?) though I don’t know the details of that thing so maybe not.
Go does in fact stack allocate variables which it can prove not to outlive their lexical scopes, but this is merely an optimization. Unless you are trying to write optimal code, there is never any reason to think about which values are stack allocated in Go.
There's not really any such thing as a 'local variable' in Go. A variable has whatever scope it has, but there's nothing special, semantically speaking, about variables defined inside functions or inside loops.
If the use of & in the example code is puzzling, it's probably because you're expecting Go to have some C-like concept of an automatic (i.e. stack allocated) variable – but it just doesn't.
>That history makes it feel like "taking the address" is a really trivial operation - returning a numerical value that the compiler had access to at that point anyway.
It is in fact a trivial operation in Go too, as I hope the above has clarified.
---
† Strictly speaking 'semantically heap allocated' is nonsense, but hopefully you know what I mean. There is no way to declare a variable in Go in such a way as to force it to be deallocated at the end of a particular lexical scope. A variable's lexical scope and its lifetime are entirely divorced (as is typical in a GCed language).I never used the word "special". As you say, adding a reference will mean the garbage collector won't deallocate it (until that reference is removed). In other words... its lifetime is extended. That's exactly what I meant.
> If the use of & in the example code is puzzling, it's probably because you're expecting Go to have some C-like concept of an automatic (i.e. stack allocated) variable ...
Not at all. In C++, you can use & on a reference variable and it will return the address of the object being referred to, regardless of whether it is allocated on the stack or the heap (or even statically allocated). Even in C, you can do &*x on a pointer to any object (which is silly by itself, but useful when combined with pointer arithmetic e.g. &x[3] translates to &*(x+3)).
> It [the & operator in Go] is in fact a trivial operation in Go too, as I hope the above has clarified.
Maybe I should have avoided the word "trivial" as its meaning is subjective, but I was careful to define what I meant by it: "returning a numerical value that the compiler had access to at that point anyway". Your comment just confirms that, as I said, it does more than that – it also adds a reference to the object.
---
To be clear, I'm not saying that it's bad or wrong that Go uses the & operator to mean this. Once you're familiar with the language, you probably get used to it very quickly. My point was just that it's a surprise initially if you're not familiar with the language, that's all.
It simply evaluates to the address of the object, just as it does in C. if you think the & operator is doing something in addition to this, I think that must just be based on a misunderstanding.
I am not quite sure what you mean by 'adding a reference' to the object.
Let's take this function:
func foo() *int {
var x int
return &x
}
All that happens is the following:- An integer is allocated (and initialized to zero).
- The address of this integer is returned.
If we dig into the implementation, we'll see that the integer is allocated on the heap. As far as Go's language semantics are concerned, everything is allocated on the heap and left to the GC to clean up.
As an implementation detail, values that provably don't outlive their containing functions are (sometimes) stack allocated. As x outlives its containing function, it won't be stack allocated. That's it. There is no special operation of 'adding a reference' or 'extending a lifetime'. Nor does the compiler even analyze lifetimes except for the purposes of applying an optional optimisation which has no effect on the semantics of the program. If you turned this optimisation off (which you totally could) then there'd be no need for the compiler to worry about x's lifetime at all.
> - An integer is allocated (and initialized to zero).
> - The address of this integer is returned.
That is not all that happens, at least down at the C/assembler level.
Let me illustrate what I mean. Consider this function, which also does both of these things (cobbled together from Google searches so please excuse incorrect syntax):
func foo() uintptr{
var x int
return uintptr(unsafe.Pointer(&x))
}
All that function does is allocate an integer (and initialise to zero) and return the address of that integer. Exactly the same as your function, right? Except it's obviously not - it doesn't extend the lifetime of the integer variable.So why not? The GC somehow knows to ignore the number returned from my function, even though, under the hood, it's still stored in a register or stack location or whatever in exactly the same way as the address returned from your function. So how does the GC know to ignore it? Is that number somehow marked in a way that says "GC, when you're scanning memory looking for address-like numbers, don't pay attention to this one"? No. It doesn't look at the number in the first place because it hasn't been told to look at it.
In contrast, in your example, the memory address is not just returned from the function (in the C sense that it's put in a register for the caller to receive). It, additionally, somehow registers that memory address with the GC to let it know that there's another reference to that variable location. That is the extra thing that your function does that mine doesn't. And that magic happens (or at least starts) at the moment you use the & operator.
The Go GC isn't a reference counting implementation. It traces the values of variables on the stack and it knows their types (because it knows which function any given stack frame corresponds to and it knows which variables that function allocates). Thus it knows that if a variable is of type *int and has a non-nil value then its value references an int. (And so on for fields of structs that are stored in stack variables, etc.) The & operator does not need to do anything special. The & operator merely takes the address of the object. When that address is stored in a pointer variable (or array member, or struct field...), that's when it becomes visible to the GC as a reference.
It's not just this bizarre gotcha (the fact that C# had it too doesn't make it OK). It's that Go has so many of these cases where they took very strong positions on things and then later reversed their position only after many, many years:
- Generics
- Only one gc knob
- No backwards incompatible language changes
Also, Go has been around quite a long time now and we've all read quite a few rants about its surprising cases. How comes this one never came up before? The thread provides evidence that it bites people regularly. It suggests to outsiders that you can't easily evaluate Go by reading about it because there will be sharp edges that people aren't talking about simply due to the quantity of things that are even worse.
Personally after using it for 10 years I've been bitten very rarely by weird corners of the language and have enjoyed using it. My complaints are more around things I'd rather see removed (struct tags, panic, nils) and inconsistencies (built-in generics were quite limited, I quite like the design for generics they came up with though so I guess that is resolved once they update the stdlib).
Overall it's still my favourite language compared to others I'm forced to work in, I particularly like the decision to eschew inheritance.
But Go did have closures initially, and worse yet, they already had C# as an example of how closures and loops interact. So they definitely had the opportunity to learn from that mistake, and I don't think it's unreasonable to ask why they did not.
But this is an important point: They were aware of it; it was by design; they had the chance to change it before 1.0 and didn't, and now they are showing willingness to change. So many languages are resistant to change on the one hand (I mean Go is), yet not resistant to keep adding features (e.g. Java / JS).
let is preferred nowadays in JS and doesn’t have the weird hoisting behaviour that var does/did. JS has neither “range” nor pointers though so I’m not sure what you mean by that.
I have been in situations where I have had to expand a macro to figure out what is going on. Not having it in a situation like this (where a for loop is obviously just a goto or a tail call) is usuallly a pain in the ass. If it was translated into the same language it would also be easier to define what it should translate to.
In CoffeeScript (which I think solves this nicely), it looks something like
for i in [0...3] # 3 3 3
setTimeout (() -> console.log "#{i}"), 100
for i in [0...3] # 0 1 2
setTimeout (do (i) -> () -> console.log "#{i}"), 100
for i in [0...3] # 0 1 2
do (i) ->
setTimeout (() -> console.log "#{i}"), 100It's up to the developers to capture loop variables or not.
In the cases where the new value isn't appended, the compiler can easily optimize it away, and reuse the same memory location in a register or on the stack. My guess is that the optimizer portion of the compiler is at a place nowadays where this will happen, with no further change needed. When the compiler was new, it might have been a regression in efficiency.
The alternative can be that whenever address of iteration variable is used inside the loop the variable is per iteration and otherwise it is per loop. This way it is not breaking the old code and have new semantics.
It's one per module though. In large enough modules you'll have tens of x:=x though. I assume this also opens the option of doing more such changes through the same system in the future.
The “static analysis” section says that it is impossible to catch all cases where address is used which is true. However if the analysis checks for whether address is TAKEN, then it is trivial. I would like to propose that as an alternative- whenever address is taken the variable is per iteration. Otherwise per loop.
This seems the simplest and logical modification.
That wouldn't fix the closure problem where the address isn't taken.
Would it be possible to add closure use as another case?
also anyone know of any proposals around generics or pattern matching?
Generics were released in March. There are no active proposals for pattern matching.
If you track the issues closely on GitHub, there's a constant low-level discussion going on about it, but so far nothing seems to be both important enough and large enough to justify the version change, like, not even close.
Closures btw is a sucha horrible pattern. Added in many langs. Always banned by corporate guidelines. The fib example on the GO site is a good example on how confusing it can be. It is right up there with promises and other trash.
What does this do? That should be a no-op by default in any sane language.
var item *Item = item
It creates a locally scoped variable that shadows the existing one.On the left side, "item" resolves to the "item" in inner scope (that is declared on the same line). On the right side it resolves to the "item" in outer scope.
Then I understand what's happening though I'm still not convinced it's sane behaviour. I would expect all identical names in a single scope to refer to the same variable.
That is a no-op
a := b
Declares a new variable a and assigns it. It's sugar for:
var a = b
So read it as:
var item = item
Many languages have scopes and variable shadowing. You can argue this is bad practice or whatever, but it's only confusing if you think := is assignment which is covered in the first page of the Go language tutorial.
var item = item
should still be a noop imo. If it does some magic like dereferencing, make it explicit. And no, other languages making the same mistake is not a good excuse. Especially not for a language focused on inexperienced newgrads.That said, it looks and feels dirty and buggy and it's a known workaround for a Go issue, so I'm glad they're at least talking about fixing it.
What does the print function exactly do if I pass a variable to it.
The statement by itself doesn't produce any visible side effect.
However it creates a new logical variable in the abstract machine, and that can have real consequence in the real underlying machine, depending on which statements happen next.
In particular, the new variable and the old variable are independent and thus the new variable may require to allocate some storage location (e.g. on the stack, a register, ...) to keep track of further mutations to its value (I say "may" because an optimising compiler may do without that extra location)
item := item
item := item
the two lines doing different things even though looking exactly the same.I'm a Scala dev and in Scala we have a similar thing, where a new import or definition in the middle of the code can change line semantics between two equal lines.
But the difference is that this is not a mistake but a concious design to switch "contexts" and is heavily guarded by the typesystem, where GOs typesystem is incredibly weak in comparison, which is not an inherent drawback, but here it is.
I do feel that I'll stick to my feelings that this design is confusing and should have been avoided from the beginning. Ku udos to the go team for breaking compatibility here. This is necessary for a language to not become another COBOL.
item := item
item := item
"no new variables on left side of :="You can't declare the same variable twice in the same scope.
item := item
works only because it declares a new variable in the current scope, initialized with a variable with the same name (but different variable) from the outer scope.Go distinguishes declarations from assignments. In this example, the second assignment is a no-op indeed.
item := item
item = item let x = ... x ...
and every such line is a declaration of a new variable that shadows the preceding one. let x = 0 in
let x = x + 1 in
... let foo = foo.clone();
let bar = bar.into_inner();
let baz = baz.to_owned();
let quz = quz.unwrap();
... my_goose : Goose = get_a_goose();
... lets us give my_goose an explicit type, while we can write: my_goose := get_a_goose();
... and leave the compiler to infer that my_goose is a Goose because that's what get_a_goose() returns and we didn't specify.The idea here is that we have this flexibility but we didn't burden the language with what feels like an extra feature (like C++ auto) with its own special rules. Given that := starts out as a single token this is in some sense revisionist, but so long as you're trying to learn to program, not studying the history of programming, that's fine.
It reminds me of how eggcorns can become language features. A person incorrectly analyses a word or phrase they've heard, e.g. they think the things which fall off an oak tree must be named "eggcorns" because they look a bit like eggs. They apply this analysis, and, if the results are successful and out-competes existing correct analysis, it can dominate, next thing you know† your spelling correction tool says "Did you mean eggcorn?" when you type acorn.
† In reality these transitions usually take generations
One shadows the variable by a variable of the same name. The key is that the new variable is scoped only to the inside of the loop, whereas the original variable is scoped to the outside of it, thus assigning to the original variable in a next iteration of the loop updates that variable again.
The major case where this is an issue is if one somehow took a reference to the variable, that reference thus now points to the new value rather than the old.
Taking a reference to the new, more closely scoped variable that is local to the iteration of the loop does not have this problem.
of Turing's Tarpit
I giggle a bit
http://www.lispworks.com/documentation/lw51/CLHS/Body/m_loop...
(defvar refs
(loop for i from 1 to 10
collect (lambda () (format t "~d~%" i))))
(loop for f in refs
do (funcall f))TBH, with my full sincerity, this isn't really a way to solve the problem. Currying has been used for decades for reasons. For-loop isn't the only place where you introduce outer-scope variables into an immediate function. You can always have other variables involved, and you'll still have to be careful. Even though this change will reduce the absolute number of bugs, people will still have the same issue.
What's that have to do with for loops?
Maybe you meant partial function application? Even that doesn't seem relevant, given that the function is actually being called immediately in the first example.
I also think "currying" and "partial application" kinda-sorta make sense. Unbound variables in a "immediate function" are merely hidden parameters of the function if analyzed semantically. The question is whether you pass them by-ref or by-value.
Partial application is the better solution in imperative languages.
It is also not relevant to this discussion in the slightest, because the behavior of allocation and how those are passed into closures is independent of the question of how you spell those closures in code. You can easily build currying or some syntactic partial application into Go but leave this problem there, and you can easily fix this problem without building currying or partial application in.
This seems like a weird take. Currying was a completely natural thing to do in SML for example, why would it only be "useful" in Haskell ?
I agree that currying isn't likely to be the solution to any problem you have in say, Rust or Java. But "only useful in Haskell" jumped out as a weird claim.
(*) please don't reply with "then you should keep the version below 1.20"
That is literally the solution. Just remove your go.mod or keep it under 1.20, and you’ll keep the old behaviour forever.
It’s the same as e.g. Rust’s edition, if you hate your life you can keep using Rust 2015 semantics to this day, and forever.
Think for a second, if the go.mod stricture locked in the entire thing, you’d just be told to not upgrade.
The entire point of the system is that it’s possible to update language-level semantics while allowing for the rest to progress for everybody.
Setting go.mod means you get stdlib updates, but the semantics of the for loop does not change. And if the designers decide to add similar breaking changes in the future (change string literals or whatever), you’ll also be ignoring those.
> Setting go.mod means you get stdlib updates, but the semantics of the for loop does not change.
So you miss out on new features, or have to update your code base. We're back to square 1.
You know what it does, but you don't grok it.
Clearly you feel that setting go.mod is for other people and the language should match what you want.
What the fuck?
var all []*Item
for _, item := range items {
all = append(all, &item)
}
I can't think of anything that isn't terribly contrived for the sake of argument.Another comment says that Russ Cox did the experiment, and did encounter a few problems. Not many, but they do exist.
> Of the failures, 36 (62%) were tests not testing what they looked like they tested because of bad interactions with t.Parallel: the new semantics made the tests actually run correctly, and then the tests failed because they found actual latent bugs in the code under test.
And if you want to keep your old bugs in under-tested programs, you can, that's why the new behaviour is opt-in.
So far as I know, they didn't find any. They did, however, find a lot of already-broken code that authors didn't realize was broken, and that would be fixed by the change.
With that in mind, while it's technically a backwards-incompatible change, I'd not go as far as call it a breaking change, since it was a net positive to make the change.
C# did it 10 years ago and it did not cause any significant problems in the community, mainly gratitude (as the C# guy in the replies mentions).