Common Gotchas in Go
deadbeef.me
deadbeef.me
My favorite gotchya is that assigning a nil pointer to an interface will not give you a nil interface.
Once you know what's going on, it totally makes sense, but usually you don't start out with that deep an understanding of how interfaces work, and it violates natural-looking assumptions about algebraic equality, which makes for very easy counter-intuitive behavior.
Example:
package main
import "fmt"
func main() {
var x *ErrorX = nil
var err error = x
if err == nil {
fmt.Println("nope, this won't happen ")
} else {
fmt.Println("here's my non-nil error:", x)
}
}
type ErrorX struct{}
func (x *ErrorX) Error() string {
return "hi i'm ErrorX"
}
https://play.golang.org/p/BB1b2_HDb6IThere are of course also the usual gotchas when capturing things in a closure (is it by value or by reference?). But this is not Go specific.
You explained the main Go-specific Gotcha in my opinion: The dual-nullability of interface types, and the fact that non-null interface objects might contain null pointers to structs that implement the interface.
Nah. They wanted references, they got references. The issues were an oversight related to scope.
Anything “makes sense” from an operational point of view, so long as you internalize the rules. But does Go's handling of `nil` denote anything you ever actually want to express, regardless of language?
> it violates natural-looking assumptions about algebraic equality
No, Go's handling of `nil` doesn't violate anything about equality: everything (besides floats, but you can't fault Go for those) is equal to itself and to nothing else. But the syntax (i.e. using a single unadorned token `nil`) does a very poor job of reflecting the language's semantics (the existence fo multiple `nil` values, all different from one another).
And I wasn't trying to say nil-to-interface violated anything real, just that it violated natural-to-make-but-actually-untrue assumptions (which is I guess pretty much the definition of a gotchya):
If we have
a = nil
b = a
it's natural (but incorrect) to assume that we have b == nil
Basically, I agree with you that the syntax is being poorly suggestive of the semantics. I like the "unset" suggestion because then we'd have: a = nil
b = a
which wouldn't falsely suggest b == unset
Too late now, though.Under the covers, interfaces are implemented as two elements, a type and a value. The value, called the interface's dynamic value, is an arbitrary concrete value and the type is that of the value. For the int value 3, an interface value contains, schematically, (int, 3).
An interface value is nil only if the inner value and type are both unset, (nil, nil). In particular, a nil interface will always hold a nil type. If we store a nil pointer of type * int inside an interface value, the inner type will be * int regardless of the value of the pointer: (* int, nil). Such an interface value will therefore be non-nil even when the pointer inside is nil.
BTW, in fact, the ", _" in the author's code can be omitted.
for idx := range zoo {
zoo[idx].legs = 999
} for i, x := range xs
to have x repeatedly be a new reference to each of xs members, it seems like we'd avoid 2 gotchyas (capture of x by a closure inside the for loop not doing what we want, and #1 in the original article).I think the downside would be the concept of references doesn't exist anywhere else in Go (so the "for ... range" loop would have to be its own thing, rather than explainable as a transformation to simpler code) & maybe also implementation concerns I don't know about.
But if we were willing to pay the complexity cost it would make the language nicer, IMO. (I'm not saying we should be willing to pay the complexity cost though -- a lot of Go's charm is how much the language leaves out.)
I'm not convinced. I believe what you'd end up with is that someone will write a "Common Gotchas in Go" article about how, if you think you are operating on a copy in the loop, you are actually operating on a reference.
Really, it seems very non-obvious to me, why one would be a less surprising behavior than the other.
(The fact that the loop variables are not scoped to the loop body - i.e. closures will share the references - is another issue and that I would pretty unambiguously call a gotcha that should be fixed…)
It actually pushes you to design things in ways you might otherwise not. E.g., instead of []X you'll have []*X just so it'll be easier to modify from inside a loop.
Feels differently to me :)
> It actually pushes you to design things in ways you might otherwise not. E.g., instead of []X you'll have []* X just so it'll be easier to modify from inside a loop.
Never do that. Instead I use indexes when I actually want to access the element.
(Where I do do that is in maps, but not because of range, but because index-expressions over maps are not addressable)
Yeah, it's pretty clumsy though.
Most programming language gotchas are literally, "just spend {small time increment} learning about {thing}." Those things build up and get lost/misplaced. Programming "nickels and dimes you to death." Design is about the emergent interaction of myriad details. Programming language design is exactly that as well.
the slice examples doesn't qualify because it's integral to the fundamental idea of what a slice is. they only way this could confound your expectation is if you'd literally spent zero time learning what a go slice is.
(so i guess it is a gotchya if you see "slice", know python slices, and assume the concepts are the same and don't bother to investigate any further.)
by the way, i agree with you that "Design is about the emergent interaction of myriad details"; by that criteria go is very well-designed...
Go lacks const, so there's no way to prevent your caller from modifying the slice you return, or the callee from modifying the slice you pass it. Also the result of append() may or may not alias its parameter, depending on the allocator's mood that day.
Do Go APIs just defensively copy, or pepper APIs with "do not modify" warnings, or rely on runtime tests? Given the fact that append() may mask aliasing bugs, is there a way to make append() alias maximally?
Append will reuse the current backing array when it has enough capacity to hold whatever you're appending. So to alias maximally like you suggested, you could try creating all your arrays/slices with lots of extra capacity. That would be pretty expensive, but you could try doing it one-off just to see what happens in your test suite?
Apart from the issue above I've found that it's usually quite easy to reason about though. Usually when you take a slice you either use it temporarily without modifying or you throw away the original one. In the rare case where you do still need both slices to outlive the same function you just copy one of them manually defensively.
The `append` builtin may allocate a new array and copy, or reuse the same array, but the result should almost always be assigned back to the original variable. The only time you wouldn't assign it back to the same variable is when you are carefully managing the slices yourself for some specific reason, which should never be exposed via an API.
Makes sense and probably is a given for most programming languages.
1. Don't use private attributes for struct types if you ever want to serialize/encode them (exception: Mutex and Channel Attributes). Just had to refactor a whole lot of code just, because I decided I wanted to save the struct to disc. Btw. the type itself can and probably should be private.
2. Don't use Pointers in map key structs. As nesting maps is a pain the simple solution is to use structs as keys. But if you do so please remember not to use any pointer fields within those structs. Again when you encode and decode maps with pointers within the key structs those pointers will bite you.
I often use the "internal" when I have a package with a lot of internal-only marshaling behaviors, so that the godoc for the package isn't made up of 80% internal implementation details and 20% payload.
"As nesting maps is a pain"
Spend a few years in Perl with its autovivication and it'll seem a lot less painful. However, this is just an observation, not an actual recommendation.
for idx, _ := range zoo {
zoo[idx].legs = 999
}
to for idx := range zoo {
zoo[idx].legs = 999
}I hope Go supports "append(nil, data[:2]...)", which will make code much cleaner.
Doesn't mean that "it's more characters!" isn't still a really silly complaint.
$ go test -bench=.
goos: darwin
goarch: amd64
Benchmark_AllocWithMake-4 1000 1626664 ns/op
Benchmark_AllocWithAppend-4 1000 1574720 ns/op $ cat /proc/cpuinfo | grep 'model name' | uniq
model name: Intel(R) Core(TM) i7-5820K CPU @ 3.30GHz
$ go version
go version go1.9.1 windows/amd64
$ go test -bench . --benchtime 10s
goos: windows
goarch: amd64
Benchmark_AllocWithMake-12 20000 870349 ns/op
Benchmark_AllocWithAppend-12 30000 561120 ns/opI want to find/build a linter that catches such cases...
I've never used this construct, but I was unaware of its consequences. You've piqued my interest.
If the capacity is less than the new length, a new backing array will be be allocated. This results in x != y
However, if the capacity is sufficient to contain the new values then x == y.
Normal usage of append is overwriting the initial slice variable so you don't need to worry about this, x = append(x, ...). If you use two variables, as here, you can potentially have two different slices used as though they were equivalent in later code.
http://devs.cloudimmunity.com/gotchas-and-common-mistakes-in...
Else, the way mentioned in the article, or (&article).legs = 99 (ugly) would work. Take your pick.
(&article).legs = 99
This would not work actually because it would only be modifying the copy that exists in the inner loop of the for loop. You'd be taking a pointer to the for loop stack variable and modifying it.If they were already references as you first mention that way simply article.legs = 99 would work yes.
for _, ele := range eles { go func() { // use ele } }
Since ele is a loop variable it is not captured as you might think it is. You need to copy it before the closure.
for _, ele := range eles { go func(ele string) { // use ele }(ele) } for _, elem := range elems { elem := elem; go func() { /* use elem */ }() }https://golang.org/ref/spec#Blocks
The problem you are trying to solve is, that with a statement like
for i := 0; i < n; i++ {
doAThing(i)
}
you want `i` to be valid inside the loop body, in the for-clauses but not outside the for-statement. Go solves this by saying that an if/for/switch statement has an implicit block surrounding them and that block is what scopes the loop variables. AFAIK no one considered, that this would have this consequence in relation to closures.Performance wouldn't really matter, because compilers tend to be pretty good at optimizing these kinds of things. They already reorder when they check for the condition and how they jump non-intuitively and the naive instruction sequence would involve freeing some stack-space at the end of the loop, reallocating it in the next iteration and then writing the new value to it. Figuring out that you can save the actual stack-pointer operations isn't that hard.
for i := 0; i < n; i++ {
i := i
doAThing(i)
}I basically just wanted to point out that the problem isn't so much an implementation question (or about performance). But that no one thought of it and it's now mainly a question of how to best phrase the spec for this :)
At some point I tend to just ignore the second part of the range and only use the index...
and then after that I just switch back to a damn for loop like god intended.