The Beauty of Concurrency in Go
pragprog.com
pragprog.com
1. Goroutines are not threads
2. type inference allows you to elide types in var declarations: var host = flag.String(...
3. Go's convention is to use camel case, not underscores.
4. Calling os.Exit all over the place is unusual imo - it may be better to panic().
5. fmt.Fprintf exists os.Stderr.WriteString(fmt.Sprintf...
6. An explanation of why the standard log package isn't suitable would be nice, although I see the format used is slightly different.
7. Ignored errors when writing. Why do you wish to sync at every packet?
8. Massive race condition by re-using b. Line 62 overwrites b. b is then passed to c.logger (68) and c.binary_logger (70) for them to process asynchronously. c.logger is then passed another byte slice (72), which forces it to finish using b. c.binary_logger is not passed anything else, allowing it to delay until the next time something is sent on that channel, which would be after b is overwritten by the next packet. I think that simply moving line 58 to between 61,62 would fix this.
The author has not quite grokked the concept of "don't share memory". b is shared, and undefined behaviour results :(
Doing that would erase the whole benefit of allocating a static array in the first place (you really don't want to be allocating memory for every packet that comes through). The funny thing is the author correctly handles this type of synchronization problem elsewhere, see lines 97-100. The right thing to do is to wait for acknowledgement from both loggers before reusing shared memory, but that sort of defeats his broader argument: if you are explicitly guarding every shared memory access with a barrier, then you might as well have been using any shared memory language besides Go. Buffered channels in both directions used to implement barriers are every bit as error-prone and complex as using a reader-writer lock explicitly.
The benefit is dubious. Performance is unlikely to be important in a single-connection logger. My advice is to lean on the garbage collector. Allocate away.
Buffered channels are certainly more complex, but if you're never sharing memory, shouldn't hurt you.
Edit: Oh, apparently you all mean OS threads. So say so. (For example, in Haskell they're called threads without any implication that each one is an OS thread. Haskell's not unusual that way.)
It's an important distinction between using many (read: hundreds) of OS threads in a single process isn't performant, while using hundreds or even tens of thousands of goroutines is just fine.
Also, you can easily run 10s of thousands of kernel threads (provided you decrease the threads' stack size).
In this case, threads refer to not-a-process units of execution in an operating system. For whatever its worth, Wikipedia defines threads as:
the smallest sequence of programmed instructions that
can be managed independently by an operating system
scheduler.
I think that's also the commonly accepted definition: that threads are related to the OS scheduler.But that's just a detail. In reality, goroutines are threads; they're just userland, non-preemptive threads. Similar constructions have been available, even to C programmers, for well over a decade (and probably much longer).
Of course they aren't called threads there, either (probably for the same reason as they aren't called threads in Go)
Programming language support for threading predates direct operating system support (at least in mainstream operating systems) by a lot of years, from what I can tell.
As I mentioned, knowing the difference is actually a good thing so your post is really helpful. My point is just that the two concepts are so intertwined and overlapping that I think it is silly for someone to correct someone's terminology on this (as Jabbles did, though I admit I share some of his(?) other concerns with the OP), particularly in cases where it is clear the person is pretty familiar with goroutines and how they operate.
However the more salient point here is the OP was either being lazy by saying "goroutines are not threads", or actually doesn't understand the concept of lightweight threads. Along the lines of Rob Pike's own explanation on the matter, he maybe should have said "goroutines provide concurrency but not parallelization".
If they are multiplexed on several threads, then they also provide parallelization.
It is actually common for people to talk about green threads (http://en.wikipedia.org/wiki/Green_threads) in the same vain as OS threads.
They are not kernel-level threads, but they are threads in every other meaningful sense; they have their own stack and execution pointer, and data operations in goroutines are not guaranteed to be atomic with respect to other goroutines.
While the current official Go implementation never preemptively schedules goroutines except on I/O and on runtime.Gosched(), nothing in the spec precludes a different, more kernel-thread-like scheduling system.
http://en.wikipedia.org/wiki/Green_threads http://en.wikipedia.org/wiki/Thread_(computing)#M:N_.28Hybri...
They are just a subclass of threads.
When I read material from them, I would have assumed I'd be seeing "idiomatic code from an expert"...not "hey, I'm a newbie hacking my way around in a new language, here's what I typed out in 2 days of playing with it".
I mean, I understand they're big on the "learning a new language on a regular basis thing", but when I see code published like this, specifically from their brand, I go into it thinking "okay, this should be high quality, idiomatic code."
...guess I'll be more careful with that assumption.
(EDIT: If this is just some personal blog, I guess I'd understand; I saw "magazine" in the URL and got the impression this was published as part of their magazine/PDF series.)
; Rough sketch: def defines a var (pretend it's a reference)
; @ is used to dereference the future and block to wait for the result.
(def f
(future
(Thread/sleep 10000) (println "done") 100))
user=> @f
done
100
;; Dereferencing again will return the already calculated value.
=> @f
100
http://clojuredocs.org/clojure_core/clojure.core/futureEdit: And more importantly, there are wrappers for the standard data structures designed around different concurrency use-cases (sync, async, coordinated, uncoordinated)
Refs are for Coordinated Synchronous access to Many Identities".
Atoms are for Uncoordinated synchronous access to a single Identity.
Agents are for Uncoordinated asynchronous access to a single Identity.
Vars are for thread local isolated identities with a shared default value.
http://stackoverflow.com/questions/9132346/clojure-differenc...
And you can use all (all!) of the Java concurrency tooling as desired, including raw threads (for which Clojure has a wrapper as well).
Part of the reason I use Clojure rather than Go is because it doesn't try to force you into a one-size-fits-all method for handling concurrency. I have no problem with CSP but it doesn't fit everything I do. Sometimes I just want to defer work or wrap it in a future. Or I want to use an intelligent coordinated data structure rather than trying to meld flesh and bone to steel in order to make a concurrency-naive data structure behave how I want in a concurrent environ.
If I can avoid those unnecessary battles, I will.
So - Clojure.
int x = std::async([]()->int{
return 100;
});
std::cout << x.wait() << std::endl:
Which isn't as nice as closure. Go could probably benfit from having a standard futures tool. I guess something along the lines of: type Future {
Wait() interface{}
}
func newFuture(func interface{},
args ...interface{}) Future type future struct {
completed bool
result interface{}
ch <-chan interface{}
}
Edit: Or just check if the channel is closed func (f *future) Wait() interface{} {
v, ok := <- ch
if ok { f.result = v; return v }
else { return f.result }
}It has more than one (you can do erlang-style share-nothing style or Java/C++ style of using mutexes to protect shared state from concurrent access).
I know nothing about Closure so it's possible it has more features but it's not necessarily a good thing. Is the complexity of 4 different solutions worth it? (by "it" I mean: a programmer has to learn all of them and when to use what; the implementor has to implement them; write wrappers for all standard data structures (what about third party libraries?) etc.).
Feature bloat has a cost.
Go gives you all you need to easily write concurrent programs and it does it with refreshingly simple design (both for people to learn and to implement).
Go people would probably object that the things that Go adds, and that node.js lacks, aren't just window dressing -- sure, there are situations where a node.js-style fast event loop that avoids blocking operations is all you need, but there are also situations where you want something more like real threads, because the problem demands it.
I'm a lisper but not a Clojure expert - but I'd assume that Clojure people don't consider the existence of e.g. Actors to be "feature bloat". My impression is more that the difficult/special concurrency-enabling feature of the language is STM, and language-level support for different concurrency paradigms, implemented on top of STM, are probably low-hanging fruit once you've got it.
If I weren't sick, I'd submit a new version that:
1. Didn't reimplement io.Copy
2. Didn't avoid io.TeeReader
3. Didn't do weird things to avoid regular channel ranges.
4. Didn't do non-standard date formatting.
5. Didn't reinvent the log package.
6. Didn't try to convince anyone runtime.GOMAXPROCS(runtime.NumCPU()) was a good idea (it's not)
In fact, maybe I will anyway. brb
So what is an appropriate GOMAXPROCS? As someone who has only dabbled in a few Go tutorials, I would imagine that you would want GOMAXPROCS to be NumCPU() (or even greater) so the goroutine thread pool could "fire on all pistons". Why does Go's scheduler default to GOMAXPROCS=1 instead of NumCPU()?
Do you believe users shouldn't have any control over the number of cores any particular application consumes?
Have you measured the CPU contention of the application and determined that using more cores is worth the overhead of increased overhead of multi-thread exclusions (vs. more simple things happening directly in the scheduler)?
Overall, it has nothing to do with this article and now even more people are going to copy it in more unnecessary places as a cargo-cult "turbo button" for their programs.
If you are going to use an idiom like that, the least you could do is check for the GOMAXPROCS environment variable and only do this as a default when the user hasn't specified otherwise.
At other times you should think about the number of processors you want to occupy. If the objective is to behave like an appliance, then 1:1 schedulers:cpus is not a bad ballpark.
The best number of processes to use is equal to the parallelism of the solution. Even with highly concurrent problems, this is still most often 1. If you get it wrong performance will suffer. But in practical terms we have more to worry about, and if you're talking to the disk and the network more than you're computing, parallelism will only increase the contention on those resources. The extra processes will consume more CPU without doing any more useful work.
So the default is pretty good.
By the way 1:1 isn't the limit either. Sometimes you will want more. If the problem truly is parallel enough to exceed your CPUs, you may want additional processes anyway. This will keep things up to speed thanks to the host's scheduler which is typically preemptive, unlike Go's. This sometimes works much better if you can pick and choose which routines run on which schedulers, and I'm not sure if Go exposes that.
They then went away and implemented their own language with lightweight processes and message passing, but missed the fact that actors are the price you have to pay for the benefits of not sharing mutable data.
And Go completely skipped that part (the most important part).
You don't need to share state data between your goroutines if you don't want to either just like you don't have to use mnesia to share state between erlang processes if you don't want to.
I don't think you can really accuse go of being a cargo cult language either, Rob Pike has implemented CSP multiple times (http://swtch.com/~rsc/thread/).
values := []string{"a", "b", "c"}
for _, v := range values {
go fmt.Println(v)
}
Each of these goroutines shares the same variable v, so this code contains a serious race condition. values := []string{"a", "b", "c"}
for _, v := range values {
go func(){fmt.Println(v);}()
}
Does have a race condition. values := []string{"a", "b", "c"}
for _, v := range values {
go func(s string){fmt.Println(s);}(v)
}
Does not have a race condition.The parent's post has no race condition, as v is evaluated before the goroutine starts.
The top example of my post has a race condition because there is no guarantee when v will be evaluated wrt to the loop.
The bottom example has no race condition because v is evaluated on every iteration and assigned to s, which is used by the goroutine at some point afterwards.
Nope: The original Communicating Sequential Processes model[24] published by Tony Hoare differed from the Actor model because it was based on the parallel composition of a fixed number of sequential processes connected in a fixed topology, and communicating using synchronous message-passing based on process names (see Actor model and process calculi history). Later versions of CSP abandoned communication based on process names in favor of anonymous communication via channels, an approach also used in Milner's work on the Calculus of Communicating Systems and the π-calculus.¹
¹: https://en.wikipedia.org/wiki/Actor_model#Contrast_with_othe...
Let me tell you, that problem is hard. Go coped pretty well, but the final thing is a mess of global states, and it's pretty elegant for what the problem is. I was hoping to avoid having many moving parts, but it ended up needing a lot of shared state between all processes.
Some problems are just hard, and, no matter how well-designed the language is, they'll still be hard. Something I miss from the language after implementing that is the ability to, say, monitor one goroutine from another to see if it returns (so the former can return as well). I know it's possible with channels, but when one goroutine is blocked on network Recv(), there's not much of a chance to listen to channels.
Anyway, yes, parallelism in Go is great, but not everything is magically all rainbows and unicorns (that's Django). Some problems will be messy and dirty, and even more so when you use channels.
Could you elaborate a bit more on this point? This seems like a natural pattern for a channel/goroutine. Spin up a goroutine that reads from a connection and send whatever is read on a channel. Then other goroutines can synchronize on the channel.
The best solution that I've found is to just close the connection, the .Recv() will error out immediately.
I think you're framing the problem wrong. You're not asking to shutdown a goroutine, you're asking "How do I abort from a synchronous read from a network connection that is blocked?"
http://golang.org/ref/spec#Select_statements
http://blog.golang.org/2010/09/go-concurrency-patterns-timin...
http://talks.golang.org/2012/concurrency.slide#1
Video for the slides: http://www.youtube.com/watch?v=f6kdp27TYZs
Does sound like you want to use channels, and break your logic into small independent parts.
You will have one goroutine using blocking Read() in a loop and feeding data to some channel. When it's done, you write to another channel that exists only for signaling:
defer {
doneChannel <- true
}
for {
data := make([]byte, 65535)
_, err := conn.Read(data)
if err != nil {
errorChannel <- err
break;
}
dataChannel <- data
}
and then in your other goroutine: for {
select {
case data := <- dataChannel:
// Handle data
case <- doneChannel:
// Other goroutine is done
}
}
The only things shared here are the channels.The way to avoid too much state and moving parts is to break the problem into isolated, manageable parts that communicate with channels. Often you will have hierarchical relationships like this, where one piece of dumb code exists to pass data from something lower down to somewhere higher up.
It's not hard, although some of the code gets a bit ugly and disjoint at times, especially in how anything synchronous has to use channels and goroutines. For example, today I wrote a simple worker pool implementation that runs a given function in parallel via goroutines and can adjust the number of workers dynamically at runtime. That function has to be declared not just as "func()" but as "func(abortChannel chan bool)", and the worker function has to honour the abort signal when it arrives from the pool. So channels do leak everywhere, even into APIs. (Yeah, I know I can use "chan struct{}" to avoid any storage, but I think "<- true" looks nicer than "<- struct{} {}".)
What is harder is to intelligently handle complex cascading failures. That's what Erlang, with its supervisor tree, is good at. Go's goroutines are "fire and forget" and cannot even be terminated programmatically from elsewhere in the program.
Having cascading failures in Go would be fantastic. The problem, unfortunately, is not very amenable to elegant solutions. I might do a writeup at some point.
I disagree. struct{}{} tells me that the value isn't important. Whenever I use map as a set, rather than a key-value store I use map[string]struct{} (say), rather than map[string]bool. Then I am forced to use the double assignment to check for membership of the set. And that's exactly what I want. I'm able to make my intent more obvious in the code I write. No one will ever look at it and say "but what if it's false?" - I dislike using booleans instead of empty structs in the same way I dislike other C programmers using integers as booleans.
Eh? If you have `map[keyType]bool`, then a key lookup is simply the set membership function. If a key exists, it returns true. Otherwise, false. That certainly doesn't seem analogous to abusing integers as booleans...
And what looks even better:
defer close(doneChannel)
(it's also syntactically correct -- you ca't just have a defer block without a function invocation) select {
case _, _ := <- doneChannel
// Other goroutine is now done
It's so implicit that you pretty much have to add a comment to the effect of "this will trigger when the channel is closed", whereas the "case <- doneChannel" is so obvious it doesn't need explaining.Also, I rather prefer the supervising goroutine to "own" the channel, so it should be the one to close it.
> you ca't just have a defer block without a function invocation
Yeah, I was not thinking Go there for a moment. Should have been "defer func() { doneChannel <- true }".
> the supervising goroutine now looks a bit odd:
This is totally valid: select {
case <-doneChannel:
//I am a person who likes to scan articles, I'm busy and generally make a read now, read later, read never decision. The code from first scan was unreadable, short 1 character variable names, "why is there a hardcoded date marked 2006.01.02-15.04.05 there??", etc.
Readable code takes a little more time - but it's worth it!
Further, the entire example seems contrived. Am I right in thinking that simply firing up wireshark would solve this problem? Why is the author continuing to write something in nearly every language when a tool exists for exactly this purpose, is multi-featured and pluggable?
Even further, message passing! With the rise in recent years of message passing libraries in nearly every language, multi-threaded, distributed applications are becoming trivial to write. I do not see the Go code presented as anything other than messy, I have seen C++ code utilizing message passing libs that are smaller, prettier and infinitely more maintainable - again with no mutex or conditional variables!
If you want to sell me on Go, make the code pretty, and present a USP.
I hope these talks can help sell you on Go http://blog.golang.org/2013/01/two-recent-go-talks.html
To be more explicit:
Short variable names generally reflect the idea that you know what a variable is for just by knowing its type. Thus you have a file named f, a time t, a variadic argument called v. When the type is not enough, a longer name is recommended. Naming things is hard though...
The hardcoded date is a wonderful piece of the time package, which I fully appreciate will look bizarre at first. (And therefore isn't a great thing to use in a first look at Go, without explanation). See the official documentation http://golang.org/pkg/time/#pkg-constants
I really like how Go does message passing, and other than channels being first class types, I don't think that's really the "killer feature" of Go. The killer feature is that goroutines are green threads, scheduled in M:N fashion on to OS threads. This encourages concurrent programming because spinning up a goroutine is comparatively cheap to spinning up an OS thread. It's difficult to do this kind of programming in most other languages (sans Erlang, Rust and Haskell).
Joe Armstrong made this argument years ago. He compared the limited ability to start processes in most languages as being similar to limiting how many objects you could create in your program.
If you want to see examples of concurrent programming in Go, go straight to the source: http://golang.org --- The tour is good, there are some codewalks, talks, articles, etc.
Date is not hardcoded, it's Go's convention of formatting date [1]
That's actually how datetime patterns are defined in Go. I shit you not.
A good illustration that the Go designers didn't think their ideas through. This is a real pain in the butt when you are writing code and regularly commenting in and out sections of code while you are testing things. And every time you do this, you need to remove or restore the imports. And since Go's tooling is nonexistent, there is no IDE to do this automatically for you.
This kind of thing belongs in a compiler plug-in (if it was designed with such a thing in mind, which is not the case for Go), macros (if the languages supports them, ideally the hygienic and statically typed kind) or an external tool, not in the compiler.
Please don't misconstrue disagreement as sloppiness. If you read the mailing list, it's pretty clear they thought it through.
> This is a real pain in the butt when you are writing code and regularly commenting in and out sections of code while you are testing things.
Not for me. I love it, actually.
> And since Go's tooling is nonexistent, there is no IDE to do this automatically for you.
Vim does this for me.
Also, Go has some of the most wonderful tools of any programming language I've ever used.
> This kind of thing belongs in a compiler plug-in (if it was designed with such a thing in mind, which is not the case for Go), macros (if the languages supports them, ideally the hygienic and statically typed kind) or an external tool, not in the compiler.
I think reasonable people can disagree on this point.
A trivial workaround to silence the compiler errors during development is to use blank identifiers (http://golang.org/doc/effective_go.html#blank_unused). I use them all the time.
You may disagree with their decision, but it's disingenuous to claim that it's because they didn't think it through.
It's impractical on many levels. Thus the need for kludgy solutions like blank identifiers. I'd rather see a strict mode or some other type of compiler flag.
Not really, if anything, Go shows that its designers have a lot of inexperience when it comes to modern language design.
Go would have been a kille language in the late 90's but it seems to ignore everything that we've learned about language design in the past decade.
Maybe one could make a table stating the language feature and which language provided it for the first time.
Even if the language is like that, if it helps improving the situation where young developers learn that strong typing does not have anything to do with VMs, I find it quite positive.
How could it be that bad?
It's as simple as a compile, a click on the error message and adding a one line comment.
No worse than the standard practice of making sure your c/c++ code compiles without warnings with the maximimum warning level set.
Development has several modes. One mode is "hacking", just hashing out what you want until it works and is elegant enough as a solution, perhaps changing your mind frequently when you see how it works in practice. Another is "polishing", carefully annotating, cleaning up, documenting, burning off loose threads, making sure the test coverage is top notch, etc.
The problem is that Go's compile-time strictness lends itself to the "polishing" phase, but not to the "hacking" phase.
> Development has several modes. One mode is "hacking",
> just hashing out what you want until it works and is
> elegant enough as a solution, perhaps changing your mind
> frequently when you see how it works in practice.
> Another is "polishing", carefully annotating, cleaning
> up, documenting, burning off loose threads, making sure
> the test coverage is top notch, etc.
>
> The problem is that Go's compile-time strictness lends
> itself to the "polishing" phase, but not to the
> "hacking" phase.
When I write Go, or indeed in any programming language, I generally start with, and stay in, what you call the "polishing" phase. Experimentation occurs in my head, and what makes it through to my fingers is the polished form of that experiment.That Go is not conducive to writing sloppy (or "hacking" phase) code is I think only a good thing.
Anecdotally, I have had colleagues who always plan ahead meticulously, using pen and paper and diagrams and plenty of note-taking before ever writing a single line of code, and the first line of code is often a test. And yet those people were terrible programmers. They take a long time to produce working code, and it's often deeply flawed. They will spend half a day or an entire day trying to hunt down a bug that I found to be trivially obvious even without knowing the codebase. Meticulousness does not imply quality.