Are Go maps sensitive to data races?
dave.cheney.net
dave.cheney.net
Since a concurrent map (aka. Dictionary aka. Hashtable) is a lot more complex and slow than a non concurrent one, it's very very unlikely that Go would ship with its standard map type being a concurrent one. It would be a huge waste.
More likely there is a separate concurrent map type. It's exactly the same in other languages with mutable collections (Java, .NET, C++, ...).
We should develop a system where maps are unsynchronised by default, but when you access them from a second thread they are then converted transparently by the runtime to a thread-safe implementation.
A half-solution could be using very lightweight synchronization (e.g. a MVar containing an unsynchronized map) that's later converted as you describe, but it would still incur some cynchronization overhead (even for single-threaded use).
You could then profile allocation sites so that if a map is frequently converted to concurrent, you then start allocating it as concurrent in the first place.
Although I could imagine a similar "signaling" mechanism, where each thread would periodically read a single snychronized variable, to check if there's any additional "re-synchronisation" work it needs to do. But I'm guessing that CPUs are implemented that way anyways.
No nobody needs to wait for a GC, but we use the same mechanism that the GC does.
> Although I could imagine a similar "signaling" mechanism, where each thread would periodically read a single synchronised variable
And that's how the GC already works. Except instead of a variable normally a 'test' instruction is used on a page of memory, and instead of setting the variable, the permission on the page are changed triggering a page fault.
Which GCs do that? The only one I know is Azul's Pauseless JVM GC, but that's a kind of a special case, given that it needs support from the kernel.
http://dl.acm.org/authorize.cfm?key=N98613
Also note that Azul almost certainly has the same mechanism to do things like stop the world for dynamic class loading, even if it isn't using it for GC when running the pauseless collector.
And Section 4 and 5 of this paper talks a bit more about it's implemented http://chrisseaton.com/rubytruffle/icooolps15-safepoints/saf....
Multithreading isn't something that should be implicit. Either you design your whole module/library/application around to be used from multiple threads (with all the extra work for synchronization and typically also well defined thread boundaries) or you stay with single threading.
The fixes for this are subtle and annoying. Either you have to lock around your data, make it somehow immutable, or force copying rather than referencing (by making it "value" data).
Go would really benefit from a port of Clojure's data structures.
That's called a channel in Go.
Something like the disruptor (https://lmax-exchange.github.io/disruptor/) is more along the lines of what I was thinking of. Without generics support you have to write it over and over again for each use case.
Lines 28, 152, 401
"Getting your locking wrong will corrupt the internal structure of the map."
Ahem.
If you have a non-concurrency-safe map implementation in a language that's advertised as being good for concurrency, there is definitely something wrong.
Do you know what they do? They throw a ConcurrentModificationException.
Do you know what they don't do? Silently corrupt themselves.
So even while Java tries to find concurrency bugs, it does not promise it.
The whole reason there are two implementations of the dictionary (Concurrent and Non concurrent) is that the performance cost of the concurrent implementation is too high to carry for the non-concurrent one. Depending on the cost of checking for concurrent access, a lot of the gain of using a simpler non-concurrent one could be lost.
A check for single thread access (which is even stricter than non-concurrent access which allows several threads as long as they aren't used concurrently) would be pretty cheap: store the thread ID on creation and then verify on reads and writes.
Failing on concurrent access otherwise on all reads and writes would basically mean that you add a lock to the write operation, and fail any reads while anyone is writing. This however is close enough to making it a full concurrent map, so it isn't worth it.
https://github.com/golang/go/commit/50c5042047be3af36e7bb478...
A Java HashMap is not threadsafe and is documented to be so, and can silently corrupt data, can go into an infinite loop etc. if you mess with it from several threads.
Both leaves you with a lot of potential for races in application logic though.
"If you have a non-concurrency-safe map implementation in a language that's advertised as being good for concurrency, there is definitely something wrong".
If we accept his premise, then the fact that "they wasn't designed for that" is no excuse. It's like selling a kids toy that has tiny choke-inducing parts. In that case, just saying: "It wasn't designed to be eaten by kids" is not really an excuse.
Second, that premise is completely untrue. Go, and most other languages give you more-or-less orthogonal primitives and let you combine them. It's okay for your types not to be thread safe, if by combining them with some other primitive you can make them thread safe.
Every other language does this. Some offer you a concurrent map along with a non-concurrent map. For many reasons I won't get into, Go only has one kind of map and lets the user do the rest (which is an extremely easy idiom in Go). Note that even though many languages conveniently offer concurrent maps, no languages with mutable state that I know of offer concurrent integer and concurrent structs. The user is still responsible for ensuring the safety of those.
This is perfectly fine because the grandparent's premise is wrong. In concurrent language, by far the most common case is by data to be owned by a goroutine/thread/whatever. It's very easy to reason about code this way, and it's the way you are encouraged to write code.
Only when you need to share data you need to worry about concurrency-safety, and in that scenario you have available all the tools to ensure it.
Java does for integer http://docs.oracle.com/javase/8/docs/api/java/util/concurren...
You can find similar classes in the java.util.concurrent.atomic package. http://docs.oracle.com/javase/8/docs/api/java/util/concurren...
I'm not sure what requirements you have for a concurrent struct, but Java has classes to atomically manipulate int and long fields of a class.
I'm not arguing your general point (in fact, I agree with it), I'm just supplying some extra information of a language that you apparently don't know.
You could argue that if most Go code is concurrent then the regular collections should be the thread safe ones, and for code that is known to be single threaded there would be simpler/faster non-concurrent collections. That is a completely valid argument to make, but I think that design would be too confusing for new Go developers, since in most other languages the standard collections are the non-concurrent ones, and the concurrent ones are special, and it's an explicit design goal of Go to be easy to adopt coming from another language.
You know, two thirds of the way through the second decade of the 21st century, it's okay to spend a few cycles not corrupting your data structures.
I can agree with an argument that thread-safe should be default and specialized/fast collections should be optional though.
Maps are not concurrency-safe, you know what else isn't? Integers, floats, pointers, structs. And not only in Go, but in pretty much almost all languages with mutable state, like C and C++.
The map implementation in Go is consistent with the rest of the language, plus it allows for great performance under some rather common scenarios.
But I guess it's fashionable to snark on message boards than rather understand all these facts.
In Java there is also the plethora of thread safe data structures available in java.util.concurrent. They are written by experts and are pretty darn fast (TM). It might have been reasonable to use one of these as the basis for Go map implementations, but the language designers chose to assume unsynchronized access to shared maps is a design bug in applications.
Yet manages to be faster than golang.
Also, the two need not be mutually exclusive as you're implying.
Not that I disagree with your comment in general, but atomic values do exist in many languages.
When you use "regular" operations, you don't get atomicity, but what you get is well-defined, and explained here: https://golang.org/ref/mem.
I recommend Go users to install the boot utility, and playing around with Clojure & core.async. These communities can share many things.
For instance, in Go making a blocking system call inside of a goroutine "just works" as Go will create additional goroutines as necessary. But if you do the equivalent in Clojure, you risk thread starvation.
"However if you are using go blocks and blocking IO calls you're in trouble. You will in fact often get worse performance than using threads (in the normal case) since you will quickly hog all the threads in the go block thread pool and block out any other work! ... Since the go block thread pool is quite small, it's easy to block all the threads and thus stopping all 'go processing'."[0]
Now there are some ways around this; I think that you can use promises for blocking IO inside of a Clojure coroutine. But this is one area where you don't have to worry about this kind of thing in Go, even if you do have to know that the default map implementation isn't thread safe. :)
[0] http://martintrojer.github.io/clojure/2013/07/07/coreasync-a...
You need to make sure that once you've done with the item and passed it on, you don't have any overt, tacit, or deeply-buried pointers into the item you just let go of.
Which types in Go are goroutine safe? Do we need to use sync.Mutex for all global variables which are read and written in goroutines? My understanding is only read-only variables are goroutine safe.