Like others, I've been telling Clojure newbies something like: When you need Clojure STM, default to using an atom unless you really know you need to use something else.
Like others, I've been telling Clojure newbies something like: When you need Clojure STM, default to using an atom unless you really know you need to use something else.
In Haskell I can add a handful of lines to parallelize an "embarrassingly parallel" computation. For example, there are 66,960,965,307 atomic lattices on 6 atoms. A 20 core Mac Studio can figure this out by reverse search in just over four minutes; divvy the work up into piles, and have everyone count the work in their pile. For the problems I care about, everyone's looking for needles in a haystack; they can report what they found with no concern for what anyone else is doing.
So what's the dumbest "try this first" approach in Clojure?
All of these options are, IME, relatively easy to drop in to some extant data processing pipeline to parallelize it, and probably require a similar level of finagling to what you're used to in Haskell.
[0] (->> some-lazy-seq (pmap ...) (reduce ...)) goes pretty far, but nesting/composing pmaps or doing i/o doesn't always work particularly well since the JVM [currently] uses OS threads rather than something like the lightweight threads GHC provides.
[1] https://clojure.org/reference/reducers
For example right now in Clojure, when you have concurrent access to a map, you're forced to choose between either atomic, but entirely serial writes (wrap the map in an atom) or per-key concurrency, but no inter-key atomicity (either use nested atoms or use ConcurrentHashMap).
But this false dilemma has all the hallmarks of complection. It's an all-or-nothing choice brought on by an overly coarse idea of atomicity. You could instead use a map structure built on top of STM to get exactly the amount of atomicity you need. If you need two keys to be modified in the same transaction then they get modified atomically. If you need another key to be modified in parallel, then that can happen. The amount of atomicity you need is specified dynamically and on the fly, instead of bound inextricably to a predetermined choice.
I find myself wishing for this kind of tool whenever I have concurrent write contention on a map. Yes I usually bite the bullet and just accept forced serial writes to an atom, but I do so begrudgingly.
Other ecosystems with more ergonomic STM systems use this to great effect (e.g. Haskell's stm-containers library: https://hackage.haskell.org/package/stm-containers).
I guess it depends upon what you mean by "per key", but if your map is a ref, and the values are refs, you can get per-value concurrency with inter-value atomicity.
You couldn't concurrently add new keys though, since that's changing the map's ref.
I think single atom with serial writes is still better. Firstly, I think it’s more performant still, serial updates on atom are fast. And secondly, it’s much easier to reason about such code, much easier to test and debug it.
> When you need Clojure STM, you usually don't, there is usually another way. But if you really, really do need it, default to using an atom unless you really know you need to use something else.
| | ref | agent | atom | var |
| coordinated | x | | | |
| asynchronous | | x | | |
| retriable | x | | x | |
| thread-local | | | | x |
Turns out I have never needed coordinated, synchronous stuff. I have dabbled with agents, but just for an Advent of Code problem.I do like that 63% (!!) of clojure repos have no mutable references at all - that tracks very strongly with my experience. And that the average number of mutable references is less than 2! Immutability can carry you a long ways, and I love that I can trust that contract. On the other hand, it's nice that I can opt in to mutation really easily if I need it.