Clojure from a Schemer's Perspective
more-magic.net
more-magic.net
Also, the post states that you can't have nested "tail-call optimized" loops in clojure, which is only partially correct: Using nested loops is commonly done in Clojure, the only thing you can't do is have the loops call each other mutually-recursively, which is a far less common use case (but still a feature I hope to see in Clojure some day).
Thanks for the kind words. I tried to keep my biases out of it as much as possible and tried to avoid making it too ranty. And on the whole, I like Clojure; it's just not really my first choice as far as Lisps go.
> It's a bit surprising they didn't get into the whole "hygenic macro" business, which seems like the most obvious differentiator between the two languages to me.
I debated including a bit about macros, but in the end decided that this would be a bit too in-depth. Clojure uses namespaces in a somewhat clever way to work around the bulk of the hygiene issues you get in CL with defmacro, but the macro system itself is rather uninteresting. I think their custom syntax for gensym is a nice touch, even though it's yet more syntax.
Scheme's macro system is very advanced and relatively complicated indeed, and the R5RS/R7RS standard system is only pattern matching/rewriting based.
Maybe I'll do a separate piece on the macro system if people are really interested.
But unrelated to that, I wanted to comment on this observation:
> In my current project, it takes almost 30 seconds to boot up a development REPL. Half a minute!
I work mostly with JVM languages and the JVM startup nowadays is actually very fast for a VM, under 100ms for your code in main to be running from a cold start. 30 seconds to boot is completely unheard of in my experience. If whatever you're using is open source, I wouldn't mind having a look to figure out what the problem may be.
I also shaved 35% of my "time to CIDER REPL" by passing a few parameters to the JVM (-client, -Xverify:none, plus two tieredcompilation-related settings I copy/pasted).
... Yet ;)
I think reading your article, familiarity seemed to be the source of a lot of your pain points. I wonder if the more you use it, to the point it becomes as familiar as Scheme, if you'd change your choice. Maybe not, but I think that's a possibility.
I've wrapped my head around the CL/Clojure sort of macro but never really grasped the essence of Scheme's hygenic system, always happy to read more about Lisp macros.
I'd like to throw in a "me too" too.
I'm comfortable enough using Scheme's syntax-rules macros, and I've seen a couple examples where syntax-case macros do something beyond, but I'd love a walkthrough of how syntax-rules maintains hygiene and a concise explanation of how Clojure nearly accomplishes the same goal with namespaces.
I guess you are referring to recur (https://clojuredocs.org/clojure.core/recur), and I also think that it's a quite cheap way to get ~90% of the benefit of TCO without actually implementing TCO. It's also more explicit than TCO which is a plus over TCO.
For mutual recursions there is always trampoline (https://clojuredocs.org/clojure.core/trampoline), FWIW.
This is not a rare situation. Am I missing something about how Clojure does it?
My Lisp dialect (Cant) has a construct like the named LET, but with the name being optional (defaulting to 'loop'). This makes simple loops as concise as Clojure's, and fancier ones as flexible as Scheme's.
Slight typo but it's "hygienic", not "hygenic" AFAIK.
This is the biggest criticism of clojure I agree with the author on. Trying to compose threading macros is extremely annoying even though it’s not that difficult to refactor. Somewhere in the thread, you’ll want to put the return value of the previous function in a different argument place. I am not sure the reasoning behind designing the macros this way. With how well the language is designed, this annoyance has always seemed out of place and somewhat of an after thought. I made my own macro [0] as an experiment to circumvent this annoyance... which was an interesting experience (both writing and using the macro). Still, I don’t think it’s a good enough solution so I’m still trying to think of better ways to implement thread macros in clojure because I love the idea.
Functions that take a map should take it as first argument, functions that take collections should take the collection as last argument. This is how the clojure core functions work, eg. -> `assoc`, `conj`, `dissoc` vs ->> `map`, `filter`, `reduce`, `some`.
If you follow this rule in your code, threading is nice and looks good. If you want to force different argument position, use a lambda, like so (->> coll (#(my-fn %))).
However, I recommend against it. Use a binding and then use the other threading macro if necessary.
(-<> x (foo bar) (map baz <>))
Found here: https://github.com/rplevy/swiss-arrows
One of the best tech talks I've ever seen.
I haven't had the pleasure of using Clojure in production, but I'm somewhat obsessed with the language and follow it fairly closely. Someone on r/Clojure[1] mentioned Specter[2] with a link to this video[3]. One of the first things the video covers is how the library deals with this concern.
[1] https://www.reddit.com/r/Clojure/comments/loz77v/just_came_a...
Specter just does this for you, under the hood. In vanilla Clojure, the natural way to transform some nested data might be something like update-in and possibly some nested transformation functions and maintaining the types by hand can become quite cumbersome. In specter, you don't have to think about it. Its neat.
But more generally lazy sequences (and their prominent/default usage in builtins) seem to make code execution less predictable without delivering big benefits. One of the less successful features in Clojure IMO.
Of course, that’s not everything specter can do. I’d say it’s probably the least interesting of specter’s features. Much more interesting is how powerful the path selector system is. For example, you have a bested structure, let’s say a vector of maps and one key of the map is a vector of numbers. In spectre you could get at all of the numbers in all of the maps, take just the odd ones, sort them, reverse and put them back into the various different maps in the vector. That is incredibly hard to do by hand!
Ie: [{:a [1 2 3 4]} {:a [5 6 7]}] => [{:a [7 2 5 4]} {:a [3 6 1]}]
Regarding lazy sequences. I agree that they’re not as useful as perhaps touted. I’m not sure I’d call them failed, but definitely they didn’t turn out quite as useful as maybe was expected. Having said that, they do have their moments where they are super useful. The question is whether they could have been made optional, so you use them only when needed... Also, (map foo (vec whatever)) will still lazily map over whatever even though it’s a vector, so devolving into a lazy sequence isn’t a necessity to enable laziness, but rather it’s just that Clojurescript sequence functions don’t maintain the type. I think Clojure‘s sequence functions should (by default, perhaps with a way to turn on current behaviour instead if there’s a performance or other reason) maintain the input type. Even if internally it works on a lazy seq, it should check the type of input and make sure the output is the same type. conj is generic (works for lots of types) and maintains its input type, map, filter etc could too. This would make them behave more like you would expect coming from another language without getting rid of laziness in any way.
Ideally this would be done without converting to and from lazy sequences internally, but if that’s necessary then a way to turn it off for performance would be good, or a way to fuse multiple operations together (via transducers maybe?). I can see how this idea would add complexity if you want the best performance and perhaps that’s why it was decided to just return lazy sequences and metro the programmer decide what to do. Clojure often does sell itself as a language for professionals and avoids doing things just because it’s easier for beginners. Still...
IIRC the necessity of loop/recur for TCO is a result of being hosted on the JVM and not a direct design decision for Clojure.
If I really wanted to play devil's advocate there is an argument to be made that needing to use loop/recur could make an novice programmer more conscientious about tail calls (since it won't work if the recur isn't actually in tail position), whereas in Scheme you get TCO for free, but might not be aware when you're getting it or not.
In any case I kind of enjoy the loop/recur syntax since I don't have to name an intermediate tail-recursive function in a let (this could be naivety WRT writing scheme on my part, but I assume that's what usually happens).
The fact loop/recur cannot be nested is something that I've run into a few times and is a fair criticism. I also find that I end up converting anonymous lambdas to fn all the time, and I'm not the biggest fan.
IMO the most ergonomic lambda syntax I've worked with is Scala. I miss it in every other language that has lambdas.
and from the article:
> I really started noticing mistakes that make additional features appear necessary: for example, there's a special macro called loop to make tail recursive calls. This uses a keyword recur to call back into the loop.
While it may be necessary due to the JVM, having them was also a conscious design choice and many people (myself included) actually prefer being explicit about TCO, because normally when its used, its not used as an optimization but as a required feature to make the semantics work correctly (eg a recursion-based loop), with "recur", the intent that it be tail-call "optimized" is clear and if it cannot be tail-call optimized, then its an error (recur won't compile if its in a non-tail position). So, I personally would use recur even if normal recursion did have TCO.
Also, TCO is more general than "recur" and I believe having recur was also somewhat meant to make the semantics clear. For example, recur can only recurse to the closest loop or function. Proper TCO can also handle cases like mutual recursion between two or more functions (ie f1 calls f2 which calls f1. If both calls are tail calls, then TCO can be used here. This is actually what the JVM has problems with: making TCO work only for the cases that recur covers is easily possible on JVM, but having recur makes it clear -- at least to me -- that its a specific feature with specific semantics rather than general TCO).
So anyway, OP calls it a mistake, I call it a feature that I like. There are a few other places where OP says something is unnecessary or wrong or a mistake, but me, someone who has used Common Lisp and Scheme, but never in earnest, but has been using Clojure for ten years, I find them to be nice features, or make the code look cleaner to my eyes, or are otherwise good design decisions to me.
This is a good article, highlighting a couple of the warts, which I think is a missing element for most people who like a language.
I completely agree with the nil punning observation. One of the issues pointed out in the article that resonated most with me is functions in the standard library usually (but not always) nil pun. You get used to nil, so when the standard library does throw a runtime exception, you're surprised. For example:
(:foo nil) => nil
(int nil) => exception
Overall great article.
As in, Clojure uses int rather than Integer, so it simply can't be nil aka null?
(some-> nil :foo)
(some-> nil int)
Of course this fails if your nil is actually a string, but that's a different issue.
The biggest area where the JVM leaks through is exceptions and error messages.
With that said, you can still use React hooks via Reagent if you want to, it's described over here: https://github.com/reagent-project/reagent/blob/master/doc/R...
Personally, I use reagent (with re-frame) and the only times I need to do JS interop is when using JS libraries (which shadow-cljs has made easy -- I know that Clojurescript has better support now too, but I've not used it because I was already a shadow-cljs user).
The biggest advantage of Clojure over more succinct Lisps is that the extra verbage and built-ins promote a larger shared base and patterns on how things should be done. This is less likely to turn each project into its own language. Clojure skills are much more likely to transfer between companies.
The other thing about being in the JVM ecosystem is that it promotes sharing so there's less reinventing the parts of the wheel that you need for 'this' project that's unsuitable for other use-cases.
But again, that's just me.
It's true that you can extend a vector on either end
user=> (def xs [2 3])
#'user/xs
user=> (conj xs 4)
[2 3 4]
user=> (into [1] xs)
[1 2 3]
but only the first operation will be O(1). The second operation is O(n) and should be avoided in hot loops.My comment was intended to say that Clojure vectors have some operations that are performant and some that aren't. It is part of Clojure's philosophy to encourage the operations that are natural for the respective data type. `conj` appends to the end when applied to a vector but inserts at the head when applied to a list.
So you want to pick the appropriate data structure for the job and make sure only natural operations are used. For a vector appending is natural whereas inserting at the head is not.
I've actually been curious about this myself. No one seems to scream that functional languages are slow, so I'm curious what's going on under the hood. Like the author, I would presume they don't fully copy the hash table every time I add an entry.
Is Clojure smart enough to know that the previous hash map is now out of scope and can be GCed, so it just mutates the hash map in place under the hood? In the tiny amount of Clojure I've written, that seems possible. On the other hand, it makes performance much harder to reason about since mutating a variable that stays in scope after the mutation presumably would still trigger a copy. Do they just implement some kind of a "fall-through"? I.e. if I add a key to a hashmap, it creates a new empty hashmap, adds that value and a reference to the original hashmap. And then when I retrieve values, it searches the new empty hashmap, and then the "parent" hashmap if it isn't found?
I'm curious because a lot of functional programming seems like the kind of thing that would thrash memory in a GCed language. It seems to an outsider like you can't do anything without allocating memory, which you're often done using almost immediately. It would seem like "y = x + 1" would always take more time than "x = x + 1" because of the allocation.
Maybe GC is just a lot better than I think. Or maybe the functional style, with it's less complicated scoping, makes GC trivial enough that it offsets the additional allocations. Does anyone have any idea, or maybe have a link handy? I'm not a language designer, nor a Java programmer, so I fear the source code may not be terribly useful to me. I'm also a terrible Clojure programmer, if Clojure is written in Clojure these days (although I'd like to get better one day, it seems like a really fun language).
For example, if you have a vector of 100 items, and you "mutate" that by adding an item (actually creating a new vector), the language doesn't allocate a new 101-length vector. Instead, we can take advantage of the assumption of immutability to "share structure" between both vectors, and just allocate a new vector with two items (the new item, and a link to the old vector.) The same kind of idea can be used to share structure in associative data structures like hash-maps.
I'm no expert on this, so my explanation is pretty anemic and probably somewhat wrong. If you're curious, the book "Purely Functional Data Structures" [0] covers these concepts in concrete detail.
[0]: https://www.amazon.com/Purely-Functional-Data-Structures-Oka...
https://en.wikipedia.org/wiki/Persistent_data_structure
Maybe GC is just a lot better than I think.
Yes, Clojure relies heavily on the JVM's GC being very good. It trashes it like there is no tomorrow and it would be a lot of work and extremely hard for an implementation of Clojure from scratch to match the performance of Clojure in the JVM because of how good the JVM's GC is.
Having said that, I have move 3 projects (10k-20k LoC) to JS from Clojure and don't plan creating new ones in Clojure, the JS projects ended up being faster, shorter and easier to understand. Idiomatic Clojure is very slow, as soon as you want to squeeze any little performance out it your code base will get ugly really fast. Learning Clojure is nice for the insights but I'll will pick nodejs first any day for new projects. Even if I need the JVM, my first choice probably will be Kotlin and then Clojure.
Of course there are many more downsides to using Clojure. No ecosystem, the cognitive overhead of doing interop with over-abstracted over-engineered Java libraries(because of no ecosystem ;)), the horrible startup times, the cultist community and the interop is really not that good, sometimes you have to write a Java wrapper over the Java lib to make it usable from Clojure. The benefits over JS are minimal but the overhead and downsides are too much. Worth learning it but not worth using it for real production projects.
There are a few large projects written in Clojure. I'd say it's a "data first" approach, not uniquely "map first".
There's this famous 4 minutes rant by Rich Hickey (the creator of Clojure) where he talks about maps and about how Java does OOP (and how maps are served through getters but you can't see easily that they're maps), it's golden:
My point was that without having a schema (which is essentially what a case class or ADT is) it makes to very hard to understand what objects are being passed around.
That’s not how Java or c# is coded anymore.
I heavily use destructuring so that functions access the bits they care about, threading macros like (-> stuff (assoc :foo x) ...) and spec to make sure the maps have the structure and keys I expect and rely on non-existent keys returning nil by default. I also use the sequence functions heavily. Mapping, filtering and reducing over vectors of maps, group-by if I need them grouped and so on.
I personally find it to be a very pleasant way to program.
Also good design and timely refactoring of said data structures to keep your functions from having to reach far into data.
You also decide what part of your tree should live in a database, supported by a query language. You can use SQL-based thins or Datomic and the many open source implementations of the same model. In the latter case it feels more like a "turtles all the way down" kind of thing.
Aside from the rest of the article, I find this one weird.
Of course you can do that, you just supply a function which has whatever you need in it. That's how you'd ideally be operating in Scheme as well.
-> is effectively _function_ chaining, not "do a bunch of things in this block" chaining.
The problem is any sizable Lisp needs quite a lot of stuff available at runtime. A trivial program would still likely[0] be a sizable blob with a memory penalty.
If anything, I'd say Clojure (cljs especially) actually bypasses this problem by leveraging a platform's privileged language. Purely theoretically, there's no reason why you couldn't have a mobile device built around CL. But if you want to target what's already out there, a parasitic design works much better.
Also, on the popularity front I guess it's a bit of a glass half full situation. Compared to other Lisps, Clojure is no doubt dominating. Among what I'd call "alternative languages" it's a tough call.
[0] Have to mention there's a lot of diversity in CL implementations so this is certainly not universally true. But I would assume this is true for something that has all the conveniences you'd want.
Even disregarding the perf problems the scale of implementation effort and resulting wasm code size would be prohibitive. Not to mention losing easy low friction interop with JS ecosystem libraries.
Moreover an immutable language like Clojure doesn't necessarily need an GC - one could probably leverage ARC to great effect.
I don't think JS will rule the roost a decade from now. It's simply too slow and most folks actively dislike the ecosystem of NPM hell.
This isn't so much a language designer choice as a consequence of targeting the JVM. Tail recursion on the JVM has been a problem for decades.