Clojure from a Schemer's perspective (2021)
more-magic.net
more-magic.net
I don't think that's true. A persistent array map is a single data structure. It's just a persistent data structure optimized for non-destructive (functional) updates. On a high-level it's not too different from finger trees which may be more famous. When I used to write Clojure, I didn't feel any need to reason about its internals for performance reasons or any other reasons. It simply feels natural.
So like you say, your experience as an "end user" makes it seem like it's just one data structure at least in interface, though in implementation it's actually doing some performance optimization stuff which, I think I'm willing to grudgingly grant to the author, does make performance harder to reason about in principle. Though in my (very unserious) experience, I think I've enjoyed languages that take a possibly incomplete stab at getting performance right over ones that draw a line leaving it entirely out of scope and declare victory.
I'll equivocate a little and share that I have been bitten by these internals before, though the fault was mine: one time, I was unwittingly relying on an implementation detail of PersistentArrayMaps, which is that they preserve insertion order. PersistentHashMaps don't, and my code was mysteriously breaking when number of keys grew beyond 16 until several hours later I learned much of what I've just shared above.
[1]: https://www.javadoc.io/static/org.clojure/clojure/1.10.1/clo...
[2]: https://stackoverflow.com/questions/16516176/what-is-the-dif...
[3]: https://github.com/clojure/clojure/blob/5ffe3833508495ca7c63...
the thing is, you don't reason about performance. you measure it.
Once you've obtained measurements, you can make decisions based on the outcome of said measurements.
If the Runtime (and not just talking about clojure here) makes a decision on performance, and it's suboptimal, your measurements _should_ reveal it - for example, by increasing your input size, the performance graph should show a clear inflexion of where this happens.
The real question is whether you _can_ force the Runtime to do something at the programmer's behest, rather than a hidden decision that you cannot control or change. I believe in clojure, you can change the backing data structure, but i'm not sure if you could do it after the fact?
You that "fun person" at parties right ? :)
This is only half-true. You do reason about it and later profile. For example, I want to know if my assinment operator in C++ will run in constant time.
In example, in C++ std::maps are implemented as RB trees, are::vectors instantiate 50% (sometimes 100%) more memory than requested, etc. All of that happens under the bonnet in order to fulfill the performance guarantees of the language.
IIRC, one of the most widely quoted memes in Clojure space goes along the lines of "simple != easy". Or better (not= :simple :easy)
And then some clever C++ developers that have code that depends on rb tree's behaviours, would cry that their cheese was moved and their code doesn't work as before, even though nothing prevented funky zombie trees to be shipped instead.
At a high level there's really just a few things you need to keep in mind, the biggest being the big O guarantees of the core DS interfaces, which are listed here https://clojure.org/reference/data_structures. Those guarantees hold across the different underlying implementations, and there are plenty of things you can tune to get a 2x, 10x, 50x improvement in performance before resorting to things that depend on which specific DS is being used under the hood
Edit: Fixed timestamp for video link
I first tried todo gerbil scheme, then Ubuntu seems to have the wrong version of gambit (gerbil-scheme runs optop of gambit-scheme for those that don't know). That took a few hours of compiling, reading, trying diff versions always ended in some deadlock-issue ??
Then I tried Guile,was a little better and easier, but soon got stumped when I wanted to access mysql-db yes there are packages but those packages I couldn't install for the life of me !
In the end chicken-scheme was surprisingly easy and complete solution for a hit-the-ground-running scheme.
Now I'm prob a noob as it comes to scheme, but not to fiddle and massage packages and programs on linux. Trying to learn scheme but spending a few hours just to get it installed was what killed it for me.
Clojure is certainly not perfect, but you could get up and running much faster in my opinion. Oh and it uses square brackets for parameters :)
not an expert but i believe you should use guix for anything guile related
> Oh and it uses square brackets for parameters :)
for many lispers this is a negative because they see as a big positive the manimal (no) syntax value proposition of lisp. besides if you use emacs you can easily (if you are a lisper) make it do something like having parameters highlighted without introducing additional syntax in the underlying language
At least it didn't go the way beautiful Standard ML went!
It blows my mind how under appreciated MLs are.
But Standard ML is the cleanest of the bunch, and creators kind of suppressed its development until it faded away completely.
But if you want a "download, unpack it, run" Scheme, you go with Racket [1]. Tons of documentation, a good-enough JIT compiler, and usually everything works. You get an amazing IDE with it (DrRacket) with many advanced features yet to be found in modern IDEs for popular languages.
> Oh and it uses square brackets for parameters
Most modern Scheme implementations make square brackets and parenthesis equivalent. E.g.
(let ((a 1)
(b 2)) ...)
;; is the same as
(let ([a 1]
[b 2]) ...)
[1] https://racket-lang.org/It takes literally just a few minutes to download, install, and get going with Racket, Elixir, and F#. It’s the main reason I have not got into Clojure yet, despite trying.
The “best” advice I have for Clojure beginners is to follow this guide: https://calva.io/get-started-with-clojure/, which will ultimately land you in a solid VSCode-based IDE environment for Clojure.
That’s not how I personally like to approach a new language mind you (REPL from the command line plz), but I’ve pretty much given up trying to get Clojure beginners started there as there are just too many moving parts that can go wrong, and unjustifiable frictions.
Just install it with your distro's package manager.
I have no idea why the official instructions only involve brew (wtf?) but it's in the repo of most distros now.
The other upside of using chicken is that their FFI interop with C is really streamlined. From one of my projects; in order to expose the `getpass` C stdlib function in my program I only had to write the following line
(bind "char* getpass(const char *prompt);")It's been a while since I last used Clojure, so I don't know if the situation has improved in the meantime.
(defn foo [a b]
(def a a)
(def b b)
(/ a b))
If the function fails, I have captured its input parameters right before the exception occurred and now I can inspect it from the REPL if 'b' is 0. From the REPL I can change the value of 'b' with simple `(def b 42)` or change the very definition of 'foo' and re-evaluate it again on the REPL with simple `(foo a b)`.I even made myself an elisp function for inserting the '(def a a)' into the code https://github.com/Bost/corona_cases/blob/master/.dir-locals... (enjoy).
Moreover, this approach with defs allows you to inspect every function, even those macro generated. And if committed, such a "debug session" is persistent across every instance where the code is deployed, from any development machine, through all test and staging machines right into production. (So when combined with some remote REPL access... you see what I mean?)
Another thing is that I like when my code strongly express the notion of compositionality, i.e. when it looks something like this:
(comp
...
(partial map ...)
foo3
(partial reduce ...)
(fn [p] (def s3 p) p)
(partial map ...)
(fn [p] (def s2 p) p)
(partial apply ...)
(partial map ...)
foo2
(fn [p] (def s1 p) p)
foo1
(partial apply ...)
(partial map ...))
here I introduce and comment in/out the '(fn [p] ...)' expressions as needed. That kills two birds with one stone. 's1' captures the output of 'foo1' which is also the input of 'foo2' at the same time.informal coffee break answer: abstracted layers of abstracted abstraction within parametrized abstractions just don't feel right to lispers.
My browser is cpp program. Say, 10 layers are typical. Then, there's a libc + ui linux code. Usually in c, so even less layers. Say, 10 more. Then, there are sycalls annoying the kernel from time to time. Pretty shallow, btw, 5.
That's 30 roughly speaking. Maybe 40.
Java + clojure combo is not impressed, not at all.
But if you're already familiar with any of those, it'll be the same as what you're used too.
- Language-exclusive build chain: Not that Make is the be-all and end-all of build tools but it worked with everything without prejudice. To understand the basics of a build system, you only needed to know one tool. All attempts at fixing Make's flaws fail to understand this fundamental recipe for success. Now when you switch languages, you need to relearn that layer on top of the language itself. And the tools all do the same thing but with their own syntax, commands, and eccentricities.
- Language-exclusive packages and repositories. Once upon a time, you learned your OS's package manager and you were done. Now you need to learn a new package system for every language out there. Again with their own syntax, commands, versioning, and eccentricities.
- Needing language-aware parsers to determine dependencies. You can't depend on spatial locality or compiler-output dependency trees to know whether to rebuild, reinstall, etc. You had to parse out the org.whatever's and resolve that back to a jar or class, which was clumsy and mostly wrong.
- The OS doesn't know how to run your program. What's so hard about building your executable so it can be natively run by the OS? Everyone has to write the same `#!/bin/sh exec java -jar ...` wrapper script. Java programs worked best by assuming it was going to be a pariah hanging out in /opt and never actually integrate with anything.
- Runtime versioning and installation was painful. I disliked having to suddenly manage environment variables just to run something. It continues today with python2 vs python3, GOPATH, etc. On top of that, JRE's weren't a simple "curl|bash" and no one was allowed to repackage the JRE to make it more user friendly.
> Language-exclusive build chain
I'm not sure I agree here. Bazel / Make / Ninja still exist. Most of the tooling that exists extant to that usually leverage some knowledge of the conventions of both the community as well as actual knowledge from the code itself (Cargo for example understands the module system in Rust). You kind of touch on this by mentioning "language-aware parsers" for dependency management, but this existed in C and C++ long before the JVM came along, so it's weird to place the blame there.
Make has many things wrong with it but I wouldn't say that a "one build tool to rule them all" is something it did well. Even with stuff like CMake or GNU Autotools the final makefiles are almost exclusively used for C projects. I rarely see Make used in any other context, which I'm not convinced is a machination of language-specific tooling rather than some intrinsic property of Make itself.
> Language-exclusive packages and repositories
The proliferation of packaging systems is frustrating, but would you rather learn pip (PyPI) / cargo (crates.io) / npm (npmjs.com) across several OS's and platforms or learn 5 different OS / platform package managers? I know for certain there's nothing like that on Windows!
This isn't the JVM's fault, and I think you're looking at it in the wrong order here. These language-specific package managers and repos are made to make the language run consistently in different environments. Most repository maintainers do not and cannot have the bandwidth to package everything under the sun for every programming language.
> The OS doesn't know how to run your program
Oof. The program you're running is the JVM, which as an input is taking a jar file and as output does some work. The OS knows how to run that just fine. Doesn't this criticism basically disqualify every single useful program that isn't a compiled C binary on Linux? ELF isn't that hard but also I'm not sure we need to tie all our boats to that sail either.
If I'm trying to be charitable, I think you're experiencing a lot of frustration with your tooling and expecting something out of the JVM and programming in general that isn't necessarily true. You may just want to stick to C / C++ and tooling built 40+ years ago, but dismissing most of the tools that people use every day seems a bit "cut off your nose to spite your face."
¯\_(ツ)_/¯
There is this very common reaction to Clojure (also reflected throughout the original post): a balking at its pragmatic decisions. But Clojure's pragmatism is precisely what makes so significant and wonderful.
On multiple measures Clojure is the most used Lisp, while also being the youngest, which I think speak a lot to the adoption benefits of being on the JVM and on JS.
Maybe it'll bite it long term, hard to say, but it definitely allowed it to quickly catch up to CL and Scheme and pass them in popularity and adoption.